{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### American Express Default Prediction\n\n#### Competition Objective\nAmerican Express is a globally integrated payments company and the largest payment card issuer in the world. They provide customers with access to products, insights, and experiences that enrich lives and support business success.\n\nCredit default prediction is central to managing risk in a consumer lending business. The objective of this competition is to **predict the probability that a customer will not pay back their credit card balance in the future**, based on their monthly customer profile.\n\n---\n\n#### Data Overview\n- The **target** is a binary variable, defined by observing an 18-month performance window after the latest credit card statement.  \n- A **default event** occurs if the customer does **not pay the due amount within 120 days** after their latest statement date.\n\n- The dataset contains **aggregated customer profile features** for each statement date.  \n- Features are **anonymized and normalized** and fall into the following categories:\n\n| Prefix | Feature Type        |\n|--------|------------------|\n| D_*    | Delinquency variables |\n| S_*    | Spend variables      |\n| P_*    | Payment variables    |\n| B_*    | Balance variables    |\n| R_*    | Risk variables       |\n\n- **Categorical features** include:\n`B_30, B_38, D_63, D_64, D_66, D_68, D_114, D_116, D_117, D_120, D_126`","metadata":{"_uuid":"4506dcab-64a0-41eb-9340-3d3c8b2f7062","_cell_guid":"f146f6e2-5272-4b7c-a6b3-c672d5369743","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"#  0. Libraries & Global Settings\nThis section imports essential libraries and sets reproducibility options.","metadata":{"_uuid":"14c2c895-131a-4f20-92b7-94942fb6a702","_cell_guid":"f3a851d4-5035-4c44-9def-4fe5705aa682","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 0) Setup\n# ============================================================\nimport os, gc, warnings, time, math\nwarnings.filterwarnings(\"ignore\")\n\nimport numpy as np\nimport pandas as pd\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport plotly.graph_objects as go\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\n\nimport lightgbm as lgb\nimport xgboost as xgb\n\nSEED = 42\nnp.random.seed(SEED)\n\npd.set_option(\"display.max_columns\", None)\npd.set_option(\"display.width\", 1200)\nsns.set_style(\"whitegrid\")","metadata":{"_uuid":"2a5a0c8b-25a0-48b1-8368-d51459fc9bff","_cell_guid":"e6769698-6146-4302-8836-d63e40745328","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T15:52:04.764054Z","iopub.execute_input":"2026-06-05T15:52:04.764336Z","iopub.status.idle":"2026-06-05T15:52:04.76962Z","shell.execute_reply.started":"2026-06-05T15:52:04.764298Z","shell.execute_reply":"2026-06-05T15:52:04.768806Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 1) Paths + Columns\n# ============================================================\nTRAIN_PARQUET = \"/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet\"\nTEST_PARQUET  = \"/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet\"\nLABELS_CSV    = \"/kaggle/input/amex-default-prediction/train_labels.csv\"\n\nID_COL   = \"customer_ID\"\nDATE_COL = \"Statement_Date\"\nTARGET   = \"target\"\n\nCAT_FEATURES = [\n    \"Balance_30\", \"Balance_38\", \"Delinquency_63\", \"Delinquency_64\",\n    \"Delinquency_66\", \"Delinquency_68\", \"Delinquency_114\",\n    \"Delinquency_116\", \"Delinquency_117\", \"Delinquency_120\", \"Delinquency_126\"\n]","metadata":{"_uuid":"e9ed0e5d-780c-4142-b27c-8b79c62ce851","_cell_guid":"8862b391-b6d7-46ce-a305-cefb1fd84de8","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T15:52:04.77051Z","iopub.execute_input":"2026-06-05T15:52:04.771046Z","iopub.status.idle":"2026-06-05T15:52:04.781995Z","shell.execute_reply.started":"2026-06-05T15:52:04.771013Z","shell.execute_reply":"2026-06-05T15:52:04.781137Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#  1. Load Raw Data & Rename Columns\n\nAMEX features use prefixes like `D_`, `S_`, `P_`, `B_`, etc.\nWe rename them to human‑readable groups:\n\n| Prefix | New Name | Meaning |\n|--------|----------|---------|\n| D_ | Delinquency_ | delinquency metrics |\n| S_ | Spend_       | spending patterns |\n| P_ | Payment_     | payments |\n| B_ | Balance_     | balances |\n| R_ | Risk_        | risk indicators |\n\n**Insights:**\n- Renaming helps when using deep learning later (TFT/Informer) because they require grouped variable understanding.\n- Sorting by `(customer_ID, Statement_Date)` is *essential* to preserve sequence order.","metadata":{"_uuid":"d79afdce-2076-4f61-abfa-d25a566e1f92","_cell_guid":"09a6d705-323d-45ad-bd92-2c460835abd5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 2) Load + Robust rename + Date handling\n# ============================================================\ndef rename_columns(df: pd.DataFrame) -> pd.DataFrame:\n    \"\"\"Robust renaming: map S_2 -> Statement_Date before prefix renames.\"\"\"\n    df = df.copy()\n\n    if \"S_2\" in df.columns:\n        df = df.rename(columns={\"S_2\": DATE_COL})\n\n    df.columns = df.columns.str.replace(r\"^D_\", \"Delinquency_\", regex=True)\n    df.columns = df.columns.str.replace(r\"^S_\", \"Spend_\",       regex=True)\n    df.columns = df.columns.str.replace(r\"^P_\", \"Payment_\",     regex=True)\n    df.columns = df.columns.str.replace(r\"^B_\", \"Balance_\",     regex=True)\n    df.columns = df.columns.str.replace(r\"^R_\", \"Risk_\",        regex=True)\n\n    # Backward-compatible\n    if \"Spend_2\" in df.columns and DATE_COL not in df.columns:\n        df = df.rename(columns={\"Spend_2\": DATE_COL})\n\n    return df\n\ndef ensure_datetime_inplace(df: pd.DataFrame, date_col=DATE_COL) -> None:\n    if date_col in df.columns:\n        df[date_col] = pd.to_datetime(df[date_col], errors=\"coerce\")\n\ntrain_raw = pd.read_parquet(TRAIN_PARQUET)\ntest_raw  = pd.read_parquet(TEST_PARQUET)\nlabels    = pd.read_csv(LABELS_CSV)\n\ntrain_raw = rename_columns(train_raw)\ntest_raw  = rename_columns(test_raw)\n\nensure_datetime_inplace(train_raw)\nensure_datetime_inplace(test_raw)\n\ntrain = train_raw.merge(labels, on=ID_COL, how=\"left\")\ntrain = train.sort_values([ID_COL, DATE_COL], kind=\"mergesort\")\ntest  = test_raw.sort_values([ID_COL, DATE_COL], kind=\"mergesort\")\n\ndel train_raw, test_raw\ngc.collect()\n\nprint(\"train shape:\", train.shape, \"| test shape:\", test.shape)\nprint(\"date null pct:\", float(train[DATE_COL].isna().mean()))","metadata":{"_uuid":"1864c7d4-f6ae-4dc2-9ecf-abc778b5bc57","_cell_guid":"402773cf-fd9f-49e3-bb33-d48602be02ab","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T15:52:14.9985Z","iopub.execute_input":"2026-06-05T15:52:14.999011Z","iopub.status.idle":"2026-06-05T16:01:42.425709Z","shell.execute_reply.started":"2026-06-05T15:52:14.998983Z","shell.execute_reply":"2026-06-05T16:01:42.425009Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"_uuid":"4f19b050-b065-4fee-97f6-7f0ef23265f8","_cell_guid":"11e15ec9-b99a-48ce-b0b2-cba47a6b7cd3","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:04:24.439194Z","iopub.execute_input":"2026-06-05T16:04:24.440026Z","iopub.status.idle":"2026-06-05T16:04:24.554435Z","shell.execute_reply.started":"2026-06-05T16:04:24.439997Z","shell.execute_reply":"2026-06-05T16:04:24.553704Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"_uuid":"c190e2b4-8dda-4959-ab7f-2105d689b23f","_cell_guid":"b0b0e70b-9125-4e1b-9eca-709d7d3f42cf","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:04:28.596686Z","iopub.execute_input":"2026-06-05T16:04:28.597565Z","iopub.status.idle":"2026-06-05T16:04:28.667743Z","shell.execute_reply.started":"2026-06-05T16:04:28.597534Z","shell.execute_reply":"2026-06-05T16:04:28.667141Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train Dimensions","metadata":{"_uuid":"5558fd1c-bdcf-4197-a829-678566102cbb","_cell_guid":"d17c6bd6-641c-4d2b-a3d8-3db17a3d737f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"print(\"Train Dimensions:\" , train.shape)","metadata":{"_uuid":"3657bf49-d7bd-4cb4-b963-bb77946b50cc","_cell_guid":"089381fc-b82b-47a4-a4e7-d9a913c26a3b","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:04:38.987207Z","iopub.execute_input":"2026-06-05T16:04:38.988112Z","iopub.status.idle":"2026-06-05T16:04:38.991998Z","shell.execute_reply.started":"2026-06-05T16:04:38.988072Z","shell.execute_reply":"2026-06-05T16:04:38.991173Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test Dimensions","metadata":{"_uuid":"b56dde51-d1cc-44ca-86d9-d67f12dfcf1e","_cell_guid":"b47d599d-5da9-4203-9e1b-b83ae53fa429","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"print(\"Test Dimensions:\" , test.shape)","metadata":{"_uuid":"bdf67ec9-fc53-4ad8-8152-0d664674516a","_cell_guid":"9bdb2318-c47c-4394-a4aa-e72014e79758","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:04:42.485679Z","iopub.execute_input":"2026-06-05T16:04:42.486495Z","iopub.status.idle":"2026-06-05T16:04:42.49044Z","shell.execute_reply.started":"2026-06-05T16:04:42.486464Z","shell.execute_reply":"2026-06-05T16:04:42.489729Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Dataset Information Summary","metadata":{"_uuid":"b38a9870-aa2f-4dc7-8749-566fb1c9498e","_cell_guid":"18742a5c-3248-474b-a3b8-15e96495b1d2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"train.info(max_cols=200, show_counts=True)","metadata":{"_uuid":"7c32c957-9306-4a03-8260-44db4e025444","_cell_guid":"b118b3fd-ccb5-4128-bb06-8abdf661c671","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:04:45.695264Z","iopub.execute_input":"2026-06-05T16:04:45.695539Z","iopub.status.idle":"2026-06-05T16:04:47.308586Z","shell.execute_reply.started":"2026-06-05T16:04:45.695516Z","shell.execute_reply":"2026-06-05T16:04:47.30788Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.info(max_cols=200, show_counts=True)","metadata":{"_uuid":"3c712e55-472f-4e98-be62-96b6990d6e23","_cell_guid":"33f9b5b9-8f3f-45bb-8778-a734fdb8f97c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:05:32.732969Z","iopub.execute_input":"2026-06-05T16:05:32.733627Z","iopub.status.idle":"2026-06-05T16:05:35.906617Z","shell.execute_reply.started":"2026-06-05T16:05:32.733599Z","shell.execute_reply":"2026-06-05T16:05:35.905867Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Display Dataset Columns","metadata":{"_uuid":"c64c0be6-181f-49d6-b1d9-5ac844a0c0a6","_cell_guid":"1251315c-b45c-4406-8f05-f64a483eaef6","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"\ndisplay(train.columns.tolist())","metadata":{"_uuid":"05ed1340-70ad-4c21-b644-b5f9b22738b6","_cell_guid":"e83fb322-d00f-40ff-893e-29058423ecc9","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:05:44.173746Z","iopub.execute_input":"2026-06-05T16:05:44.174195Z","iopub.status.idle":"2026-06-05T16:05:44.181665Z","shell.execute_reply.started":"2026-06-05T16:05:44.174164Z","shell.execute_reply":"2026-06-05T16:05:44.180933Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ndisplay(test.columns.tolist())","metadata":{"_uuid":"6ccb11d5-2abc-40f3-9785-69af6de9f16d","_cell_guid":"f3ea05a2-a2ce-4aad-8edc-53264ce75ddc","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:05:49.903863Z","iopub.execute_input":"2026-06-05T16:05:49.904665Z","iopub.status.idle":"2026-06-05T16:05:49.910547Z","shell.execute_reply.started":"2026-06-05T16:05:49.904635Z","shell.execute_reply":"2026-06-05T16:05:49.909776Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#  2. Exploratory Data Analysis (EDA)\n\n### Distribution of Target Variable\n\nThis code visualizes the overall distribution of the `target` variable using an interactive donut chart in Plotly. \n\n- **Default accounts** are represented in orange.\n- **Paid accounts** are represented in blue.\n- The chart shows the **percentage** of each category.","metadata":{"_uuid":"0b7357af-629f-43ee-ac0a-81cf05f20a0c","_cell_guid":"5473f1ba-ad8c-4e9d-ac84-5003d90477df","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"## Counting Statements per Customer\n\nWe can count how many rows (credit card statements) exist for each customer. About **84% and 88% of customers have 13 statements**, while the remaining 16% and 12% have between 1 and 12 statements for train and test, respectively.\n\n**Insight:** The model will need to handle **variable-length inputs per customer**, unless we simplify by using only the most recent statement or aggregating statistics across all statements.","metadata":{"_uuid":"c25bfaba-17b9-452d-8233-26acb1c56e08","_cell_guid":"53dc4dd4-561e-4c91-814f-b3ac0dd6243a","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 3) Quick EDA (lightweight)\n# ============================================================\ntarget_dist = (\n    train[TARGET].dropna().value_counts(normalize=True).rename(index={1:\"Default\", 0:\"Paid\"})\n)\ntotal_obs = int(train[TARGET].notna().sum())\n\nfig = go.Figure()\nfig.add_trace(go.Pie(\n    labels=target_dist.index,\n    values=target_dist.values * 100,\n    hole=0.45,\n    sort=False,\n    marker=dict(colors=[\"#8DBAE2\",\"#EDD3B3\"], line=dict(color=\"#016CC9\", width=2.0)),\n))\nfig.update_layout(\n    template=\"plotly_white\",\n    title=dict(text=\"Distribution of Target Variable (Default vs Paid)\", x=0.5),\n    annotations=[dict(text=f\"Total<br><b>{total_obs:,}</b>\", x=0.5, y=0.5, showarrow=False)]\n)\nfig.show()\n\ntrain_sc = train[ID_COL].value_counts().value_counts().sort_index(ascending=False)\ntest_sc  = test[ID_COL].value_counts().value_counts().sort_index(ascending=False)\n\nfig, axes = plt.subplots(1, 2, figsize=(12, 6))\naxes[0].pie(train_sc, labels=train_sc.index, startangle=30); axes[0].set_title(\"Train statements per customer\")\naxes[1].pie(test_sc,  labels=test_sc.index,  startangle=30); axes[1].set_title(\"Test statements per customer\")\nplt.tight_layout()\nplt.show()","metadata":{"_uuid":"6367897e-ce57-4b64-b171-3dccdc7f765f","_cell_guid":"44f63bc2-e9e9-467f-ad59-fe395c4fbe3c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:05:56.393094Z","iopub.execute_input":"2026-06-05T16:05:56.393767Z","iopub.status.idle":"2026-06-05T16:06:00.98357Z","shell.execute_reply.started":"2026-06-05T16:05:56.393735Z","shell.execute_reply":"2026-06-05T16:06:00.982668Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Categorical Features\n\nBased on the [data description](https://www.kaggle.com/competitions/amex-default-prediction/data), there are **eleven categorical features**. We visualize their distributions for `target = 0` and `target = 1`.\n\nFor the ten features with missing values, **missing entries are shown as the rightmost bar** in each histogram.","metadata":{"_uuid":"57a24aed-7c07-47d6-96c9-e2baba5498f1","_cell_guid":"881788b2-35dd-475e-b99b-159423d994b6","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# --------------------------------------------\n# Categorical features distribution by target\n# --------------------------------------------\n\nplt.figure(figsize=(16, 16))\n\nfor i, f in enumerate(CAT_FEATURES):\n    plt.subplot(4, 3, i + 1)\n    \n    # Value counts for target=0\n    temp0 = train[train.target == 0][f].value_counts(dropna=False, normalize=True).sort_index()\n    plt.bar(temp0.index - 0.15, temp0.values, width=0.3, alpha=0.5, label='target=0')\n    \n    # Value counts for target=1\n    temp1 = train[train.target == 1][f].value_counts(dropna=False, normalize=True).sort_index()\n    plt.bar(temp1.index + 0.15, temp1.values, width=0.3, alpha=0.5, label='target=1')\n    \n    plt.xlabel(f)\n    plt.ylabel('Frequency')\n    plt.legend()\n    plt.xticks(temp0.index, temp0.index.astype(str))  # ensure tick labels are string\n\nplt.suptitle('Categorical Features Distribution by Target', fontsize=20, y=0.93)\nplt.tight_layout(rect=[0, 0, 1, 0.95])\nplt.show()","metadata":{"_uuid":"a759e51c-d1eb-4141-9c45-41d656027c97","_cell_guid":"a835312c-9fb1-49f5-afa0-cea1a0c0637b","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:06:06.018889Z","iopub.execute_input":"2026-06-05T16:06:06.019486Z","iopub.status.idle":"2026-06-05T16:06:29.380643Z","shell.execute_reply.started":"2026-06-05T16:06:06.019455Z","shell.execute_reply":"2026-06-05T16:06:29.379933Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## The Binary Features\n\nTwo features are binary:\n\n- `Balance_31` contains only 0 or 1. \n- `Delinquency_87` contains only 1 or missing values.","metadata":{"_uuid":"dca552ac-3a98-481e-8fda-83fd9acf7bd8","_cell_guid":"e70055de-55b4-4a3b-a95b-4baf50e78158","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ---------------------------\n# Binary features\n# ---------------------------\nbin_features = ['Balance_31', 'Delinquency_87']\n\nplt.figure(figsize=(16, 4))\n\nfor i, f in enumerate(bin_features):\n    plt.subplot(1, 2, i + 1)\n    \n    # Value counts for target=0\n    temp0 = train[train.target == 0][f].value_counts(dropna=False, normalize=True).sort_index()\n    plt.bar(temp0.index - 0.1, temp0.values, width=0.2, alpha=0.5, label='target=0')\n    \n    # Value counts for target=1\n    temp1 = train[train.target == 1][f].value_counts(dropna=False, normalize=True).sort_index()\n    plt.bar(temp1.index + 0.1, temp1.values, width=0.2, alpha=0.5, label='target=1')\n    \n    plt.xlabel(f)\n    plt.ylabel('Frequency')\n    plt.xticks(temp0.index, temp0.index.astype(str))\n    plt.legend()\n\nplt.suptitle('Binary Features Distribution by Target', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.95])\nplt.show()","metadata":{"_uuid":"680b19d5-41bd-4ccd-8683-d67c3acd9f26","_cell_guid":"022c15bd-1467-4adc-8c72-f95c7aa8e9dc","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:06:54.057024Z","iopub.execute_input":"2026-06-05T16:06:54.058009Z","iopub.status.idle":"2026-06-05T16:06:58.196693Z","shell.execute_reply.started":"2026-06-05T16:06:54.057968Z","shell.execute_reply":"2026-06-05T16:06:58.196058Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## The Numerical Features\n\nIf I plot histograms for the 175 numerical features, I can see that they exhibit a wide variety of distributions.","metadata":{"_uuid":"a3dfae45-f4ec-44ac-bfbb-59660f44258b","_cell_guid":"7e719bae-648c-4181-adb9-b35dfbd09f60","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ----------------------------------------\n# Continuous (numerical) features\n# -----------------------------------------\ncont_features = sorted([\n    f for f in train.columns\n    if f not in CAT_FEATURES + bin_features + ['customer_ID', 'target', 'Statement_Date']\n])\nprint(f\"Number of continuous features: {len(cont_features)}\")\n\n# Plot continuous features in batches of 4 per row\nncols = 4\nfor i, f in enumerate(cont_features):\n    # Start a new figure every ncols features\n    if i % ncols == 0:\n        if i > 0:\n            plt.show()\n        plt.figure(figsize=(16, 3))\n        if i == 0:\n            plt.suptitle('Continuous Features', fontsize=20, y=1.02)\n    \n    plt.subplot(1, ncols, i % ncols + 1)\n    plt.hist(train[f], bins=200, color='skyblue', edgecolor='black')\n    plt.xlabel(f)\n    plt.ylabel('Frequency')\n\nplt.tight_layout()\nplt.show()","metadata":{"_uuid":"57f50d81-5249-4343-b061-5f26a892ce49","_cell_guid":"2fec089d-46b6-4a98-b125-573c3331511a","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:07:02.356364Z","iopub.execute_input":"2026-06-05T16:07:02.356766Z","iopub.status.idle":"2026-06-05T16:07:57.102413Z","shell.execute_reply.started":"2026-06-05T16:07:02.356739Z","shell.execute_reply":"2026-06-05T16:07:57.101774Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Feature Grouping","metadata":{"_uuid":"8df4fe82-88ab-44b0-87ed-09acc9e697b5","_cell_guid":"407a93ab-085b-436f-92f7-807dbfcb56df","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"DELINQ_COLS = [c for c in train.columns if c.startswith(\"Delinquency_\")]\nBALANCE_COLS = [c for c in train.columns if c.startswith(\"Balance_\")]\nSPEND_COLS = [c for c in train.columns if c.startswith(\"Spend_\")]\nPAYMENT_COLS = [c for c in train.columns if c.startswith(\"Payment_\")]\nRISK_COLS = [c for c in train.columns if c.startswith(\"Risk_\")]\n\nNUMERIC_COLS = (\n    DELINQ_COLS + BALANCE_COLS + SPEND_COLS + PAYMENT_COLS + RISK_COLS\n)","metadata":{"_uuid":"fd29891e-b9ea-4ff4-80b0-9b21575804dd","_cell_guid":"ee9af5bb-8efa-4350-ad5a-c650b520e0ea","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:09:36.830972Z","iopub.execute_input":"2026-06-05T16:09:36.83159Z","iopub.status.idle":"2026-06-05T16:09:36.836933Z","shell.execute_reply.started":"2026-06-05T16:09:36.831562Z","shell.execute_reply":"2026-06-05T16:09:36.836193Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Automatic Feature Counting\n# -----------------------------\n\nfeature_counts = {\n    \"customer_ID\": 1 if ID_COL in train.columns else 0,\n    \"Statement_Date\": 1 if DATE_COL in train.columns else 0,\n    \"Delinquency_*\": len(DELINQ_COLS),\n    \"Spend_*\": len(SPEND_COLS),\n    \"Payment_*\": len(PAYMENT_COLS),\n    \"Balance_*\": len(BALANCE_COLS),\n    \"Risk_*\": len(RISK_COLS),\n}\n\nfeature_counts[\"Total\"] = sum(feature_counts.values())","metadata":{"_uuid":"545dfbab-34f6-4254-b51c-c6ca68ceb85a","_cell_guid":"0d51887c-2974-42df-bf13-541dadab5996","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:09:52.918029Z","iopub.execute_input":"2026-06-05T16:09:52.918425Z","iopub.status.idle":"2026-06-05T16:09:52.923055Z","shell.execute_reply.started":"2026-06-05T16:09:52.918397Z","shell.execute_reply":"2026-06-05T16:09:52.922275Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Summary Table of Features\n\nThis interactive table gives a **clear overview of dataset statistics**, helping to quickly spot patterns, anomalies, or missing data before modeling.","metadata":{"_uuid":"3e90094b-ddbd-4115-b376-c8364a915738","_cell_guid":"cfc66761-905c-4de1-b345-47160e6562fa","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ---------------------------\n# Build summary table\n# ---------------------------\ntable_I = pd.DataFrame(\n    [feature_counts],\n    index=[\"Train\"]\n)\n\n# ---------------------------\n# Create Plotly table\n# ---------------------------\nfig = go.Figure(\n    data=[\n        go.Table(\n            header=dict(\n                values=[\"Attributes\"] + list(table_I.columns),\n                fill_color=\"#F2F2F2\",\n                align=\"center\",\n                font=dict(size=14, color=\"black\"),\n                line_color=\"black\",\n            ),\n            cells=dict(\n                values=[\n                    table_I.index.tolist(),\n                    *[table_I[col].tolist() for col in table_I.columns]\n                ],\n                fill_color=\"white\",\n                align=\"center\",\n                font=dict(size=13, color=\"black\"),\n                line_color=\"black\",\n            ),\n        )\n    ]\n)\n\n# ---------------------------\n# Layout\n# ---------------------------\nfig.update_layout(\n    title=dict(\n        text=\"Interpretation of Anonymized Time-Series Behavioral Data\",\n        x=0.5,\n        font=dict(size=18, color=\"black\"),\n    ),\n    width=1000,\n    height=300,\n    margin=dict(l=20, r=20, t=80, b=20),\n)\n\nfig.show()","metadata":{"_uuid":"a365113f-9099-487d-925c-8305bdc89555","_cell_guid":"c2bea3f0-6457-49de-95fd-9951f8123bb0","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:09:56.994333Z","iopub.execute_input":"2026-06-05T16:09:56.994967Z","iopub.status.idle":"2026-06-05T16:09:57.050041Z","shell.execute_reply.started":"2026-06-05T16:09:56.994938Z","shell.execute_reply":"2026-06-05T16:09:57.049268Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =====================================================================\n# Helper function: compute counts + percentages for \"13 records\" vs \"1-12\"\n# =====================================================================\ndef behavior_record_stats(df, id_col=\"customer_ID\"):\n    customer_records = df.groupby(id_col).size()\n\n    records_13 = int((customer_records == 13).sum())\n    records_1_12 = int((customer_records < 13).sum())\n    total_customers = int(customer_records.shape[0])\n\n    pct_13 = (records_13 / total_customers) * 100 if total_customers else 0\n    pct_1_12 = (records_1_12 / total_customers) * 100 if total_customers else 0\n\n    return {\n        \"customer_records\": customer_records,\n        \"records_13\": records_13,\n        \"records_1_12\": records_1_12,\n        \"total_customers\": total_customers,\n        \"pct_13\": pct_13,\n        \"pct_1_12\": pct_1_12\n    }\n\n# Compute stats for Train and Test\ntrain_stats = behavior_record_stats(train, id_col=\"customer_ID\")\ntest_stats  = behavior_record_stats(test,  id_col=\"customer_ID\")\n\n# =====================================================================\n# Plot settings\n# =====================================================================\ncolors = [\"#4472C4\", \"#ED7D31\"]\nexplode = (0.05, 0)\ncategories = [\"13 records\", \"1-12 records\"]\n\n# Create 2x2 layout: (Train pie, Train bar) on top; (Test pie, Test bar) below\nfig, axes = plt.subplots(2, 2, figsize=(16, 10))\nfig.patch.set_facecolor(\"white\")\n\ndef plot_panel(ax_pie, ax_bar, stats, title_prefix):\n    # ----- Pie -----\n    percentages = [stats[\"pct_13\"], stats[\"pct_1_12\"]]\n\n    wedges, texts, autotexts = ax_pie.pie(\n        percentages,\n        labels=categories,\n        autopct=\"%1.1f%%\",\n        startangle=90,\n        colors=colors,\n        explode=explode,\n        textprops={\"fontsize\": 12, \"fontweight\": \"bold\"}\n    )\n\n    for autotext in autotexts:\n        autotext.set_color(\"white\")\n        autotext.set_fontsize(14)\n\n    ax_pie.set_title(\n        f\"{title_prefix}: Distribution of Behavior Records\\nPer Customer\",\n        fontsize=13, fontweight=\"bold\", pad=20\n    )\n    ax_pie.set_facecolor(\"white\")\n\n    # ----- Bar -----\n    counts = [stats[\"records_13\"], stats[\"records_1_12\"]]\n    x = np.arange(len(categories))\n    width = 0.55\n\n    bars = ax_bar.bar(\n        x, counts, width,\n        color=colors, alpha=0.9,\n        edgecolor=\"black\", linewidth=0.5\n    )\n\n    ax_bar.set_facecolor(\"white\")\n    ax_bar.set_xlabel(\"Behavior Records Category\", fontsize=11, fontweight=\"bold\")\n    ax_bar.set_ylabel(\"Number of Customers\", fontsize=11, fontweight=\"bold\")\n    ax_bar.set_title(\n        f\"{title_prefix}: Customer Counts by Record Type\",\n        fontsize=13, fontweight=\"bold\", pad=20\n    )\n    ax_bar.set_xticks(x)\n    ax_bar.set_xticklabels(categories)\n    ax_bar.grid(True, axis=\"y\", alpha=0.3, linewidth=0.5, color=\"gray\", linestyle=\"-\")\n    ax_bar.tick_params(labelsize=10)\n\n    # Spine styling\n    for spine in [\"top\", \"right\", \"bottom\", \"left\"]:\n        ax_bar.spines[spine].set_linewidth(0.5)\n\n    # Value labels\n    for i, bar in enumerate(bars):\n        height = bar.get_height()\n        ax_bar.text(\n            bar.get_x() + bar.get_width()/2.,\n            height,\n            f\"{int(height):,}\\n({percentages[i]:.1f}%)\",\n            ha=\"center\", va=\"bottom\",\n            fontsize=10, fontweight=\"bold\"\n        )\n\n# Plot Train (row 0)\nplot_panel(axes[0, 0], axes[0, 1], train_stats, \"Train\")\n\n# Plot Test (row 1)\nplot_panel(axes[1, 0], axes[1, 1], test_stats, \"Test\")\n\nplt.tight_layout(rect=[0, 0.04, 1, 1])\nplt.savefig(\"train_test_behavior_records.png\", dpi=300, bbox_inches=\"tight\", facecolor=\"white\")\nplt.show()\n\n# =====================================================================\n# Print summary statistics\n# =====================================================================\ndef print_stats(stats, label):\n    print(\"\\n\" + \"=\"*60)\n    print(f\"{label} Behavior Records Summary\")\n    print(\"=\"*60)\n    print(f\"Total Customers: {stats['total_customers']:,}\")\n    print(f\"  - 13 records: {stats['records_13']:,} ({stats['pct_13']:.1f}%)\")\n    print(f\"  - 1-12 records: {stats['records_1_12']:,} ({stats['pct_1_12']:.1f}%)\")\n    print(f\"\\nAverage records per customer: {stats['customer_records'].mean():.2f}\")\n    print(f\"Median records per customer: {stats['customer_records'].median():.0f}\")\n    print(\"=\"*60)\n\nprint_stats(train_stats, \"Train\")\nprint_stats(test_stats, \"Test\")","metadata":{"_uuid":"a0ef035d-fd25-4d76-9714-306513fceb88","_cell_guid":"dd49549c-0545-44b0-8f80-50990412a39d","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:10:01.343383Z","iopub.execute_input":"2026-06-05T16:10:01.34413Z","iopub.status.idle":"2026-06-05T16:10:05.727529Z","shell.execute_reply.started":"2026-06-05T16:10:01.3441Z","shell.execute_reply":"2026-06-05T16:10:05.726916Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Features with Highest Percentage of Missing Values\n\nThis analysis computes the percentage of missing values for each feature in the `train` dataset. \nThe proportion of missing observations is calculated by dividing the number of missing values in each column by the total number of records and converting the result to a percentage.\n\nThe features are then sorted in descending order based on their missingness, and the top 50 variables with the highest proportion of missing values are visualized. \nThis step helps identify features that may require imputation, transformation, or exclusion prior to further exploratory data analysis or model development.","metadata":{"_uuid":"3f6a920e-fef6-4127-a301-5bc85083a0cd","_cell_guid":"9f88bafd-e96d-4547-b916-b7203bb16e7b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ---------------------------\n# Missing data distribution (Top 50 features)\n# ---------------------------\nmissing_data = (\n    train.isnull()\n    .sum()\n    .div(len(train))\n    .mul(100)\n    .sort_values(ascending=False)\n)\n\ntop_50_missing = missing_data.head(50)\n\nfig_df = pd.DataFrame({\n    \"Feature\": top_50_missing.index,\n    \"Missing_Pct\": top_50_missing.values\n})\n\n# ---------------------------\n# Colors (severity-based, AmEx-style)\n# ---------------------------\ndef severity_color(pct):\n    if pct > 60:\n        return \"#C73E3A\"   # High missing\n    elif pct > 30:\n        return \"#F0A44B\"   # Medium missing\n    else:\n        return \"#4CAF50\"   # Low missing\n\nbar_colors = fig_df[\"Missing_Pct\"].apply(severity_color)\n\n# ---------------------------\n# Create figure\n# ---------------------------\nfig = go.Figure()\n\nfig.add_trace(\n    go.Bar(\n        x=fig_df[\"Feature\"],\n        y=fig_df[\"Missing_Pct\"],\n        marker=dict(\n            color=bar_colors,\n            line=dict(width=0.6, color=\"black\")\n        ),\n        hovertemplate=(\n            \"<b>%{x}</b><br>\"\n            \"Missing: %{y:.2f}%<br>\"\n            \"<extra></extra>\"\n        ),\n    )\n)\n\n# ---------------------------\n# Layout\n# ---------------------------\nfig.update_layout(\n    template=\"plotly_white\",\n    title=dict(\n        text=\"Amount of Missing Data (Top 50 Features)\",\n        x=0.5,\n        font=dict(size=20)\n    ),\n    xaxis=dict(\n        title=\"Features\",\n        tickangle=90,\n        showgrid=False\n    ),\n    yaxis=dict(\n        title=\"Missing Data (%)\",\n        gridcolor=\"rgba(0,0,0,0.1)\"\n    ),\n    legend=dict(\n        title=\"Missingness Severity\",\n        orientation=\"h\",\n        y=1.12,\n        x=0.5,\n        xanchor=\"center\"\n    ),\n    width=900,\n    height=520,\n    margin=dict(l=60, r=40, t=90, b=120)\n)\n\nfig.show()","metadata":{"_uuid":"2afe74f2-452f-4299-a812-498f35dff77a","_cell_guid":"0bc1062f-b474-4f4a-bd09-2f70ec92fd69","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:10:22.06696Z","iopub.execute_input":"2026-06-05T16:10:22.067277Z","iopub.status.idle":"2026-06-05T16:10:23.374257Z","shell.execute_reply.started":"2026-06-05T16:10:22.067252Z","shell.execute_reply":"2026-06-05T16:10:23.37357Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# credit: https://www.kaggle.com/willkoehrsen/start-here-a-gentle-introduction. \n# One of the best notebooks on getting started with a ML problem.\n\ndef missing_values_table(df):\n        # Total missing values\n        mis_val = df.isnull().sum()\n        \n        # Percentage of missing values\n        mis_val_percent = 100 * df.isnull().sum() / len(df)\n        \n        # Make a table with the results\n        mis_val_table = pd.concat([mis_val, mis_val_percent], axis=1)\n        \n        # Rename the columns\n        mis_val_table_ren_columns = mis_val_table.rename(\n        columns = {0 : 'Missing Values', 1 : '% of Total Values'})\n        \n        # Sort the table by percentage of missing descending\n        mis_val_table_ren_columns = mis_val_table_ren_columns[\n            mis_val_table_ren_columns.iloc[:,1] != 0].sort_values(\n        '% of Total Values', ascending=False).round(1)\n        \n        # Print some summary information\n        print (\"Selected dataframe has \" + str(df.shape[1]) + \" columns.\\n\"      \n            \"There are \" + str(mis_val_table_ren_columns.shape[0]) +\n              \" columns that have missing values.\")\n        \n        # Return the dataframe with missing information\n        return mis_val_table_ren_columns","metadata":{"_uuid":"1e266f99-6525-451d-bd9b-a8f062f641a4","_cell_guid":"1fb970c9-eb18-48b6-9213-68d179abd4b0","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:10:56.548143Z","iopub.execute_input":"2026-06-05T16:10:56.549313Z","iopub.status.idle":"2026-06-05T16:10:56.554683Z","shell.execute_reply.started":"2026-06-05T16:10:56.549274Z","shell.execute_reply":"2026-06-05T16:10:56.553857Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_Test = missing_values_table(train)\ndf_Test","metadata":{"_uuid":"b0216ea5-2946-4d69-b59f-cce5c19b62aa","_cell_guid":"88a60edd-dcb4-4220-a188-0afc16e4333f","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:10:59.067929Z","iopub.execute_input":"2026-06-05T16:10:59.068838Z","iopub.status.idle":"2026-06-05T16:11:01.549343Z","shell.execute_reply.started":"2026-06-05T16:10:59.068805Z","shell.execute_reply":"2026-06-05T16:11:01.548488Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Handling Missing Values","metadata":{"_uuid":"48d13636-e6bf-4313-976c-8cc6e41b4298","_cell_guid":"a7f0214b-c059-4adc-a3fc-3504b2112f9c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 4) Imputation\n# ============================================================\nNUMERIC_COLS = [c for c in train.columns if c not in (CAT_FEATURES + [DATE_COL, ID_COL, TARGET])]\n\ncat_fill = {}\nfor col in CAT_FEATURES:\n    if col in train.columns:\n        m = train[col].mode(dropna=True)\n        cat_fill[col] = m.iloc[0] if len(m) else \"Unknown\"\n\nnum_fill = {}\nfor col in NUMERIC_COLS:\n    if col in train.columns:\n        num_fill[col] = train[col].median()\n\nfor col, v in cat_fill.items():\n    if col in train.columns: train[col].fillna(v, inplace=True)\n    if col in test.columns:  test[col].fillna(v, inplace=True)\n\nfor col, v in num_fill.items():\n    if col in train.columns: train[col].fillna(v, inplace=True)\n    if col in test.columns:  test[col].fillna(v, inplace=True)\n\nmissing_train = int(train[CAT_FEATURES + NUMERIC_COLS].isna().sum().sum())\nmissing_test  = int(test[CAT_FEATURES + NUMERIC_COLS].isna().sum().sum())\nprint(f\"Missing after imputation -> Train: {missing_train:,} | Test: {missing_test:,}\")","metadata":{"_uuid":"31335a18-7e10-4cd4-91a3-3b06341cd34f","_cell_guid":"901c95a2-fe52-4008-96bc-2c7616c84e6e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:11:05.222943Z","iopub.execute_input":"2026-06-05T16:11:05.223725Z","iopub.status.idle":"2026-06-05T16:11:24.076265Z","shell.execute_reply.started":"2026-06-05T16:11:05.223696Z","shell.execute_reply":"2026-06-05T16:11:24.075367Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"_uuid":"1df08d51-4ce2-4e0d-b08c-885935034256","_cell_guid":"0d09fa66-89cb-4706-8341-a4b7fc056d6b","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:12:00.815529Z","iopub.execute_input":"2026-06-05T16:12:00.81616Z","iopub.status.idle":"2026-06-05T16:12:00.889025Z","shell.execute_reply.started":"2026-06-05T16:12:00.816132Z","shell.execute_reply":"2026-06-05T16:12:00.888415Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"_uuid":"d52a2ceb-70e7-4125-8ef8-fbe09e75a993","_cell_guid":"2411ee3e-cdd8-47ae-8f4a-2c8daa2749af","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:12:03.392295Z","iopub.execute_input":"2026-06-05T16:12:03.393037Z","iopub.status.idle":"2026-06-05T16:12:03.47136Z","shell.execute_reply.started":"2026-06-05T16:12:03.393005Z","shell.execute_reply":"2026-06-05T16:12:03.470697Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#  3. Feature Engineering\n\nThis is a **memory-safe** approach designed for AMEX scale.\n\n###  Why light aggregations?\n- AMEX train has 5.5M rows; test has 11.3M → using diff, rank, last‑k per column blows memory.\n- We only compute:\n  - **mean**\n  - **std**\n  - **last**\n- For categorical variables:\n  - **last**\n  - **nunique**\n\n###  Output:\n- One row per `customer_ID`.\n- Compact but strong baseline representation.","metadata":{"_uuid":"2733c5f1-c73d-481b-89e3-b32bde925f9b","_cell_guid":"38800df3-c5a1-48b4-8053-9bdf36cdcc39","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 5) Customer-level Aggregates (baseline features)\n# ============================================================\ndef safe_agg(df, group_col, cols, agg_funcs, prefix=\"\"):\n    if not cols:\n        return pd.DataFrame({group_col: df[group_col].unique()})\n    agg = df.groupby(group_col, sort=False, observed=True)[cols].agg(agg_funcs)\n    agg.columns = [f\"{prefix}{'_'.join(map(str, c)).strip('_')}\" for c in agg.columns]\n    return agg.reset_index()\n\ndef build_num_aggregates(df, num_cols):\n    if not num_cols:\n        return pd.DataFrame({ID_COL: df[ID_COL].unique()})\n    # IMPROVE: add \"min\",\"max\",\"sum\" if memory computer allows\n    return safe_agg(df, ID_COL, num_cols, [\"mean\",\"std\",\"last\"])\n\ndef build_cat_aggregates(df, cat_cols):\n    valid = [c for c in cat_cols if c in df.columns and c != ID_COL]\n    if not valid:\n        return pd.DataFrame({ID_COL: df[ID_COL].unique()})\n    return safe_agg(df, ID_COL, valid, [\"last\",\"nunique\"])\n\ndef build_all_features_basic(df, num_cols, cat_cols):\n    nums = build_num_aggregates(df, num_cols)\n    cats = build_cat_aggregates(df, cat_cols)\n    out = nums.merge(cats, on=ID_COL, how=\"left\", copy=False)\n    for c in out.select_dtypes(\"float64\").columns:\n        out[c] = out[c].astype(np.float32)\n    return out\n\nexclude = {ID_COL, DATE_COL, TARGET} | set(CAT_FEATURES)\nnum_cols = [c for c in train.columns if c not in exclude and pd.api.types.is_numeric_dtype(train[c])]\n\ntrain_features = build_all_features_basic(train, num_cols, CAT_FEATURES)\ntest_features  = build_all_features_basic(test,  num_cols, CAT_FEATURES)\n\nprint(\"train_features:\", train_features.shape, \"| test_features:\", test_features.shape)","metadata":{"_uuid":"5466cd07-43f4-4be8-8cd0-827d799e184f","_cell_guid":"c0664424-ee59-49ed-b150-5e75f813dd39","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:15:22.373487Z","iopub.execute_input":"2026-06-05T16:15:22.373942Z","iopub.status.idle":"2026-06-05T16:16:59.633588Z","shell.execute_reply.started":"2026-06-05T16:15:22.373913Z","shell.execute_reply":"2026-06-05T16:16:59.632768Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_features.head()","metadata":{"_uuid":"d58d62ec-e697-4dbd-ae50-6e67da57f434","_cell_guid":"223bc7d3-870a-4bb9-b574-06d377876e3b","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:16:59.63492Z","iopub.execute_input":"2026-06-05T16:16:59.635307Z","iopub.status.idle":"2026-06-05T16:16:59.888782Z","shell.execute_reply.started":"2026-06-05T16:16:59.63528Z","shell.execute_reply":"2026-06-05T16:16:59.888058Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_features.head()","metadata":{"_uuid":"91c60a9f-58a5-433c-b1db-8385f29a7f9f","_cell_guid":"6cfa94c7-4d1f-4b5e-adee-cf51f1012d0f","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:17:10.259983Z","iopub.execute_input":"2026-06-05T16:17:10.260794Z","iopub.status.idle":"2026-06-05T16:17:10.496192Z","shell.execute_reply.started":"2026-06-05T16:17:10.260762Z","shell.execute_reply":"2026-06-05T16:17:10.495488Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4. Align Labels & Stratified Train / Validation Split\n\nNow that we have:\n\n- `train_features` – one row per `customer_ID` with aggregated features.\n- `test_features` – same feature construction on the test set.\n\nWe need to:\n\n1. **Align the target labels** with `train_features`, since:\n   - `train` is at **statement** level.\n   - `train_features` is at **customer** level.\n2. Define a clean list of feature columns (`FEATURE_COLS`).\n3. Ensure `test_features` has exactly the same columns in the same order.\n4. Perform a **stratified train/validation split**:\n   - Stratified on the target (`y`) to preserve default vs paid proportions.\n   - Index-based to avoid building a huge `X` before splitting.","metadata":{"_uuid":"819cd2fd-72e2-446c-9290-7f1d2a7731c8","_cell_guid":"f007307e-47e3-4b60-8d19-528947c06e2a","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 6) Memory-SAFE ALIGN + SPLIT\n# ============================================================\nfrom pathlib import Path\n\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\nOUT_DIR.mkdir(exist_ok=True)\n\n# ------------------------------------------------------------\n# Step 1: Align labels\n# ------------------------------------------------------------\ntrain_id_target = train[[ID_COL, TARGET]].drop_duplicates(subset=[ID_COL]).set_index(ID_COL)\n\ny = train_id_target.loc[train_features[ID_COL].values, TARGET].values.astype(np.int64)\n\ndel train_id_target\ngc.collect()\n\n# ------------------------------------------------------------\n# Step 2: Feature columns\n# ------------------------------------------------------------\nFEATURE_COLS = [c for c in train_features.columns if c != ID_COL]\n\n# safer fill\ntrain_features.loc[:, FEATURE_COLS] = train_features[FEATURE_COLS].fillna(0.0)\n\ntest_features = test_features.set_index(ID_COL).reindex(columns=FEATURE_COLS)\ntest_features.loc[:, FEATURE_COLS] = test_features[FEATURE_COLS].fillna(0.0)\ntest_features = test_features.reset_index()\n\n# ------------------------------------------------------------\n# Step 3: Split ONLY indices\n# ------------------------------------------------------------\nidx = np.arange(len(train_features))\n\nidx_train, idx_valid, y_train, y_valid = train_test_split(\n    idx, y, test_size=0.2, random_state=SEED, stratify=y\n)\n\nprint(\"Index split complete:\", len(idx_train), len(idx_valid))\n\n# ------------------------------------------------------------\n# Step 4: Convert to NUMPY ONCE\n# ------------------------------------------------------------\ntrain_np = train_features[FEATURE_COLS].to_numpy(dtype=np.float32, copy=False)\ntest_np  = test_features[FEATURE_COLS].to_numpy(dtype=np.float32, copy=False)\n\n# ------------------------------------------------------------\n# Step 5: SAVE to MEMMAP\n# ------------------------------------------------------------\ntrain_path = OUT_DIR / \"tabular_train.npy\"\ntest_path  = OUT_DIR / \"tabular_test.npy\"\n\nnp.save(train_path, train_np)\nnp.save(test_path, test_np)\n\n# FREE RAM immediately\ndel train_np, test_np, train_features, test_features\ngc.collect()\n\n# ------------------------------------------------------------\n# Step 6: LOAD using MEMMAP\n# ------------------------------------------------------------\nX_all = np.load(train_path, mmap_mode=\"r\")\nX_test = np.load(test_path, mmap_mode=\"r\")\n\n# create views\nX_train = X_all[idx_train]\nX_valid = X_all[idx_valid]\n\nprint(\"X_train:\", X_train.shape, \"X_valid:\", X_valid.shape, \"X_test:\", X_test.shape)\n\n# ------------------------------------------------------------\n# Step 7: Class balance\n# ------------------------------------------------------------\npos = int((y_train == 1).sum())\nneg = int((y_train == 0).sum())\nscale_pos = neg / max(1, pos)\n\nprint(f\"Class balance -> pos={pos:,} neg={neg:,} scale_pos_weight={scale_pos:.2f}\")","metadata":{"_uuid":"bc4f802c-aa40-4513-a3a9-67f632609a8f","_cell_guid":"647bbfcd-4bf7-4bbb-9fdb-bd9059e04297","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:17:23.904988Z","iopub.execute_input":"2026-06-05T16:17:23.90566Z","iopub.status.idle":"2026-06-05T16:18:12.141647Z","shell.execute_reply.started":"2026-06-05T16:17:23.905627Z","shell.execute_reply":"2026-06-05T16:18:12.140711Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5. AMEX Metric\n\nWe use the official AMEX competition metric, which combines:\n\n1. **Top 4% capture rate**:\n- Sort customers by predicted default probability (descending).\n- Use weights: 20 for non-defaults (`y=0`), 1 for defaults (`y=1`).\n- Select the subset of customers up to **4% of total weighted population**.\n- Compute the fraction of all defaults captured in this subset.\n\n2. **Weighted Gini coefficient**:\n- Measures rank-quality of predictions vs actuals.\n- Uses the same weighting scheme.\n- Normalized by the maximum possible Gini (`Gini_max`) computed on perfect predictions.\n\nFinal AMEX score:\n\n\n$$\n\\text{AMEX} = 0.5 \\times \\left( \\frac{\\text{Gini}}{\\text{Gini}_{\\text{max}}} + \\text{Top-4\\% capture} \\right)\n$$\n\n\nWe implement this as `amex_metric(y_true, y_pred)` and use it on validation predictions.","metadata":{"_uuid":"bd3d79a0-2bac-4544-a9eb-882bd97e2518","_cell_guid":"85b5b6c8-c1c7-4847-9a2f-e50c6747955e","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# 5) Metrics Suite (AMEX + confusion + ROC/PR + calibration)\n# ============================================================\nfrom sklearn.metrics import (\n    confusion_matrix, classification_report,\n    precision_recall_curve, average_precision_score,\n    f1_score, precision_score, recall_score, accuracy_score,\n    roc_auc_score, roc_curve, brier_score_loss\n)\n\ndef amex_metric(y_true, y_pred):\n    y_true = np.asarray(y_true, dtype=np.int32)\n    y_pred = np.asarray(y_pred, dtype=np.float64)\n\n    order = np.argsort(-y_pred)\n    y_true_sorted = y_true[order]\n    weights_sorted = np.where(y_true_sorted == 0, 20, 1)\n\n    cum_weights = np.cumsum(weights_sorted)\n    cutoff = 0.04 * cum_weights[-1]\n    top_mask = cum_weights <= cutoff\n    top_four = y_true_sorted[top_mask].sum() / max(1, y_true_sorted.sum())\n\n    def weighted_gini(actual, pred, weight):\n        order_pred = np.argsort(pred)\n        actual = actual[order_pred]\n        weight = weight[order_pred]\n        cum_weight = np.cumsum(weight)\n        cum_actual = np.cumsum(actual * weight)\n        if cum_actual[-1] == 0:\n            return 0.0\n        lorentz = cum_actual / cum_actual[-1]\n        random = cum_weight / cum_weight[-1]\n        return (lorentz - random).sum()\n\n    weights_full = np.where(y_true == 0, 20, 1)\n    g = weighted_gini(y_true, y_pred, weights_full)\n    g_max = weighted_gini(y_true, y_true, weights_full)\n\n    return 0.5 * (g / (g_max + 1e-9) + top_four)\n\ndef eval_classification(y_true, y_prob, threshold=0.5, name=\"Model\"):\n    y_true = np.asarray(y_true).astype(int)\n    y_prob = np.asarray(y_prob)\n    y_pred = (y_prob >= threshold).astype(int)\n\n    auc = roc_auc_score(y_true, y_prob)\n    ap  = average_precision_score(y_true, y_prob)  # PR-AUC\n    brier = brier_score_loss(y_true, y_prob)\n    amex = amex_metric(y_true, y_prob)\n\n    acc = accuracy_score(y_true, y_pred)\n    prec = precision_score(y_true, y_pred, zero_division=0)\n    rec  = recall_score(y_true, y_pred, zero_division=0)\n    f1   = f1_score(y_true, y_pred, zero_division=0)\n\n    cm = confusion_matrix(y_true, y_pred)\n\n    print(f\"\\n=== {name} @ threshold={threshold:.2f} ===\")\n    print(f\"AMEX   : {amex:.5f}\")\n    print(f\"AUC-ROC: {auc:.5f}\")\n    print(f\"PR-AUC : {ap:.5f}\")\n    print(f\"Brier  : {brier:.5f}\")\n    print(f\"Accuracy: {acc:.5f} | Precision: {prec:.5f} | Recall: {rec:.5f} | F1: {f1:.5f}\")\n    print(\"Confusion matrix [ [TN FP], [FN TP] ]:\")\n    print(cm)\n    print(\"\\nClassification report:\")\n    print(classification_report(y_true, y_pred, digits=4))\n\n    return {\n        \"name\": name, \"threshold\": threshold,\n        \"amex\": amex, \"auc\": auc, \"pr_auc\": ap, \"brier\": brier,\n        \"accuracy\": acc, \"precision\": prec, \"recall\": rec, \"f1\": f1,\n        \"tn\": int(cm[0,0]), \"fp\": int(cm[0,1]), \"fn\": int(cm[1,0]), \"tp\": int(cm[1,1])\n    }\n\ndef best_threshold_by_f1(y_true, y_prob):\n    p, r, th = precision_recall_curve(y_true, y_prob)\n    f1 = 2 * p * r / (p + r + 1e-12)\n    best_idx = int(np.nanargmax(f1))\n    best_th = float(th[max(best_idx - 1, 0)]) if len(th) else 0.5\n    return best_th, float(f1[best_idx])\n\ndef plot_roc_pr(y_true, y_prob, title_prefix=\"Model\"):\n    fpr, tpr, _ = roc_curve(y_true, y_prob)\n    p, r, _ = precision_recall_curve(y_true, y_prob)\n\n    plt.figure(figsize=(6,5))\n    plt.plot(fpr, tpr)\n    plt.plot([0,1],[0,1], linestyle=\"--\")\n    plt.title(f\"{title_prefix} ROC Curve\")\n    plt.xlabel(\"False Positive Rate\")\n    plt.ylabel(\"True Positive Rate\")\n    plt.tight_layout()\n    plt.show()\n\n    plt.figure(figsize=(6,5))\n    plt.plot(r, p)\n    plt.title(f\"{title_prefix} Precision-Recall Curve\")\n    plt.xlabel(\"Recall\")\n    plt.ylabel(\"Precision\")\n    plt.tight_layout()\n    plt.show()","metadata":{"_uuid":"31d4248e-965a-438a-a2cd-2b400573b001","_cell_guid":"6d9a963a-66d5-485c-89ce-37d348a85baa","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:19:04.621358Z","iopub.execute_input":"2026-06-05T16:19:04.622144Z","iopub.status.idle":"2026-06-05T16:19:04.636321Z","shell.execute_reply.started":"2026-06-05T16:19:04.622116Z","shell.execute_reply":"2026-06-05T16:19:04.635461Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6) Boosting Models and Parameter tuning\n\nWe tune with early stopping and evaluate using:\n- AMEX\n- AUC\n- PR-AUC\n- Brier","metadata":{"_uuid":"58e66971-60b7-483b-9450-c39abcbb06d0","_cell_guid":"90090a83-5820-4142-a36b-56197b7d8043","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"## A. LightGBM Pipeline","metadata":{"_uuid":"41548e07-f6e0-4aa4-b833-01f78cde4b5b","_cell_guid":"123adfcb-36b8-44a4-9833-9218ed0c5afa","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# LIGHTGBM PIPELINE\n# ============================================================\n\nimport re\nimport lightgbm as lgb\nimport numpy as np\nimport pandas as pd\nimport time\nimport gc\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nfrom sklearn.metrics import (\n    roc_auc_score, average_precision_score, brier_score_loss,\n    accuracy_score, precision_score, recall_score, f1_score,\n    confusion_matrix\n)\n\n# ============================================================\n# 0) Preconditions\n# ============================================================\nrequired = [\n    \"X_train\", \"X_valid\", \"X_test\", \"y_train\", \"y_valid\",\n    \"amex_metric\", \"best_threshold_by_f1\",\n    \"eval_classification\", \"plot_roc_pr\", \"scale_pos\", \"SEED\"\n]\nmissing = [k for k in required if k not in globals()]\nif missing:\n    raise NameError(f\"Missing required objects: {missing}\")\n\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\nOUT_DIR.mkdir(parents=True, exist_ok=True)\nFEATURE_PATH = OUT_DIR / \"feature_names.npy\"\n\n# ============================================================\n# 1) REAL FEATURE NAMES\n# ============================================================\n\nfeature_names = None\n\nif \"FEATURE_COLS\" in globals() and FEATURE_COLS is not None:\n    feature_names = list(FEATURE_COLS)\n    print(f\"[INFO] Using FEATURE_COLS with {len(feature_names)} names\")\n\nelif hasattr(X_train, \"columns\"):\n    feature_names = list(X_train.columns)\n    print(f\"[INFO] Using X_train DataFrame columns with {len(feature_names)} names\")\n\nelif FEATURE_PATH.exists():\n    feature_names = list(np.load(FEATURE_PATH, allow_pickle=True))\n    print(f\"[INFO] Loaded feature names from {FEATURE_PATH}\")\n\nelse:\n    raise RuntimeError(\n        \" Real feature names are not available.\\n\\n\"\n        \"Run ONE of these before this pipeline:\\n\"\n        \"1) FEATURE_COLS = list(original_dataframe.columns_after_final_feature_engineering)\\n\"\n        \"or\\n\"\n        \"2) np.save('/kaggle/working/seq_artifacts/feature_names.npy', \"\n        \"original_dataframe.columns_after_final_feature_engineering.values)\\n\\n\"\n        \"This pipeline intentionally refuses to use meaningless feature_* names.\"\n    )\n\n# Convert to strings + trim spaces\nfeature_names = [str(c).strip() for c in feature_names]\n\n# Detect generic fallback names like feature_0, feature_1, ...\ngeneric_flags = [bool(re.fullmatch(r\"feature_\\d+\", f)) for f in feature_names]\nif all(generic_flags):\n    raise RuntimeError(\n        \" The resolved feature names are still generic (feature_0, feature_1, ...).\\n\\n\"\n        \"That means the pipeline is loading fallback names, not the real AMEX feature names.\\n\"\n        \"Go back to the step where your final design matrix was still a DataFrame and run:\\n\\n\"\n        \"FEATURE_COLS = list(X_train.columns)\\n\"\n        \"np.save('/kaggle/working/seq_artifacts/feature_names.npy', FEATURE_COLS)\\n\\n\"\n        \"Then re-run this pipeline.\"\n    )\n\n# Persist clean\nnp.save(FEATURE_PATH, np.array(feature_names, dtype=object))\nprint(f\"[INFO] Saved verified feature names to {FEATURE_PATH}\")\n\n# ============================================================\n# 2) SANITY CHECK FEATURE COUNT\n# ============================================================\nn_features = X_train.shape[1]\nif len(feature_names) != n_features:\n    raise RuntimeError(\n        f\" Feature name count mismatch: {len(feature_names)} names vs {n_features} columns.\\n\"\n        \"Ensure FEATURE_COLS / feature_names.npy come from the SAME final design matrix used here.\"\n    )\n\n# ============================================================\n# 3) MEMORY OPTIMISATION\n# ============================================================\n# Preserve names first, then convert arrays\nX_train_np = X_train.astype(np.float32)\nX_valid_np = X_valid.astype(np.float32)\nX_test_np  = X_test.astype(np.float32)\n\n# ============================================================\n# 4) TRAINING FUNCTION\n# ============================================================\ndef train_lgb_once(params, X_tr, y_tr, X_va, y_va):\n    dtr = lgb.Dataset(X_tr, label=y_tr, free_raw_data=False)\n    dva = lgb.Dataset(X_va, label=y_va, reference=dtr, free_raw_data=False)\n\n    model = lgb.train(\n        params,\n        dtr,\n        num_boost_round=2000,\n        valid_sets=[dva],\n        valid_names=[\"valid\"],\n        callbacks=[\n            lgb.early_stopping(100),\n            lgb.log_evaluation(100)\n        ]\n    )\n\n    pred = model.predict(X_va, num_iteration=model.best_iteration)\n    score = amex_metric(y_va, pred)\n\n    return model, pred, score\n\n# ============================================================\n# 5) PARAMETER SAMPLING\n# ============================================================\ndef sample_lgb_params(seed):\n    rng = np.random.default_rng(seed)\n\n    return {\n        \"objective\": \"binary\",\n        \"metric\": \"auc\",\n        \"boosting_type\": \"gbdt\",\n\n        # faster convergence\n        \"learning_rate\": float(rng.choice([0.03, 0.05, 0.07])),\n\n        # AMEX-friendly ranges\n        \"num_leaves\": int(rng.choice([64, 96, 128])),\n        \"min_data_in_leaf\": int(rng.choice([500, 1000, 2000])),\n\n        # regularisation / subsampling\n        \"feature_fraction\": float(rng.choice([0.7, 0.8])),\n        \"bagging_fraction\": float(rng.choice([0.7, 0.8])),\n        \"bagging_freq\": 1,\n\n        \"lambda_l2\": float(rng.choice([1.0, 2.0, 5.0])),\n\n        # imbalance handling\n        \"scale_pos_weight\": float(scale_pos),\n\n        # extra controls\n        \"max_bin\": 255,\n        \"min_gain_to_split\": 0.01,\n\n        # stability\n        \"verbosity\": -1,\n        \"seed\": seed,\n        \"n_jobs\": 4\n    }\n\n# ============================================================\n# 6) TUNING LOOP\n# ============================================================\nN_TRIALS_LGB = 3\n\nbest_lgb = None\nbest_lgb_score = -1\nbest_lgb_pred = None\nbest_lgb_params = None\n\nstart = time.time()\n\nfor t in range(N_TRIALS_LGB):\n    print(f\"\\n--- Starting Trial {t+1}/{N_TRIALS_LGB} ---\")\n\n    params = sample_lgb_params(SEED + t)\n\n    model, pred, score = train_lgb_once(\n        params, X_train_np, y_train, X_valid_np, y_valid\n    )\n\n    print(f\"[LGB {t+1}] AMEX={score:.5f} | iter={model.best_iteration}\")\n\n    # early skip weak configs\n    if t >= 2 and score < best_lgb_score - 0.01:\n        print(\"Skipping weak configuration\")\n        del model\n        gc.collect()\n        continue\n\n    if score > best_lgb_score:\n        best_lgb_score = score\n        best_lgb = model\n        best_lgb_pred = pred\n        best_lgb_params = params\n\n    del model\n    gc.collect()\n\nprint(\"\\nBest LGB AMEX:\", best_lgb_score)\nprint(\"Best LGB params:\", best_lgb_params)\nprint(\"Tuning time (min):\", (time.time() - start) / 60)\n\nif best_lgb is None or best_lgb_pred is None:\n    raise RuntimeError(\" No valid LightGBM model was retained.\")\n\n# ============================================================\n# 7) EVALUATION\n# ============================================================\nth, f1_val = best_threshold_by_f1(y_valid, best_lgb_pred)\n\nprint(f\"Best threshold (F1): {th:.4f} | F1: {f1_val:.5f}\")\n\neval_classification(y_valid, best_lgb_pred, threshold=th, name=\"LightGBM\")\nplot_roc_pr(y_valid, best_lgb_pred, \"LightGBM\")\n\n# Save validation outputs\nnp.save(OUT_DIR / \"lgb_val_prob.npy\", best_lgb_pred.astype(np.float32))\nnp.save(OUT_DIR / \"lgb_val_true.npy\", np.asarray(y_valid).astype(np.int8))\n\n# Structured metrics\ny_hat = (best_lgb_pred >= th).astype(np.uint8)\ntn, fp, fn, tp = confusion_matrix(y_valid, y_hat).ravel()\n\nlgb_metrics = {\n    \"name\": \"LightGBM\",\n    \"amex\": float(amex_metric(y_valid, best_lgb_pred)),\n    \"auc\": float(roc_auc_score(y_valid, best_lgb_pred)),\n    \"pr_auc\": float(average_precision_score(y_valid, best_lgb_pred)),\n    \"brier\": float(brier_score_loss(y_valid, best_lgb_pred)),\n    \"accuracy\": float(accuracy_score(y_valid, y_hat)),\n    \"precision\": float(precision_score(y_valid, y_hat, zero_division=0)),\n    \"recall\": float(recall_score(y_valid, y_hat, zero_division=0)),\n    \"f1\": float(f1_score(y_valid, y_hat, zero_division=0)),\n    \"tn\": int(tn),\n    \"fp\": int(fp),\n    \"fn\": int(fn),\n    \"tp\": int(tp),\n    \"threshold\": float(th)\n}\n\nprint(\"\\n[LightGBM Metrics]\")\nprint({\n    \"amex\": round(lgb_metrics[\"amex\"], 5),\n    \"auc\": round(lgb_metrics[\"auc\"], 5),\n    \"pr_auc\": round(lgb_metrics[\"pr_auc\"], 5),\n    \"brier\": round(lgb_metrics[\"brier\"], 5)\n})\n\n# ============================================================\n# 8) INTERPRETABILITY\n# ============================================================\n# Gain importance\nimp = best_lgb.feature_importance(importance_type=\"gain\")\nimp = imp / (np.sum(imp) + 1e-8)\n\n# Top features\ntopk = min(15, len(imp))\ntop_idx = np.argsort(imp)[::-1][:topk]\n\nplt.figure(figsize=(10, 5))\nplt.bar(range(topk), imp[top_idx])\nplt.xticks(range(topk), [feature_names[i] for i in top_idx], rotation=45, ha=\"right\")\nplt.ylabel(\"Normalized importance\")\nplt.title(\"LightGBM: Top Feature Importance (Gain)\")\nplt.tight_layout()\nplt.show()\n\n# ============================================================\n# 9) COVERAGE + ENTROPY\n# ============================================================\nsorted_imp = np.sort(imp)[::-1]\ncum = np.cumsum(sorted_imp)\n\nk80 = int(np.argmax(cum >= 0.80) + 1)\nk90 = int(np.argmax(cum >= 0.90) + 1)\n\nplt.figure(figsize=(7, 4))\nplt.plot(cum, marker=\"o\")\nplt.axhline(0.80, linestyle=\"--\", alpha=0.6, label=\"80%\")\nplt.axhline(0.90, linestyle=\"--\", alpha=0.6, label=\"90%\")\nplt.axvline(k80, linestyle=\":\", alpha=0.6, label=f\"k80={k80}\")\nplt.axvline(k90, linestyle=\":\", alpha=0.6, label=f\"k90={k90}\")\nplt.title(\"LightGBM: Feature Coverage Curve\")\nplt.xlabel(\"Top-k features\")\nplt.ylabel(\"Cumulative importance\")\nplt.grid(alpha=0.3)\nplt.legend()\nplt.tight_layout()\nplt.show()\n\np = imp / (imp.sum() + 1e-8)\nentropy = -np.sum(p * np.log(p + 1e-8))\ntop3 = float(np.sort(p)[-3:].sum())\ntop5 = float(np.sort(p)[-5:].sum())\n\nprint(\"\\n[LightGBM Interpretability]\")\nprint(\"Entropy:\", round(entropy, 4))\nprint(\"Top-3 mass:\", round(top3, 4))\nprint(\"Top-5 mass:\", round(top5, 4))\nprint(\"k80:\", k80, \"| k90:\", k90)\n\n# ============================================================\n# 10) GROUPED IMPORTANCE (AMEX FEATURE FAMILIES)\n# ============================================================\n\nhas_amex_style = any(\"_\" in f for f in feature_names)\n\nif has_amex_style:\n    groups = [f.split(\"_\")[0] if \"_\" in f else \"OTHER\" for f in feature_names]\n\n    group_imp = {}\n    for g in sorted(set(groups)):\n        group_imp[g] = sum(\n            imp[i] for i in range(len(imp)) if groups[i] == g\n        )\n\n    print(\"\\n[Grouped Feature Importance]\")\n    for k, v in sorted(group_imp.items(), key=lambda x: -x[1]):\n        print(f\"{k}: {round(v, 4)}\")\n\n    plt.figure(figsize=(6, 4))\n    plt.bar(group_imp.keys(), group_imp.values())\n    plt.title(\"LightGBM: Grouped Feature Importance\")\n    plt.ylabel(\"Normalized importance\")\n    plt.tight_layout()\n    plt.show()\n\n    # Save grouped importance as CSV\n    group_imp_df = (\n        pd.DataFrame({\"Group\": list(group_imp.keys()), \"Importance\": list(group_imp.values())})\n        .sort_values(\"Importance\", ascending=False)\n        .reset_index(drop=True)\n    )\n    group_imp_df.to_csv(OUT_DIR / \"lgb_grouped_importance.csv\", index=False)\nelse:\n    print(\n        \"\\n[Grouped Feature Importance] skipped because feature names do not contain AMEX-style prefixes \"\n        \"(e.g., D_, B_, P_, S_).\"\n    )\n\n# ============================================================\n# 11) SAVE INTERPRETABILITY ARTIFACTS\n# ============================================================\nnp.save(OUT_DIR / \"lgb_feature_importance.npy\", imp.astype(np.float32))\nnp.save(OUT_DIR / \"lgb_entropy.npy\", np.array([entropy], dtype=np.float32))\nnp.save(OUT_DIR / \"lgb_topk_mass.npy\", np.array([top3, top5], dtype=np.float32))\nnp.save(OUT_DIR / \"lgb_k_coverage.npy\", np.array([k80, k90], dtype=np.int32))\n\n# Save full ranked feature table for traceability\nranked_idx = np.argsort(imp)[::-1]\nfeature_table = pd.DataFrame({\n    \"feature\": [feature_names[i] for i in ranked_idx],\n    \"importance\": imp[ranked_idx]\n})\nfeature_table.to_csv(OUT_DIR / \"lgb_feature_importance_table.csv\", index=False)\n\n# ============================================================\n# 12) TEST PREDICTION\n# ============================================================\nlgb_test_pred = best_lgb.predict(\n    X_test_np,\n    num_iteration=best_lgb.best_iteration\n)\nnp.save(OUT_DIR / \"lgb_test_pred.npy\", lgb_test_pred.astype(np.float32))\n\nprint(\"\\n LightGBM pipeline complete\")\nprint(\"Artifacts saved to:\", OUT_DIR)","metadata":{"_uuid":"759cab24-4f8c-4a8d-a145-b19df050fde5","_cell_guid":"5b4478b7-d7dd-49cd-9078-50bb833aed85","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:19:30.48405Z","iopub.execute_input":"2026-06-05T16:19:30.484797Z","iopub.status.idle":"2026-06-05T16:38:59.357607Z","shell.execute_reply.started":"2026-06-05T16:19:30.484764Z","shell.execute_reply":"2026-06-05T16:38:59.355798Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## B. XGBoost Pipeline","metadata":{"_uuid":"1b900d84-8607-4102-bb37-8c139e7da19c","_cell_guid":"e3fec661-890f-4f35-8cb7-0328047125e3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# XGBOOST PIPELINE\n# ============================================================\n\nimport re\nimport xgboost as xgb\nimport numpy as np\nimport pandas as pd\nimport time\nimport gc\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nfrom sklearn.metrics import (\n    roc_auc_score, average_precision_score, brier_score_loss,\n    accuracy_score, precision_score, recall_score, f1_score,\n    confusion_matrix\n)\n\n# ============================================================\n# 0) Preconditions\n# ============================================================\nrequired = [\n    \"X_train\", \"X_valid\", \"X_test\", \"y_train\", \"y_valid\",\n    \"amex_metric\", \"best_threshold_by_f1\",\n    \"eval_classification\", \"plot_roc_pr\", \"scale_pos\", \"SEED\"\n]\nmissing = [k for k in required if k not in globals()]\nif missing:\n    raise NameError(f\"Missing required objects: {missing}\")\n\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\nOUT_DIR.mkdir(parents=True, exist_ok=True)\nFEATURE_PATH = OUT_DIR / \"feature_names.npy\"\n\n# ============================================================\n# 1) FEATURE NAMES (STRICT)\n# ============================================================\n\nif \"FEATURE_COLS\" in globals() and FEATURE_COLS is not None:\n    feature_names = list(FEATURE_COLS)\n\nelif hasattr(X_train, \"columns\"):\n    feature_names = list(X_train.columns)\n\nelif FEATURE_PATH.exists():\n    feature_names = list(np.load(FEATURE_PATH, allow_pickle=True))\n\nelse:\n    raise RuntimeError(\" Missing real feature names\")\n\nfeature_names = [str(c).strip() for c in feature_names]\n\nif all(bool(re.fullmatch(r\"feature_\\d+\", f)) for f in feature_names):\n    raise RuntimeError(\" feature_* detected\")\n\nnp.save(FEATURE_PATH, np.array(feature_names, dtype=object))\n\n# ============================================================\n# 2) MEMORY OPTIMISATION\n# ============================================================\n\nX_train_np = X_train.astype(np.float32)\nX_valid_np = X_valid.astype(np.float32)\nX_test_np  = X_test.astype(np.float32)\n\n# CREATE DMATRIX ONCE\ndtrain = xgb.DMatrix(X_train_np, label=y_train)\ndvalid = xgb.DMatrix(X_valid_np, label=y_valid)\n\ndel X_train_np, X_valid_np\ngc.collect()\n\n# ============================================================\n# 3) TRAIN FUNCTION\n# ============================================================\n\ndef train_xgb_once(params):\n\n    model = xgb.train(\n        params,\n        dtrain,\n        num_boost_round=2000,\n        evals=[(dvalid, \"valid\")],\n        early_stopping_rounds=100, \n        verbose_eval=100,\n        maximize=True            \n    )\n\n    pred = model.predict(\n        dvalid,\n        iteration_range=(0, model.best_iteration + 1)\n    )\n\n    score = amex_metric(y_valid, pred)\n\n    return model, pred, score\n\n# ============================================================\n# 4) PARAM SAMPLING\n# ============================================================\n\ndef sample_xgb_params(seed):\n\n    rng = np.random.default_rng(seed)\n\n    return {\n        \"objective\": \"binary:logistic\",\n        \"eval_metric\": \"auc\",\n        \"tree_method\": \"hist\",\n\n        \"eta\": float(rng.choice([0.03, 0.05])),\n\n        \"max_depth\": int(rng.choice([5, 6])),\n\n        \"subsample\": 0.8,\n        \"colsample_bytree\": 0.8,\n\n        \"lambda\": 1.0,\n        \"alpha\": 0.0,\n\n        \"min_child_weight\": 10.0,\n        \"scale_pos_weight\": float(scale_pos),\n\n        \"seed\": seed,\n        \"nthread\": 4\n    }\n\n# ============================================================\n# 5) TRAIN LOOP\n# ============================================================\n\nN_TRIALS_XGB = 3 \n\nbest_xgb = None\nbest_score = -1\nbest_pred = None\nbest_params = None\n\nstart = time.time()\n\nfor t in range(N_TRIALS_XGB):\n\n    print(f\"\\n--- Starting Trial {t+1}/{N_TRIALS_XGB} ---\")\n\n    params = sample_xgb_params(SEED + t)\n\n    model, pred, score = train_xgb_once(params)\n\n    print(f\"[XGB {t+1}] AMEX={score:.5f} | iter={model.best_iteration}\")\n\n    if score > best_score:\n        best_score = score\n        best_xgb = model\n        best_pred = pred\n        best_params = params\n\n    gc.collect()\n\nprint(\"\\n Best XGB AMEX:\", best_score)\nprint(\"Best params:\", best_params)\nprint(\"Time (min):\", (time.time() - start) / 60)\n\n# ============================================================\n# 6) EVALUATION\n# ============================================================\n\nth, f1_val = best_threshold_by_f1(y_valid, best_pred)\n\nprint(f\"Best threshold: {th:.4f} | F1={f1_val:.5f}\")\n\neval_classification(y_valid, best_pred, threshold=th, name=\"XGBoost\")\nplot_roc_pr(y_valid, best_pred, \"XGBoost\")\n\nnp.save(OUT_DIR / \"xgb_val_prob.npy\", best_pred.astype(np.float32))\n\n# ============================================================\n# 7) INTERPRETABILITY\n# ============================================================\n\nimp_dict = best_xgb.get_score(importance_type=\"gain\")\n\nfull_imp = np.zeros(len(feature_names))\nfor k, v in imp_dict.items():\n    full_imp[int(k[1:])] = v\n\nfull_imp /= (full_imp.sum() + 1e-8)\n\ntop_idx = np.argsort(full_imp)[::-1][:15]\n\nplt.figure(figsize=(10,5))\nplt.bar(range(15), full_imp[top_idx])\nplt.xticks(range(15), [feature_names[i] for i in top_idx], rotation=45)\nplt.title(\"XGBoost Feature Importance\")\nplt.show()\n\nprint(\"\\n XGBoost pipeline complete\")","metadata":{"_uuid":"a308d60d-1908-4042-aa74-3368a0d4d8cd","_cell_guid":"aec22081-eaed-41d2-ac23-7d0f29a07968","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T16:38:59.358916Z","iopub.execute_input":"2026-06-05T16:38:59.359151Z","iopub.status.idle":"2026-06-05T17:02:29.423772Z","shell.execute_reply.started":"2026-06-05T16:38:59.359129Z","shell.execute_reply":"2026-06-05T17:02:29.422971Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7) Sequential Models (Transformer / \"Informer-style baseline\")\n\nWhy we do this:\n- The tree baselines use **aggregated** (customer-level) features.\n- Transformers/TFT require **sequences** to model temporal dependencies.\n\nI build a fixed-length sequence of 13 statements per customer (pad-left with zeros for shorter histories) and scale features for deep learning.","metadata":{"_uuid":"60d24268-62c8-4b2e-b3e4-bb338c6a8378","_cell_guid":"fefced27-1bd1-4044-b190-1d4556a28ba0","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# TRUE LOW-MEMORY SEQUENCE PIPELINE (CHUNKED BUILD)\n# - no giant feature DataFrame copy\n# - sort only ID/date arrays\n# - build memmap in chunks\n# - chunked scaling\n# ============================================================\n\nimport os, gc\nimport numpy as np\nimport pandas as pd\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import train_test_split\n\n# ============================================================\n# CONFIG\n# ============================================================\nSEQ_LEN = 13\nOUT_DIR = \"/kaggle/working/seq_artifacts\"\nos.makedirs(OUT_DIR, exist_ok=True)\n\nTRAIN_PATH = os.path.join(OUT_DIR, \"X_train_seq.dat\")\nTEST_PATH  = os.path.join(OUT_DIR, \"X_test_seq.dat\")\nY_PATH     = os.path.join(OUT_DIR, \"y_seq.npy\")\nMEAN_PATH  = os.path.join(OUT_DIR, \"scaler_mean.npy\")\nSCALE_PATH = os.path.join(OUT_DIR, \"scaler_scale.npy\")\nTRAIN_IDS_PATH = os.path.join(OUT_DIR, \"train_ids_seq.npy\")\nTEST_IDS_PATH  = os.path.join(OUT_DIR, \"test_ids_seq.npy\")\n\nROW_CHUNK = 100_000          # rows loaded at a time for sequence build\nSCALE_CHUNK_ROWS = 200_000   # rows loaded at a time for scaling\n\n# ============================================================\n# FEATURES (numeric only for sequence tensor)\n# ============================================================\nSEQ_FEATURES = [c for c in train.columns if c not in [ID_COL, DATE_COL, TARGET]]\nSEQ_FEATURES = [c for c in SEQ_FEATURES if pd.api.types.is_numeric_dtype(train[c])]\n\nif len(SEQ_FEATURES) == 0:\n    raise ValueError(\"No numeric sequence features available. Encode categoricals before sequence building.\")\n\nprint(f\"[INFO] Using {len(SEQ_FEATURES)} numeric sequence features\", flush=True)\n\n\n# ============================================================\n# HELPERS\n# ============================================================\ndef get_sorted_order(df, id_col, date_col):\n    \"\"\"\n    Returns row order sorted by [id_col, date_col] using only these two columns.\n    Much lighter than sorting a wide DataFrame.\n    \"\"\"\n    print(\"[sort] extracting ID/date arrays...\", flush=True)\n    ids = df[id_col].to_numpy()\n    dates = pd.to_datetime(df[date_col]).astype(\"int64\").to_numpy()\n\n    print(\"[sort] computing lexsort order...\", flush=True)\n    order = np.lexsort((dates, ids))  # primary=id, secondary=date\n    sorted_ids = ids[order]\n\n    return order, sorted_ids\n\ndef get_group_boundaries(sorted_ids):\n    \"\"\"\n    From sorted IDs, compute unique IDs and contiguous group sizes.\n    \"\"\"\n    n = len(sorted_ids)\n    if n == 0:\n        return np.array([], dtype=sorted_ids.dtype), np.array([], dtype=np.int64)\n\n    change = np.empty(n, dtype=bool)\n    change[0] = True\n    change[1:] = sorted_ids[1:] != sorted_ids[:-1]\n\n    start_idx = np.flatnonzero(change)\n    counts = np.diff(np.append(start_idx, n))\n    unique_ids = sorted_ids[start_idx]\n\n    return unique_ids, counts\n\ndef build_sequences_chunked(df, id_col, date_col, feature_cols, out_path, seq_len=13, row_chunk=100_000):\n    \"\"\"\n    Build sequence memmap without copying all features at once.\n    Only loads chunked slices from the source DataFrame.\n    \"\"\"\n    # Sort by ID/date without sorting the full 188-feature frame\n    order, sorted_ids = get_sorted_order(df, id_col, date_col)\n    unique_ids, counts = get_group_boundaries(sorted_ids)\n\n    n_groups = len(unique_ids)\n    n_features = len(feature_cols)\n\n    print(f\"[build] groups={n_groups:,}, seq_len={seq_len}, n_features={n_features}\", flush=True)\n\n    X = np.memmap(\n        out_path,\n        mode=\"w+\",\n        shape=(n_groups, seq_len, n_features),\n        dtype=np.float32\n    )\n\n    ids_out = np.empty(n_groups, dtype=object)\n\n    # rolling state across chunks\n    current_gid = 0\n    current_id = unique_ids[0] if n_groups > 0 else None\n    buf = np.zeros((seq_len, n_features), dtype=np.float32)\n    count_in_group = 0\n    rows_seen_in_group = 0\n    target_group_size = counts[0] if n_groups > 0 else 0\n\n    print(\"[build] streaming rows in chunks...\", flush=True)\n\n    for start in range(0, len(order), row_chunk):\n        end = min(start + row_chunk, len(order))\n        idx_chunk = order[start:end]\n\n        # load ONLY this chunk of rows and ONLY needed columns\n        chunk = df.iloc[idx_chunk][[id_col] + feature_cols]\n\n        for row in chunk.itertuples(index=False, name=None):\n            rid = row[0]\n            vals = np.asarray(row[1:], dtype=np.float32)\n\n            # sanity: if ID changed unexpectedly, flush current group\n            if rid != current_id:\n                # flush current customer\n                X[current_gid] = 0.0\n                if count_in_group >= seq_len:\n                    X[current_gid] = buf\n                else:\n                    X[current_gid, -count_in_group:] = buf[:count_in_group]\n                ids_out[current_gid] = current_id\n\n                # move to next customer\n                current_gid += 1\n                current_id = rid\n                buf[:] = 0.0\n                count_in_group = 0\n                rows_seen_in_group = 0\n                target_group_size = counts[current_gid]\n\n            # rolling fill\n            if count_in_group < seq_len:\n                buf[count_in_group] = vals\n                count_in_group += 1\n            else:\n                buf[:-1] = buf[1:]\n                buf[-1] = vals\n\n            rows_seen_in_group += 1\n\n        del chunk\n        gc.collect()\n\n        print(f\"[build] processed rows {end:,}/{len(order):,}\", flush=True)\n\n    # flush last group\n    if n_groups > 0:\n        X[current_gid] = 0.0\n        if count_in_group >= seq_len:\n            X[current_gid] = buf\n        else:\n            X[current_gid, -count_in_group:] = buf[:count_in_group]\n        ids_out[current_gid] = current_id\n\n    X.flush()\n    del order, sorted_ids, counts, unique_ids, buf\n    gc.collect()\n\n    return ids_out, n_groups, n_features\n\ndef fit_scaler_chunked_memmap(memmap_3d, n_features, chunk_rows=200_000):\n    scaler = StandardScaler()\n    X2d = memmap_3d.reshape(-1, n_features)\n    total_rows = X2d.shape[0]\n\n    for start in range(0, total_rows, chunk_rows):\n        end = min(start + chunk_rows, total_rows)\n        chunk = X2d[start:end]\n        mask = ~(chunk == 0).all(axis=1)\n        if mask.any():\n            scaler.partial_fit(chunk[mask])\n\n        if start % (chunk_rows * 10) == 0:\n            print(f\"[fit_scaler] rows {start:,}-{end:,}/{total_rows:,}\", flush=True)\n\n    return scaler\n\ndef transform_memmap_inplace(memmap_3d, scaler, n_features, chunk_rows=200_000):\n    X2d = memmap_3d.reshape(-1, n_features)\n    total_rows = X2d.shape[0]\n\n    for start in range(0, total_rows, chunk_rows):\n        end = min(start + chunk_rows, total_rows)\n        chunk = X2d[start:end]\n        mask = ~(chunk == 0).all(axis=1)\n\n        if mask.any():\n            chunk[mask] -= scaler.mean_\n            chunk[mask] /= scaler.scale_\n\n        chunk[~mask] = 0\n\n        if start % (chunk_rows * 10) == 0:\n            print(f\"[transform] rows {start:,}-{end:,}/{total_rows:,}\", flush=True)\n\n    memmap_3d.flush()\n\n\n# ============================================================\n# STEP 1: BUILD TRAIN SEQUENCES\n# ============================================================\nprint(\"Building TRAIN sequences...\", flush=True)\ntrain_ids_seq, N_train, D = build_sequences_chunked(\n    train, ID_COL, DATE_COL, SEQ_FEATURES, TRAIN_PATH, seq_len=SEQ_LEN, row_chunk=ROW_CHUNK\n)\nnp.save(TRAIN_IDS_PATH, train_ids_seq)\n\n# ============================================================\n# STEP 2: BUILD TEST SEQUENCES\n# ============================================================\nprint(\"Building TEST sequences...\", flush=True)\ntest_ids_seq, N_test, D_test = build_sequences_chunked(\n    test, ID_COL, DATE_COL, SEQ_FEATURES, TEST_PATH, seq_len=SEQ_LEN, row_chunk=ROW_CHUNK\n)\nnp.save(TEST_IDS_PATH, test_ids_seq)\n\nassert D == D_test, \"Train/test feature dimension mismatch.\"\n\n# ============================================================\n# STEP 3: BUILD TARGET\n# ============================================================\nprint(\"Building target vector...\", flush=True)\ny_map = (\n    train[[ID_COL, TARGET]]\n    .drop_duplicates(subset=[ID_COL])\n    .set_index(ID_COL)[TARGET]\n)\ny_seq = y_map.reindex(train_ids_seq).to_numpy(dtype=np.int64)\nnp.save(Y_PATH, y_seq)\n\ndel y_map\ngc.collect()\n\n# ============================================================\n# STEP 4: OPEN MEMMAPS\n# ============================================================\nX_train_seq = np.memmap(TRAIN_PATH, mode=\"r+\", dtype=np.float32, shape=(N_train, SEQ_LEN, D))\nX_test_seq  = np.memmap(TEST_PATH,  mode=\"r+\", dtype=np.float32, shape=(N_test,  SEQ_LEN, D))\n\n# ============================================================\n# STEP 5: FIT SCALER (CHUNKED)\n# ============================================================\nprint(\"Fitting scaler on TRAIN...\", flush=True)\nscaler = fit_scaler_chunked_memmap(X_train_seq, n_features=D, chunk_rows=SCALE_CHUNK_ROWS)\nnp.save(MEAN_PATH, scaler.mean_.astype(np.float32))\nnp.save(SCALE_PATH, scaler.scale_.astype(np.float32))\n\n# ============================================================\n# STEP 6: TRANSFORM TRAIN\n# ============================================================\nprint(\"Scaling TRAIN...\", flush=True)\ntransform_memmap_inplace(X_train_seq, scaler, n_features=D, chunk_rows=SCALE_CHUNK_ROWS)\n\n# ============================================================\n# STEP 7: TRANSFORM TEST\n# ============================================================\nprint(\"Scaling TEST...\", flush=True)\ntransform_memmap_inplace(X_test_seq, scaler, n_features=D, chunk_rows=SCALE_CHUNK_ROWS)\n\ndel X_train_seq, X_test_seq\ngc.collect()\n\n# ============================================================\n# STEP 8: REOPEN READ-ONLY + SPLIT\n# ============================================================\nprint(\"Preparing split...\", flush=True)\nX_train_seq = np.memmap(TRAIN_PATH, mode=\"r\", dtype=np.float32, shape=(N_train, SEQ_LEN, D))\ny_seq = np.load(Y_PATH)\n\nidx = np.arange(N_train, dtype=np.int32)\n\nidx_tr, idx_va, y_tr, y_va = train_test_split(\n    idx, y_seq, test_size=0.2, random_state=SEED, stratify=y_seq\n)\n\nprint(\"Train idx:\", len(idx_tr), \"Valid idx:\", len(idx_va), flush=True)\nprint(\"Final shapes:\", X_train_seq.shape, y_seq.shape, flush=True)\nprint(\"\\n PIPELINE READY FOR TRAINING\", flush=True)","metadata":{"_uuid":"a2e68860-7dec-47f0-b8ea-daade6f41ec0","_cell_guid":"bc60123c-bb39-4d72-ba5d-cbbc8f91a7c2","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T17:02:29.425095Z","iopub.execute_input":"2026-06-05T17:02:29.425378Z","iopub.status.idle":"2026-06-05T17:14:34.900209Z","shell.execute_reply.started":"2026-06-05T17:02:29.425355Z","shell.execute_reply":"2026-06-05T17:14:34.899639Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 15) Install dependencies (run once, then restart kernel)\n# ============================================================\n# NOTE: uncomment if not installed in your Kaggle environment.\n#!pip -q install \"pytorch-forecasting>=1.0.0\" \"pytorch-lightning>=2.0.0\" \"torchmetrics>=1.0.0\"","metadata":{"_uuid":"3f83cc7f-46e9-4ddb-8411-0c553db8841a","_cell_guid":"dbe4a94a-9d3b-4453-a983-faa323568d98","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 8) Transformer Pipeline","metadata":{"_uuid":"dc769142-a30a-4849-a91b-93afd3502eb1","_cell_guid":"56eb75f3-f890-4f77-8c06-6f3a72297305","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# TRANSFORMER PIPELINE\n# ============================================================\n\nimport re\nimport numpy as np\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom sklearn.model_selection import train_test_split\nfrom pathlib import Path\nimport matplotlib.pyplot as plt\n\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\nprint(\"DEVICE:\", DEVICE)\n\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\n\n# ============================================================\n# DATA\n# ============================================================\n\nSEQ_LEN = 13\nn_features = len(SEQ_FEATURES)\n\nX_train_seq = np.memmap(\n    OUT_DIR / \"X_train_seq.dat\",\n    mode=\"r\",\n    dtype=np.float32\n).reshape(-1, SEQ_LEN, n_features)\n\nX_test_seq = np.memmap(\n    OUT_DIR / \"X_test_seq.dat\",\n    mode=\"r\",\n    dtype=np.float32\n).reshape(-1, SEQ_LEN, n_features)\n\ny_seq = np.load(OUT_DIR / \"y_seq.npy\").astype(np.uint8)\n\n# ============================================================\n# STRICT FEATURE NAMES\n# ============================================================\nif \"SEQ_FEATURES\" in globals() and SEQ_FEATURES is not None:\n    FEATURE_NAMES = list(SEQ_FEATURES)\nelse:\n    raise RuntimeError(\n        \" Missing SEQ_FEATURES. Define:\\n\"\n        \"SEQ_FEATURES = list(original_sequence_columns)\"\n    )\n\nif len(FEATURE_NAMES) != n_features:\n    raise RuntimeError(\" Feature mismatch\")\n\nif all(bool(re.fullmatch(r\"feature_\\d+\", f)) for f in FEATURE_NAMES):\n    raise RuntimeError(\" feature_* detected — fix preprocessing\")\n\n# ============================================================\n# SPLIT\n# ============================================================\nidx = np.arange(len(X_train_seq))\nidx_tr, idx_va = train_test_split(idx, test_size=0.2, stratify=y_seq)\n\ny_va = y_seq[idx_va]\n\n# ============================================================\n# DATASET\n# ============================================================\nclass DatasetSeq(Dataset):\n    def __init__(self, X, idx, y):\n        self.X = X\n        self.idx = idx\n        self.y = y\n\n    def __len__(self):\n        return len(self.idx)\n\n    def __getitem__(self, i):\n        j = self.idx[i]\n        return (\n            torch.from_numpy(self.X[j]).float(),\n            torch.tensor(self.y[j]).float()\n        )\n\ntrain_loader = DataLoader(\n    DatasetSeq(X_train_seq, idx_tr, y_seq),\n    batch_size=128, shuffle=True\n)\n\nvalid_loader = DataLoader(\n    DatasetSeq(X_train_seq, idx_va, y_seq),\n    batch_size=256, shuffle=False\n)\n\n# ============================================================\n# MODEL\n# ============================================================\nclass TransformerCLS(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        d_model = 128\n\n        self.proj = nn.Linear(n_features, d_model)\n\n        self.cls = nn.Parameter(torch.randn(1, 1, d_model) * 0.02)\n\n        self.pos = nn.Parameter(torch.randn(1, SEQ_LEN + 1, d_model) * 0.02)\n\n        encoder_layer = nn.TransformerEncoderLayer(\n            d_model=d_model,\n            nhead=4,\n            dim_feedforward=256,\n            batch_first=True\n        )\n\n        self.encoder = nn.TransformerEncoder(encoder_layer, num_layers=2)\n\n        self.head = nn.Linear(d_model, 1)\n\n    def forward(self, x):\n        h = self.proj(x)\n\n        cls = self.cls.expand(h.size(0), -1, -1)\n        h = torch.cat([cls, h], dim=1)\n\n        h = h + self.pos[:, :h.size(1)]\n        h = self.encoder(h)\n\n        return self.head(h[:, 0]).squeeze(-1)\n\nmodel = TransformerCLS().to(DEVICE)\n\n# ============================================================\n# TRAIN\n# ============================================================\nloss_fn = nn.BCEWithLogitsLoss()\nopt = torch.optim.AdamW(model.parameters(), lr=1e-3)\n\nbest_score = -1\nbest_pred = None\n\nfor epoch in range(5):\n\n    model.train()\n\n    for xb, yb in train_loader:\n        xb, yb = xb.to(DEVICE), yb.to(DEVICE)\n\n        opt.zero_grad()\n        loss = loss_fn(model(xb), yb)\n        loss.backward()\n        opt.step()\n\n    # validation\n    model.eval()\n    preds = []\n\n    with torch.no_grad():\n        for xb, _ in valid_loader:\n            xb = xb.to(DEVICE)\n            preds.append(torch.sigmoid(model(xb)).cpu().numpy())\n\n    preds = np.concatenate(preds)\n\n    score = amex_metric(y_va, preds)\n    print(f\"Epoch {epoch} | AMEX={score:.5f}\")\n\n    if score > best_score:\n        best_score = score\n        best_pred = preds\n\n# ============================================================\n# METRICS\n# ============================================================\nth,_ = best_threshold_by_f1(y_va, best_pred)\n\nprint(\"\\n[Transformer Metrics]\")\nprint(\"AMEX:\", round(best_score,5))\n\nnp.save(OUT_DIR / \"trf_valid_pred.npy\", best_pred)\n\n# ============================================================\n# INTERPRETABILITY\n# ============================================================\n\ndef grad_importance(model, X):\n    xb = torch.tensor(X, dtype=torch.float32).to(DEVICE)\n    xb.requires_grad_()\n\n    out = torch.sigmoid(model(xb)).mean()\n    out.backward()\n\n    return (xb.grad.abs() * xb.abs()).mean(dim=0).detach().cpu().numpy()\n\nimp = grad_importance(model, X_train_seq[idx_va[:64]])\n\nimp = imp / (np.max(np.abs(imp)) + 1e-8)\n\nfeat_imp = imp.mean(axis=0)\ntime_imp = imp.mean(axis=1)\n\n# top features\ntopk = 15\ntop_idx = np.argsort(-feat_imp)[:topk]\n\nplt.figure(figsize=(10,5))\nplt.bar(range(topk), feat_imp[top_idx])\nplt.xticks(range(topk),\n           [FEATURE_NAMES[i] for i in top_idx],\n           rotation=45, ha=\"right\")\nplt.title(\"Transformer Feature Importance\")\nplt.tight_layout()\nplt.show()\n\n# time importance\nplt.plot(time_imp)\nplt.title(\"Time Importance\")\nplt.show()\n\n# ============================================================\n# TEMPORAL METRICS\n# ============================================================\np = time_imp / (time_imp.sum() + 1e-8)\nentropy = -np.sum(p * np.log(p + 1e-8))\n\ntop3 = float(np.sort(p)[-3:].sum())\ntop5 = float(np.sort(p)[-5:].sum())\n\nprint(\"\\n[Temporal Interpretability]\")\nprint(\"Entropy:\", round(entropy,4))\nprint(\"Top-3 mass:\", round(top3,4))\nprint(\"Top-5 mass:\", round(top5,4))\n\n# ============================================================\n# GROUPING\n# ============================================================\n\ngroups = [f.split(\"_\")[0] if \"_\" in f else f for f in FEATURE_NAMES]\n\ngroup_imp = {}\nfor g in set(groups):\n    group_imp[g] = sum(\n        feat_imp[i] for i in range(len(feat_imp)) if groups[i] == g\n    )\n\ntotal = sum(group_imp.values())\ngroup_imp = {k: v/total for k,v in group_imp.items()}\n\nprint(\"\\n[Grouped Feature Importance]\")\nfor k, v in sorted(group_imp.items(), key=lambda x: -x[1]):\n    print(f\"{k}: {round(v,4)}\")\n\nplt.figure(figsize=(6,4))\nplt.bar(group_imp.keys(), group_imp.values())\nplt.title(\"Grouped Feature Importance\")\nplt.show()\n\n# ============================================================\n# TEST\n# ============================================================\n\nmodel.eval()\ntest_preds = []\n\nwith torch.no_grad():\n    for i in range(0, len(X_test_seq), 256):\n        xb = torch.tensor(X_test_seq[i:i+256]).to(DEVICE)\n        test_preds.append(torch.sigmoid(model(xb)).cpu().numpy())\n\ntest_preds = np.concatenate(test_preds)\n\nnp.save(OUT_DIR / \"transformer_test_pred.npy\", test_preds)\n\nprint(\"\\n Transformer pipeline complete\")","metadata":{"_uuid":"805c59a4-36d2-43a0-96a3-36fb2051acd8","_cell_guid":"3ca005ce-b86f-4945-a9f8-3119c451b130","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T17:14:34.902895Z","iopub.execute_input":"2026-06-05T17:14:34.903124Z","iopub.status.idle":"2026-06-05T18:25:39.469616Z","shell.execute_reply.started":"2026-06-05T17:14:34.903104Z","shell.execute_reply":"2026-06-05T18:25:39.46829Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n# 9) Temporal Fusion Transformer (TFT) and Informer\n\nThis section is for **TFT** and **Informer** training/evaluation in my notebook.\n\nNotes:\n- TFT is implemented using **pytorch-forecasting** (PyTorch Lightning).\n- The informer is added as a **compact informer-style classifier** (multi-head attention encoder with a classification head).\n- Both sections include **imbalance handling**, **metrics**, and **interpretability without SHAP**.","metadata":{"_uuid":"1e86d86e-7309-492f-9da4-5176e9337ad1","_cell_guid":"254c3180-132b-4133-8bad-3de598ff0d24","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"# A) TFT Pipeline\n\n## Objective\nTrain a **Temporal Fusion Transformer (TFT)** for **multi-horizon default prediction** with **built-in interpretability**.\n\n\n## Data Design\n- **group_id** → customer/entity  \n- **time_idx** → sequence order  \n\n### Feature Types\n- **Static** → time-invariant (customer context)  \n- **Time-varying**:\n  - Known → `time_idx`\n  - Unknown → numeric + categorical features  \n\n## Model Setup\n- Encoder: `12`\n- Prediction horizon: `6` (multi-horizon)\n- Loss: Weighted cross-entropy  \n- Training: Early stopping + gradient accumulation  \n\n## Core TFT Components\n\n### 1. Variable Selection Networks (VSN)\n- Learns **feature importance**\n- Outputs:\n  - Feature × time importance\n  - Global ranking\n\n### 2. Temporal Attention\n- Learns **time-step importance**\n- Captures long-range dependencies\n\n### 3. Gating (GRN)\n- Filters noise\n- Controls information flow\n- Supports (not directly produces) interpretability  \n\n### 4. Static Context\n- Uses static features to condition predictions  \n- Enables **context-level insights**\n\n## Interpretability Outputs\n\n###  Feature Importance\n- `encoder_variables`\n- Feature × time heatmap\n\n###  Temporal Importance\n- `attention`\n- Time-step importance\n\n###  Static Importance\n- `static_variables`\n- Customer-level influence\n\n## Key Strengths\n\n-  Native interpretability (no post-hoc methods)  \n-  Multi-horizon forecasting  \n-  Handles static + temporal data jointly  \n\n## Final Conclusion\nTFT provides **built-in interpretability** via:\n- Variable selection → *feature importance*  \n- Attention → *time importance*  \n- Static context → *entity-level influence*  \n\n## One-Line Summary\nTFT delivers **full model-level interpretability (feature + time + context)** using native architecture components.","metadata":{"_uuid":"455fccb0-d77f-40fe-ae81-3b3ad3011519","_cell_guid":"8c6d0cc5-ad09-42f9-b851-b9db08896e2e","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# TFT PIPELINE\n# ============================================================\n\nimport re\nimport gc\nimport numpy as np\nimport pandas as pd\nimport torch\nimport torch.nn.functional as F\nimport lightning.pytorch as pl\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import (\n    roc_auc_score, average_precision_score, brier_score_loss,\n    accuracy_score, precision_score, recall_score, f1_score,\n    confusion_matrix\n)\nfrom lightning.pytorch.callbacks import EarlyStopping, ModelCheckpoint\nfrom pytorch_forecasting import TimeSeriesDataSet\nfrom pytorch_forecasting.models import TemporalFusionTransformer\nfrom pytorch_forecasting.data import NaNLabelEncoder\nfrom pytorch_forecasting.metrics import CrossEntropy\n\n# ============================================================\n# 0) Preconditions\n# ============================================================\nrequired = [\n    \"train\", \"ID_COL\", \"DATE_COL\", \"TARGET\", \"CAT_FEATURES\", \"SEED\",\n    \"amex_metric\", \"best_threshold_by_f1\", \"eval_classification\", \"plot_roc_pr\"\n]\nmissing = [k for k in required if k not in globals()]\nif missing:\n    raise NameError(f\"Missing required objects: {missing}\")\n\n# ============================================================\n# 1) CONFIG\n# ============================================================\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\nOUT_DIR.mkdir(parents=True, exist_ok=True)\n\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\npl.seed_everything(SEED)\n\nprint(\"[DEVICE]\", DEVICE)\n\n# Conservative settings for Kaggle / AMEX\nMAX_ENCODER_LENGTH = 12\nMAX_PRED_LENGTH    = 1\nMIN_PRED_IDX       = MAX_ENCODER_LENGTH\n\n# Reduce feature pressure to keep TFT manageable\nMAX_REAL   = 20\nMAX_CAT    = 3\nMAX_STATIC = 3\n\n# Smaller batch + accumulation = lower peak memory\nBATCH_SIZE  = 8\nACCUM_STEPS = 8\nEPOCHS      = 5\nPATIENCE    = 2\n\nNA_TOKEN = \"__NA__\"\n\n# ============================================================\n# 2) STRICT FEATURE NAMES\n# ============================================================\nFEATURE_PATH = OUT_DIR / \"feature_names.npy\"\n\nfeature_names = None\n\nif \"FEATURE_COLS\" in globals() and FEATURE_COLS is not None:\n    feature_names = list(FEATURE_COLS)\n    print(f\"[INFO] Using FEATURE_COLS with {len(feature_names)} names\")\n\nelif hasattr(train, \"columns\"):\n    # derive from current train dataframe\n    feature_names = [c for c in train.columns if c not in [ID_COL, DATE_COL, TARGET]]\n    print(f\"[INFO] Derived feature names from train DataFrame with {len(feature_names)} names\")\n\nelif FEATURE_PATH.exists():\n    feature_names = list(np.load(FEATURE_PATH, allow_pickle=True))\n    print(f\"[INFO] Loaded feature names from {FEATURE_PATH}\")\n\nelse:\n    raise RuntimeError(\n        \" Real feature names are not available.\\n\\n\"\n        \"Run ONE of these before this pipeline:\\n\"\n        \"1) FEATURE_COLS = list(original_dataframe.columns_after_final_feature_engineering)\\n\"\n        \"or\\n\"\n        \"2) np.save('/kaggle/working/seq_artifacts/feature_names.npy', \"\n        \"original_dataframe.columns_after_final_feature_engineering.values)\\n\\n\"\n        \"This pipeline intentionally refuses to use meaningless feature_* names.\"\n    )\n\nfeature_names = [str(f).strip() for f in feature_names]\n\ngeneric_flags = [bool(re.fullmatch(r\"feature_\\d+\", f)) for f in feature_names]\nif all(generic_flags):\n    raise RuntimeError(\n        \" The resolved feature names are still generic (feature_0, feature_1, ...).\\n\"\n        \"Go back to the step where your final design matrix was still a DataFrame and save real names.\"\n    )\n\nnp.save(FEATURE_PATH, np.array(feature_names, dtype=object))\nprint(f\"[INFO] Saved verified feature names to {FEATURE_PATH}\")\n\n# ============================================================\n# 3) FEATURE SPLIT\n# ============================================================\nall_cols = [c for c in train.columns if c not in [ID_COL, DATE_COL, TARGET]]\n\nnum_cols = [c for c in all_cols if pd.api.types.is_numeric_dtype(train[c])][:MAX_REAL]\ncat_cols = [c for c in CAT_FEATURES if c in train.columns][:MAX_CAT]\nstatic_candidates = (num_cols + cat_cols)[:MAX_STATIC]\n\nprint(f\"[COLS] num={len(num_cols)} cat={len(cat_cols)} static={len(static_candidates)}\")\n\n# Sanity: selected features must be real named columns\nselected_feature_names = num_cols + cat_cols\nif len(selected_feature_names) == 0:\n    raise RuntimeError(\" No usable TFT features selected.\")\n\n# ============================================================\n# 4) PREP DATAFRAME\n# ============================================================\ntft_df = train[[ID_COL, DATE_COL, TARGET] + selected_feature_names].copy()\n\ntft_df[DATE_COL] = pd.to_datetime(tft_df[DATE_COL])\ntft_df = tft_df.sort_values([ID_COL, DATE_COL]).reset_index(drop=True)\n\n# categorical handling\nfor c in cat_cols:\n    tft_df[c] = tft_df[c].astype(\"string\").fillna(NA_TOKEN).astype(\"category\")\n\n# target type\ntft_df[TARGET] = tft_df[TARGET].astype(np.int64)\n\n# group ids \nTRAIN_IDUNIQUES_PATH = OUT_DIR / \"train_id_uniques.npy\"\nif TRAIN_IDUNIQUES_PATH.exists():\n    train_id_uniques = np.load(TRAIN_IDUNIQUES_PATH, allow_pickle=True)\n    id_to_gid = {cid: i for i, cid in enumerate(train_id_uniques.tolist())}\n    tft_df = tft_df[tft_df[ID_COL].isin(id_to_gid)].copy()\n    tft_df[\"group_id\"] = tft_df[ID_COL].map(id_to_gid).astype(np.int32)\n    print(\"[GROUP] Using train_id_uniques.npy for stable mapping\")\nelse:\n    tft_df[\"group_id\"] = pd.factorize(tft_df[ID_COL])[0].astype(np.int32)\n    print(\"[GROUP] train_id_uniques.npy not found; using local factorization\")\n\ntft_df[\"time_idx\"] = tft_df.groupby(\"group_id\").cumcount().astype(np.int32)\n\n# Diagnostics\nlens_all = (tft_df.groupby(\"group_id\")[\"time_idx\"].max().astype(int) + 1)\nmin_required_len = MAX_ENCODER_LENGTH + MAX_PRED_LENGTH\nprint(\"[LEN] per-group length summary:\\n\", lens_all.describe())\nprint(\"[LEN] min_required_len =\", min_required_len)\n\n# ============================================================\n# 5) GROUP-LEVEL SPLIT\n# ============================================================\nid_target = tft_df[[\"group_id\", TARGET]].drop_duplicates(\"group_id\")\n\nif TRAIN_IDUNIQUES_PATH.exists() and all(k in globals() for k in [\"idx_tr\", \"idx_va\"]):\n    idx_tr_arr = np.asarray(idx_tr, dtype=np.int32)\n    idx_va_arr = np.asarray(idx_va, dtype=np.int32)\n\n    split_train_ids = set(train_id_uniques[idx_tr_arr].tolist())\n    split_valid_ids = set(train_id_uniques[idx_va_arr].tolist())\n\n    train_df = tft_df[tft_df[ID_COL].isin(split_train_ids)].copy()\n    valid_df = tft_df[tft_df[ID_COL].isin(split_valid_ids)].copy()\n    print(\"[SPLIT] Reused existing idx_tr / idx_va split\")\nelse:\n    tr_ids, va_ids = train_test_split(\n        id_target[\"group_id\"],\n        stratify=id_target[TARGET],\n        test_size=0.2,\n        random_state=SEED\n    )\n    train_df = tft_df[tft_df[\"group_id\"].isin(tr_ids)].copy()\n    valid_df = tft_df[tft_df[\"group_id\"].isin(va_ids)].copy()\n    print(\"[SPLIT] Created new stratified group split\")\n\n# Free original big frame early\ndel tft_df\ngc.collect()\n\n# ============================================================\n# 6) FILTER SHORT GROUPS\n# ============================================================\ntrain_len = (train_df.groupby(\"group_id\")[\"time_idx\"].max().astype(int) + 1)\nvalid_len = (valid_df.groupby(\"group_id\")[\"time_idx\"].max().astype(int) + 1)\n\ngood_train_ids = train_len[train_len >= min_required_len].index\ngood_valid_ids = valid_len[valid_len >= min_required_len].index\n\ntrain_df = train_df[train_df[\"group_id\"].isin(good_train_ids)].copy()\nvalid_df = valid_df[valid_df[\"group_id\"].isin(good_valid_ids)].copy()\n\nprint(f\"[FILTER] kept train groups={train_df['group_id'].nunique()} | kept valid groups={valid_df['group_id'].nunique()}\")\n\nif train_df[\"group_id\"].nunique() == 0:\n    raise RuntimeError(\n        \" ZERO training groups remain after filtering. \"\n        \"Reduce MAX_ENCODER_LENGTH or ensure longer per-group histories.\"\n    )\n\n# class weights\nid_target_f = train_df[[ \"group_id\", TARGET ]].drop_duplicates(\"group_id\")\npos = int((id_target_f[TARGET] == 1).sum())\nneg = int((id_target_f[TARGET] == 0).sum())\nprint(f\"[CLASS] train groups pos={pos} neg={neg}\")\n\n# ============================================================\n# 7) DATASET\n# ============================================================\ncat_encoders = {c: NaNLabelEncoder(add_nan=True) for c in cat_cols}\n\ntraining = TimeSeriesDataSet(\n    train_df,\n    time_idx=\"time_idx\",\n    target=TARGET,\n    group_ids=[\"group_id\"],\n\n    min_encoder_length=MAX_ENCODER_LENGTH,\n    max_encoder_length=MAX_ENCODER_LENGTH,\n    min_prediction_length=1,\n    max_prediction_length=MAX_PRED_LENGTH,\n    min_prediction_idx=MIN_PRED_IDX,\n\n    static_categoricals=[c for c in static_candidates if c in cat_cols],\n    static_reals=[c for c in static_candidates if c in num_cols],\n\n    # keep this minimal for memory\n    time_varying_known_reals=[\"time_idx\"],\n    time_varying_unknown_reals=num_cols,\n    time_varying_unknown_categoricals=cat_cols,\n\n    categorical_encoders=cat_encoders,\n    target_normalizer=None,\n    allow_missing_timesteps=True,\n)\n\nvalidation = TimeSeriesDataSet.from_dataset(\n    training,\n    valid_df,   # important: validation is ONLY valid_df, not full train\n    predict=True,\n    stop_randomization=True\n)\n\ntrain_loader = training.to_dataloader(\n    train=True,\n    batch_size=BATCH_SIZE,\n    num_workers=0\n)\n\nval_loader = validation.to_dataloader(\n    train=False,\n    batch_size=BATCH_SIZE,\n    num_workers=0\n)\n\n# Free split dfs if memory is tight\ngc.collect()\nif DEVICE == \"cuda\":\n    torch.cuda.empty_cache()\n\n# ============================================================\n# 8) LOSS\n# ============================================================\nclass WeightedCrossEntropy(CrossEntropy):\n    def __init__(self, class_weights: torch.Tensor, reduction: str = \"mean\"):\n        super().__init__(reduction=reduction)\n        self.register_buffer(\"class_weights\", class_weights.float())\n\n    def loss(self, y_pred: torch.Tensor, target: torch.Tensor) -> torch.Tensor:\n        B, T, C = y_pred.shape\n        ce = F.cross_entropy(\n            y_pred.reshape(B * T, C),\n            target.reshape(B * T).long(),\n            weight=self.class_weights,\n            reduction=\"none\",\n        )\n        return ce.reshape(B, T)\n\nclass_weights = torch.tensor([1.0, neg / max(1, pos)], dtype=torch.float32)\nloss = WeightedCrossEntropy(class_weights=class_weights)\n\n# ============================================================\n# 9) MODEL\n# ============================================================\nbf16_ok = bool(\n    DEVICE == \"cuda\"\n    and hasattr(torch.cuda, \"is_bf16_supported\")\n    and torch.cuda.is_bf16_supported()\n)\n\nPRECISION = \"bf16-mixed\" if bf16_ok else (\"16-mixed\" if DEVICE == \"cuda\" else 32)\nMASK_BIAS_FOR_FP16 = -float(\"inf\")\n\ntft_kwargs = dict(\n    learning_rate=3e-4,\n    hidden_size=16,\n    attention_head_size=2,\n    dropout=0.1,\n    hidden_continuous_size=8,\n    output_size=2,\n    loss=loss,\n)\n\nif PRECISION == \"16-mixed\":\n    tft_kwargs[\"mask_bias\"] = MASK_BIAS_FOR_FP16\n\ntft = TemporalFusionTransformer.from_dataset(\n    training,\n    **tft_kwargs\n)\n\n# ============================================================\n# 10) TRAINER\n# ============================================================\nckpt = ModelCheckpoint(\n    dirpath=OUT_DIR,\n    filename=\"tft_best\",\n    monitor=\"val_loss\",\n    mode=\"min\",\n    save_top_k=1\n)\n\ntrainer = pl.Trainer(\n    max_epochs=EPOCHS,\n    accelerator=\"gpu\" if DEVICE == \"cuda\" else \"cpu\",\n    devices=1,\n    precision=PRECISION,\n    gradient_clip_val=1.0,\n    accumulate_grad_batches=ACCUM_STEPS,\n    callbacks=[EarlyStopping(\"val_loss\", patience=PATIENCE, mode=\"min\"), ckpt],\n    enable_progress_bar=True,\n    log_every_n_steps=50,\n    num_sanity_val_steps=0\n)\n\n\n# ============================================================\n# 11) VALIDATION PROBABILITIES + METRICS\n# ============================================================\nraw = tft.predict(val_loader, mode=\"raw\", return_x=True)\n\nif hasattr(raw, \"output\"):\n    raw_out = raw.output\n    raw_x = raw.x\nelif isinstance(raw, tuple) and len(raw) >= 2:\n    raw_out, raw_x = raw[0], raw[1]\nelse:\n    raw_out = raw\n    raw_x = None\n\n# robust extraction of prediction tensor\nif isinstance(raw_out, dict) and \"prediction\" in raw_out:\n    pred_tensor = raw_out[\"prediction\"]\nelif hasattr(raw_out, \"prediction\"):\n    pred_tensor = raw_out.prediction\nelse:\n    pred_tensor = raw_out\n\nif isinstance(pred_tensor, (list, tuple)):\n    pred_tensor = pred_tensor[0]\n\npred_tensor = pred_tensor.detach()\n\n# logits -> positive class probability\nif pred_tensor.ndim == 3 and pred_tensor.shape[-1] == 2:\n    val_prob = torch.softmax(pred_tensor, dim=-1)[..., 1].reshape(-1).cpu().numpy()\nelif pred_tensor.ndim == 2 and pred_tensor.shape[-1] == 2:\n    val_prob = torch.softmax(pred_tensor, dim=-1)[..., 1].reshape(-1).cpu().numpy()\nelse:\n    raise RuntimeError(f\"Unexpected TFT prediction shape for classification: {tuple(pred_tensor.shape)}\")\n\n# align validation targets using returned index\nif raw_x is None:\n    raise RuntimeError(\" Could not recover raw.x from TFT predictions to align validation labels.\")\n\nval_index = validation.x_to_index(raw_x)\nval_group_ids = val_index[\"group_id\"].to_numpy()\n\ngroup_target_map = (\n    valid_df[[ \"group_id\", TARGET ]]\n    .drop_duplicates(\"group_id\")\n    .set_index(\"group_id\")[TARGET]\n)\n\ny_val_bin = group_target_map.loc[val_group_ids].to_numpy().astype(int)\n\n# evaluate\nth, f1_val = best_threshold_by_f1(y_val_bin, val_prob)\n\neval_classification(y_val_bin, val_prob, threshold=th, name=\"TFT\")\nplot_roc_pr(y_val_bin, val_prob, \"TFT\")\n\n# structured metrics\ny_hat = (val_prob >= th).astype(np.uint8)\ntn, fp, fn, tp = confusion_matrix(y_val_bin, y_hat).ravel()\n\ntft_metrics = {\n    \"name\": \"TFT\",\n    \"amex\": float(amex_metric(y_val_bin, val_prob)),\n    \"auc\": float(roc_auc_score(y_val_bin, val_prob)),\n    \"pr_auc\": float(average_precision_score(y_val_bin, val_prob)),\n    \"brier\": float(brier_score_loss(y_val_bin, val_prob)),\n    \"accuracy\": float(accuracy_score(y_val_bin, y_hat)),\n    \"precision\": float(precision_score(y_val_bin, y_hat, zero_division=0)),\n    \"recall\": float(recall_score(y_val_bin, y_hat, zero_division=0)),\n    \"f1\": float(f1_score(y_val_bin, y_hat, zero_division=0)),\n    \"tn\": int(tn),\n    \"fp\": int(fp),\n    \"fn\": int(fn),\n    \"tp\": int(tp),\n    \"threshold\": float(th),\n}\n\nprint(\"\\n[TFT Metrics]\")\nprint({\n    \"amex\": round(tft_metrics[\"amex\"], 5),\n    \"auc\": round(tft_metrics[\"auc\"], 5),\n    \"pr_auc\": round(tft_metrics[\"pr_auc\"], 5),\n    \"brier\": round(tft_metrics[\"brier\"], 5)\n})\n\n# save validation artifacts\nnp.save(OUT_DIR / \"tft_val_prob.npy\", val_prob.astype(np.float32))\nnp.save(OUT_DIR / \"tft_val_true.npy\", y_val_bin.astype(np.int8))\n\n# ============================================================\n# 12) INTERPRETABILITY\n# ============================================================\ninterp = tft.interpret_output(raw_out, reduction=\"none\")\nprint(\"[INTERP] keys:\", list(interp.keys()))\n\n# actual TFT variable names if available\nenc_var_names = getattr(tft, \"encoder_variables\", None)\ndec_var_names = getattr(tft, \"decoder_variables\", None)\nstatic_var_names = getattr(tft, \"static_variables\", None)\n\n# ------------------------------------------------------------\n# 12a) Encoder variable importance\n# ------------------------------------------------------------\nenc_global = None\nenc_names = None\n\nif \"encoder_variables\" in interp:\n    enc_imp = interp[\"encoder_variables\"].detach().cpu().numpy()\n\n    if enc_imp.ndim == 3:\n        enc_global = enc_imp.mean(axis=(0, 1))\n    elif enc_imp.ndim == 2:\n        enc_global = enc_imp.mean(axis=0)\n    else:\n        enc_global = enc_imp\n\n    if enc_var_names is not None and len(enc_var_names) == len(enc_global):\n        enc_names = list(enc_var_names)\n    else:\n        # fallback only if actual model names unavailable\n        enc_names = [f\"encoder_var_{i}\" for i in range(len(enc_global))]\n\n    topk = min(15, len(enc_global))\n    top_idx = np.argsort(enc_global)[::-1][:topk]\n\n    plt.figure(figsize=(10, 5))\n    plt.bar(range(topk), enc_global[top_idx])\n    plt.xticks(range(topk), [enc_names[i] for i in top_idx], rotation=45, ha=\"right\")\n    plt.ylabel(\"Normalized importance\")\n    plt.title(\"TFT: Top Encoder Variable Importance\")\n    plt.tight_layout()\n    plt.show()\n\n# ------------------------------------------------------------\n# 12b) Grouped importance from encoder names\n# ------------------------------------------------------------\nif enc_global is not None and enc_names is not None:\n    # group only if names contain prefixes like D_, B_, P_, S_\n    has_prefix = any(\"_\" in n for n in enc_names)\n    if has_prefix:\n        groups = [n.split(\"_\")[0] if \"_\" in n else \"OTHER\" for n in enc_names]\n\n        group_imp = {}\n        for g in sorted(set(groups)):\n            group_imp[g] = sum(\n                enc_global[i] for i in range(len(enc_global)) if groups[i] == g\n            )\n\n        total = sum(group_imp.values())\n        group_imp = {k: v / total for k, v in group_imp.items()}\n\n        print(\"\\n[TFT Grouped Feature Importance]\")\n        for k, v in sorted(group_imp.items(), key=lambda x: -x[1]):\n            print(f\"{k}: {round(v, 4)}\")\n\n        plt.figure(figsize=(6, 4))\n        plt.bar(group_imp.keys(), group_imp.values())\n        plt.title(\"TFT: Grouped Feature Importance\")\n        plt.ylabel(\"Relative importance\")\n        plt.tight_layout()\n        plt.show()\n\n        pd.DataFrame({\n            \"Group\": list(group_imp.keys()),\n            \"Importance\": list(group_imp.values())\n        }).sort_values(\"Importance\", ascending=False).to_csv(\n            OUT_DIR / \"tft_grouped_importance.csv\", index=False\n        )\n    else:\n        print(\"\\n[TFT Grouped Feature Importance] skipped because encoder names do not contain prefixes.\")\n\n# ------------------------------------------------------------\n# 12c) Temporal attention\n# ------------------------------------------------------------\nif \"attention\" in interp:\n    attn = interp[\"attention\"].detach().cpu().numpy()\n\n    if attn.ndim == 4:\n        attn_1d = attn.mean(axis=(0, 1, 2))\n    elif attn.ndim == 3:\n        attn_1d = attn.mean(axis=(0, 1))\n    elif attn.ndim == 2:\n        attn_1d = attn.mean(axis=0)\n    else:\n        attn_1d = attn.reshape(-1)\n\n    plt.figure(figsize=(10, 4))\n    plt.plot(attn_1d, marker=\"o\")\n    plt.title(\"TFT: Average Temporal Attention\")\n    plt.xlabel(\"Encoder time step\")\n    plt.ylabel(\"Attention mass\")\n    plt.grid(alpha=0.3)\n    plt.tight_layout()\n    plt.show()\n\n    p = np.abs(attn_1d) / (np.sum(np.abs(attn_1d)) + 1e-8)\n    entropy = -np.sum(p * np.log(p + 1e-8))\n    top3 = float(np.sort(p)[-3:].sum())\n    top5 = float(np.sort(p)[-5:].sum())\n\n    cum = np.cumsum(np.sort(p)[::-1])\n    k80 = int(np.argmax(cum >= 0.80) + 1)\n    k90 = int(np.argmax(cum >= 0.90) + 1)\n\n    print(\"\\n[TFT Temporal Metrics]\")\n    print(\"Entropy:\", round(entropy, 4))\n    print(\"Top-3 mass:\", round(top3, 4))\n    print(\"Top-5 mass:\", round(top5, 4))\n    print(\"k80:\", k80, \"| k90:\", k90)\n\n    np.save(OUT_DIR / \"tft_attention_time_importance.npy\", attn_1d.astype(np.float32))\n    np.save(OUT_DIR / \"tft_attention_entropy.npy\", np.array([entropy], dtype=np.float32))\n    np.save(OUT_DIR / \"tft_attention_topk_mass.npy\", np.array([top3, top5], dtype=np.float32))\n    np.save(OUT_DIR / \"tft_attention_k_coverage.npy\", np.array([k80, k90], dtype=np.int32))\n\n# ------------------------------------------------------------\n# 12d) Static variable importance\n# ------------------------------------------------------------\nif \"static_variables\" in interp:\n    s = interp[\"static_variables\"].detach().cpu().numpy()\n\n    if s.ndim > 1:\n        s = s.mean(axis=0)\n\n    if static_var_names is not None and len(static_var_names) == len(s):\n        names = list(static_var_names)\n    else:\n        names = [f\"static_var_{i}\" for i in range(len(s))]\n\n    plt.figure(figsize=(8, 4))\n    plt.bar(range(len(s)), s)\n    plt.xticks(range(len(s)), names, rotation=45, ha=\"right\")\n    plt.title(\"TFT: Static Variable Importance\")\n    plt.ylabel(\"Importance\")\n    plt.tight_layout()\n    plt.show()\n\n    np.save(OUT_DIR / \"tft_static_variable_importance.npy\", s.astype(np.float32))\n\n# save summary artifacts\nif enc_global is not None:\n    np.save(OUT_DIR / \"tft_encoder_variable_importance.npy\", enc_global.astype(np.float32))\n    pd.DataFrame({\n        \"feature\": enc_names,\n        \"importance\": enc_global\n    }).sort_values(\"importance\", ascending=False).to_csv(\n        OUT_DIR / \"tft_feature_importance_table.csv\", index=False\n    )\n\nprint(\"\\n TFT pipeline complete\")\nprint(\"Artifacts saved to:\", OUT_DIR)","metadata":{"_uuid":"a33eaa17-cf7e-43ad-92da-ea4e6253ddc9","_cell_guid":"0b4d1715-ffa0-4c62-8eae-6f3a84864ee3","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-06-05T18:25:39.473874Z","iopub.execute_input":"2026-06-05T18:25:39.475047Z","execution_failed":"2026-06-05T18:26:24.622Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# B) Informer Pipeline– Multi-Horizon Default Forecasting\n\n## Objective\nTrain an Informer-based encoder–decoder model to predict **multi-horizon default hazard** over the next `PRED_LEN` months and evaluate using **AMEX metric (any-default)**.\n\n## Input Requirements\n- `X_train_seq`: `[N, T, F]` sequence features (memmap supported)\n- Labels (one required):\n  - `default_month_idx`: `[N]` (first default month, -1 if none)\n  - `y_event_seq`: `[N, T]` (monthly event indicator)\n- Predefined:\n  - `SEQ_FEATURES`, `SEED`\n  - Evaluation utils: `amex_metric`, `best_threshold_by_f1`, etc.\n\n## Dataset Construction\n- Encoder input: `x_enc ∈ [ENC_LEN, F]`\n- Decoder input: `x_dec ∈ [LABEL_LEN + PRED_LEN, F]`\n  - Contains historical tail + zero-padded future\n- Target: `y_h ∈ [PRED_LEN]` (future default events)\n\n## Model Architecture (Informer Full)\n- **Encoder**\n  - ProbSparse Attention (Top‑U query selection)\n  - Distillation (Conv + pooling → sequence compression)\n- **Decoder**\n  - Causal self-attention\n  - Cross-attention to encoder outputs\n- **Embedding**\n  - Value + positional encoding\n- **Output**\n  - Horizon logits → sigmoid → hazard probabilities\n\n## Training Setup\n- Loss: `BCEWithLogitsLoss` (class imbalance via `pos_weight`)\n- Optimizer: `AdamW`\n- Techniques:\n  - AMP + GradScaler\n  - Gradient clipping (1.0)\n  - Early stopping (patience = 3)\n- Metric:\n  - Convert hazard → any-default probability:\n    ```\n    P(default) = 1 − ∏(1 − hazard_t)\n    ```\n  - Validate using AMEX metric\n\n## Outputs Saved\n- Validation:\n  - `val_hazard`, `val_anyprob`, `val_anytrue`\n  - Predicted default month (argmax hazard)\n- Test:\n  - `test_hazard`, `test_anyprob`\n- Stored in `/kaggle/working/seq_artifacts`\n\n## Informer-Native Interpretability\nExtracted from ProbSparse attention:\n\n- **Active Queries**\n  - Selected high-impact time steps\n- **Active Ratio**\n  - `u / L_Q` → sparsity level\n- **Distillation Trace**\n  - Sequence length reduction across encoder layers\n- **Cross-Attention Mass**\n  - Importance of encoder positions for decoder\n\nSaved:\n- Encoder active ratios\n- Decoder cross-attention ratios\n- Distillation lengths\n\n## Attribution Interpretability (Grad×Input)\nComputed on encoder inputs using any-default probability:\n\n### Feature Importance\n- Aggregated |grad × input| across time & batch\n- Top-K features identified\n\n### Temporal Importance\n- Importance per time step\n\n### Metrics\n- **Entropy** → dispersion of importance\n- **Top-k mass** → concentration (Top-3, Top-5)\n- **Coverage**\n  - Steps required to explain 80% / 90% importance\n\n### Long-Range Dependency Proxy\n- Compare early vs late sequence importance:\n  - Detect reliance on long vs recent history\n\nSaved:\n- Feature importance\n- Time importance\n- Entropy + Top-K metrics\n- Coverage (k80, k90)\n\n## Key Advantages\n- Efficient long-sequence modeling via ProbSparse attention\n- Explicit multi-horizon default timing (not just binary outcome)\n- Built-in interpretability:\n  - Structural (attention patterns)\n  - Attribution (gradient-based)","metadata":{"_uuid":"a3399ec9-a1a0-45b3-9f54-0bf6cd285ebb","_cell_guid":"5a9d4c71-351e-4da3-919c-2c653bf1054c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# INFORMER PIPELINE\n# ============================================================\n\nimport os\nimport gc\nimport math\nimport re\nimport numpy as np\nimport torch\nimport torch.nn as nn\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\nfrom sklearn.model_selection import train_test_split\nfrom torch.utils.data import Dataset, DataLoader, Subset\n\n# ============================================================\n# 0) PRECONDITIONS\n# ============================================================\nrequired = [\n    \"SEQ_FEATURES\",\n    \"SEED\",\n    \"amex_metric\",\n    \"best_threshold_by_f1\",\n    \"eval_classification\",\n    \"plot_roc_pr\",\n]\nmissing = [k for k in required if k not in globals()]\nif missing:\n    raise NameError(f\"Missing required objects: {missing}\")\n\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\nuse_amp = (DEVICE == \"cuda\")\nprint(\"DEVICE:\", DEVICE)\n\ntorch.manual_seed(SEED)\nnp.random.seed(SEED)\nif DEVICE == \"cuda\":\n    torch.cuda.manual_seed_all(SEED)\n\n# ============================================================\n# 1) PATHS + DATA LOADING\n# ============================================================\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\n\nTRAIN_PATH = OUT_DIR / \"X_train_seq.dat\"\nTEST_PATH  = OUT_DIR / \"X_test_seq.dat\"\nY_PATH     = OUT_DIR / \"y_seq.npy\"\n\nif not TRAIN_PATH.exists():\n    raise FileNotFoundError(f\"Missing {TRAIN_PATH}\")\nif not TEST_PATH.exists():\n    raise FileNotFoundError(f\"Missing {TEST_PATH}\")\nif not Y_PATH.exists():\n    raise FileNotFoundError(f\"Missing {Y_PATH}\")\n\n# sequence dimensions\nSEQ_LEN = 13\nN_FEATURES = len(SEQ_FEATURES)\nDTYPE = np.float32\n\ndef infer_num_rows(path, seq_len, n_features, dtype=np.float32):\n    total_bytes = os.path.getsize(path)\n    row_bytes = seq_len * n_features * np.dtype(dtype).itemsize\n    if total_bytes % row_bytes != 0:\n        raise RuntimeError(\n            f\"File size of {path} is not divisible by row size. \"\n            f\"Check seq_len={seq_len} and n_features={n_features}.\"\n        )\n    return total_bytes // row_bytes\n\nN_TRAIN = infer_num_rows(TRAIN_PATH, SEQ_LEN, N_FEATURES, DTYPE)\nN_TEST  = infer_num_rows(TEST_PATH,  SEQ_LEN, N_FEATURES, DTYPE)\n\nX_train_seq = np.memmap(TRAIN_PATH, mode=\"r\", dtype=DTYPE, shape=(N_TRAIN, SEQ_LEN, N_FEATURES))\nX_test_seq  = np.memmap(TEST_PATH,  mode=\"r\", dtype=DTYPE, shape=(N_TEST,  SEQ_LEN, N_FEATURES))\ny_seq = np.load(Y_PATH).astype(np.int64)\n\nif len(y_seq) != N_TRAIN:\n    raise RuntimeError(f\"y_seq length mismatch: {len(y_seq)} vs train rows {N_TRAIN}\")\n\nprint(\"X_train_seq:\", X_train_seq.shape)\nprint(\"X_test_seq :\", X_test_seq.shape)\nprint(\"y_seq      :\", y_seq.shape)\n\n# ============================================================\n# 2) STRICT FEATURE NAMES\n# ============================================================\nFEATURE_NAMES = list(SEQ_FEATURES)\n\nif len(FEATURE_NAMES) != N_FEATURES:\n    raise RuntimeError(\"Feature mismatch\")\n\nif all(bool(re.fullmatch(r\"feature_\\d+\", f)) for f in FEATURE_NAMES):\n    raise RuntimeError(\"Generic feature_* names detected. Use real feature names.\")\n\n# ============================================================\n# 3) SPLIT\n# ============================================================\nif all(k in globals() for k in [\"idx_tr\", \"idx_va\"]):\n    print(\"Using existing idx_tr / idx_va\")\n    idx_tr = np.asarray(idx_tr, dtype=np.int32)\n    idx_va = np.asarray(idx_va, dtype=np.int32)\nelse:\n    idx = np.arange(N_TRAIN, dtype=np.int32)\n    idx_tr, idx_va = train_test_split(\n        idx,\n        test_size=0.2,\n        random_state=SEED,\n        stratify=y_seq\n    )\n\ny_tr = y_seq[idx_tr]\ny_va = y_seq[idx_va]\n\nprint(\"Train idx:\", len(idx_tr), \"| Valid idx:\", len(idx_va))\n\n# ============================================================\n# 4) DATASET\n# ============================================================\nLABEL_LEN = 4  # decoder history length\n\nclass MemmapSeqDataset(Dataset):\n    def __init__(self, X_mmap, y=None):\n        self.X = X_mmap\n        self.y = None if y is None else np.asarray(y, dtype=np.uint8)\n\n    def __len__(self):\n        return self.X.shape[0]\n\n    def __getitem__(self, idx):\n        x = torch.from_numpy(self.X[idx])  # [T, F]\n        if self.y is None:\n            return x\n        return x, torch.tensor(self.y[idx], dtype=torch.float32)\n\ndef collate_fn(batch):\n    \"\"\"\n    Returns:\n      x_enc: [B, T, F]\n      x_dec: [B, LABEL_LEN+1, F]  -> last LABEL_LEN history + one zero token\n      y:     [B]\n    \"\"\"\n    if isinstance(batch[0], (tuple, list)):\n        xs, ys = zip(*batch)\n        x_enc = torch.stack(xs, dim=0).float()\n        y = torch.stack([torch.as_tensor(v) for v in ys]).float()\n    else:\n        x_enc = torch.stack(batch, dim=0).float()\n        y = None\n\n    hist_tail = x_enc[:, -LABEL_LEN:, :]                    # [B, LABEL_LEN, F]\n    zero_token = torch.zeros(x_enc.size(0), 1, x_enc.size(2))\n    x_dec = torch.cat([hist_tail, zero_token], dim=1)      # [B, LABEL_LEN+1, F]\n\n    if y is None:\n        return x_enc, x_dec\n    return x_enc, x_dec, y\n\ntrain_ds = Subset(MemmapSeqDataset(X_train_seq, y_seq), idx_tr)\nvalid_ds = Subset(MemmapSeqDataset(X_train_seq, y_seq), idx_va)\ntest_ds  = MemmapSeqDataset(X_test_seq, y=None)\n\nBATCH_SIZE = 128 if DEVICE == \"cuda\" else 64\n\ntrain_loader = DataLoader(\n    train_ds,\n    batch_size=BATCH_SIZE,\n    shuffle=True,\n    num_workers=0,\n    pin_memory=(DEVICE == \"cuda\"),\n    collate_fn=collate_fn\n)\n\nvalid_loader = DataLoader(\n    valid_ds,\n    batch_size=BATCH_SIZE,\n    shuffle=False,\n    num_workers=0,\n    pin_memory=(DEVICE == \"cuda\"),\n    collate_fn=collate_fn\n)\n\ntest_loader = DataLoader(\n    test_ds,\n    batch_size=BATCH_SIZE,\n    shuffle=False,\n    num_workers=0,\n    pin_memory=(DEVICE == \"cuda\"),\n    collate_fn=collate_fn\n)\n\n# ============================================================\n# 5) MODEL BUILDING BLOCKS\n# ============================================================\nclass PositionalEncoding(nn.Module):\n    def __init__(self, d_model, max_len=512):\n        super().__init__()\n        pe = torch.zeros(max_len, d_model, dtype=torch.float32)\n        position = torch.arange(0, max_len, dtype=torch.float32).unsqueeze(1)\n        div_term = torch.exp(\n            torch.arange(0, d_model, 2, dtype=torch.float32) * (-math.log(10000.0) / d_model)\n        )\n        pe[:, 0::2] = torch.sin(position * div_term)\n        pe[:, 1::2] = torch.cos(position * div_term)\n        self.register_buffer(\"pe\", pe.unsqueeze(0), persistent=False)\n\n    def forward(self, x):\n        return x + self.pe[:, :x.size(1), :]\n\ndef triangular_causal_mask(B, L, device):\n    mask = torch.triu(torch.ones((L, L), device=device, dtype=torch.bool), diagonal=1)\n    return mask.unsqueeze(0).unsqueeze(1).expand(B, 1, L, L)\n\nclass ProbSparseAttention(nn.Module):\n    \"\"\"\n    ProbSparse attention + built-in interpretability:\n      - active query indices\n      - active ratio\n      - sparse attention weights for selected queries (optional)\n    \"\"\"\n    def __init__(self, factor=5, dropout=0.1, output_attention=False):\n        super().__init__()\n        self.factor = factor\n        self.dropout = nn.Dropout(dropout)\n        self.output_attention = output_attention\n\n    def forward(self, Q, K, V, attn_mask=None):\n        # Q,K,V: [B,H,L,D]\n        B, H, L_Q, D = Q.shape\n        _, _, L_K, _ = K.shape\n\n        U_part = min(L_K, max(1, int(self.factor * math.log(L_K + 1))))\n        u = min(L_Q, max(1, int(self.factor * math.log(L_Q + 1))))\n\n        index_sample = torch.randint(L_K, (L_Q, U_part), device=Q.device)\n        K_sample = K[:, :, index_sample, :]  # [B,H,L_Q,U_part,D]\n\n        Q_expand = Q.unsqueeze(-2)\n        QK_sample = torch.matmul(Q_expand, K_sample.transpose(-2, -1)).squeeze(-2)  # [B,H,L_Q,U_part]\n\n        M = QK_sample.max(dim=-1).values - QK_sample.mean(dim=-1)\n        M_top = M.topk(u, dim=-1, sorted=False).indices  # [B,H,u]\n\n        Q_reduce = Q.gather(2, M_top.unsqueeze(-1).expand(-1, -1, -1, D))  # [B,H,u,D]\n        scores = torch.matmul(Q_reduce, K.transpose(-2, -1)) / math.sqrt(D) # [B,H,u,L_K]\n\n        if attn_mask is not None:\n            mask_sel = attn_mask.expand(B, H, L_Q, L_K).gather(\n                2, M_top.unsqueeze(-1).expand(-1, -1, -1, L_K)\n            )\n            scores = scores.masked_fill(mask_sel, float(\"-inf\"))\n\n        attn = torch.softmax(scores, dim=-1)\n        attn = self.dropout(attn)\n        out = torch.matmul(attn, V)  # [B,H,u,D]\n\n        context = V.mean(dim=-2, keepdim=True).expand(B, H, L_Q, D).contiguous()\n        context.scatter_(2, M_top.unsqueeze(-1).expand(-1, -1, -1, D), out)\n\n        info = {\n            \"active_ratio\": float(u) / float(max(1, L_Q)),\n            \"active_queries\": M_top.detach(),\n            \"attn_sparse\": attn.detach() if self.output_attention else None\n        }\n        return context, info\n\nclass AttentionLayer(nn.Module):\n    def __init__(self, d_model, n_heads, factor=5, dropout=0.1, output_attention=False):\n        super().__init__()\n        assert d_model % n_heads == 0\n        self.n_heads = n_heads\n        self.d_head = d_model // n_heads\n\n        self.q_proj = nn.Linear(d_model, d_model)\n        self.k_proj = nn.Linear(d_model, d_model)\n        self.v_proj = nn.Linear(d_model, d_model)\n        self.out_proj = nn.Linear(d_model, d_model)\n\n        self.inner_attention = ProbSparseAttention(\n            factor=factor,\n            dropout=dropout,\n            output_attention=output_attention\n        )\n\n    def forward(self, x_q, x_kv, attn_mask=None):\n        B, Lq, D = x_q.shape\n        _, Lk, _ = x_kv.shape\n\n        Q = self.q_proj(x_q).view(B, Lq, self.n_heads, self.d_head).transpose(1, 2)\n        K = self.k_proj(x_kv).view(B, Lk, self.n_heads, self.d_head).transpose(1, 2)\n        V = self.v_proj(x_kv).view(B, Lk, self.n_heads, self.d_head).transpose(1, 2)\n\n        context, info = self.inner_attention(Q, K, V, attn_mask=attn_mask)\n        context = context.transpose(1, 2).contiguous().view(B, Lq, D)\n        return self.out_proj(context), info\n\nclass EncoderLayer(nn.Module):\n    def __init__(self, d_model, n_heads, d_ff=256, dropout=0.1, factor=5, output_attention=False):\n        super().__init__()\n        self.attn = AttentionLayer(d_model, n_heads, factor=factor, dropout=dropout, output_attention=output_attention)\n        self.norm1 = nn.LayerNorm(d_model)\n        self.ff = nn.Sequential(\n            nn.Linear(d_model, d_ff),\n            nn.GELU(),\n            nn.Dropout(dropout),\n            nn.Linear(d_ff, d_model),\n            nn.Dropout(dropout),\n        )\n        self.norm2 = nn.LayerNorm(d_model)\n\n    def forward(self, x):\n        a, info = self.attn(x, x, attn_mask=None)\n        x = self.norm1(x + a)\n        x = self.norm2(x + self.ff(x))\n        return x, info\n\nclass ConvDistill(nn.Module):\n    \"\"\"\n    Distillation step from Informer:\n    conv + activation + maxpool halves sequence length\n    \"\"\"\n    def __init__(self, d_model):\n        super().__init__()\n        self.conv = nn.Conv1d(d_model, d_model, kernel_size=3, padding=1)\n        self.act = nn.ELU()\n        self.pool = nn.MaxPool1d(kernel_size=2, stride=2)\n\n    def forward(self, x):\n        # x: [B,L,D]\n        x = x.transpose(1, 2)\n        x = self.pool(self.act(self.conv(x)))\n        return x.transpose(1, 2)\n\nclass Encoder(nn.Module):\n    def __init__(self, d_model, n_heads, n_layers=3, d_ff=256, dropout=0.1, factor=5, distil=True, output_attention=False):\n        super().__init__()\n        self.layers = nn.ModuleList([\n            EncoderLayer(d_model, n_heads, d_ff=d_ff, dropout=dropout, factor=factor, output_attention=output_attention)\n            for _ in range(n_layers)\n        ])\n        self.distil = distil and n_layers > 1\n        self.distill_layers = nn.ModuleList([ConvDistill(d_model) for _ in range(n_layers - 1)]) if self.distil else None\n\n    def forward(self, x):\n        attn_infos = []\n        lengths = [x.size(1)]\n\n        for i, layer in enumerate(self.layers):\n            x, info = layer(x)\n            attn_infos.append(info)\n\n            if self.distil and i < len(self.distill_layers):\n                x = self.distill_layersx\n                lengths.append(x.size(1))\n\n        return x, attn_infos, lengths\n\nclass DecoderLayer(nn.Module):\n    def __init__(self, d_model, n_heads, d_ff=256, dropout=0.1, factor=5, output_attention=False):\n        super().__init__()\n        self.self_attn = AttentionLayer(d_model, n_heads, factor=factor, dropout=dropout, output_attention=output_attention)\n        self.cross_attn = AttentionLayer(d_model, n_heads, factor=factor, dropout=dropout, output_attention=output_attention)\n\n        self.norm1 = nn.LayerNorm(d_model)\n        self.norm2 = nn.LayerNorm(d_model)\n        self.ff = nn.Sequential(\n            nn.Linear(d_model, d_ff),\n            nn.GELU(),\n            nn.Dropout(dropout),\n            nn.Linear(d_ff, d_model),\n            nn.Dropout(dropout),\n        )\n        self.norm3 = nn.LayerNorm(d_model)\n\n    def forward(self, x, enc_out, self_mask):\n        a, self_info = self.self_attn(x, x, attn_mask=self_mask)\n        x = self.norm1(x + a)\n\n        c, cross_info = self.cross_attn(x, enc_out, attn_mask=None)\n        x = self.norm2(x + c)\n\n        x = self.norm3(x + self.ff(x))\n        return x, self_info, cross_info\n\nclass Decoder(nn.Module):\n    def __init__(self, d_model, n_heads, n_layers=2, d_ff=256, dropout=0.1, factor=5, output_attention=False):\n        super().__init__()\n        self.layers = nn.ModuleList([\n            DecoderLayer(d_model, n_heads, d_ff=d_ff, dropout=dropout, factor=factor, output_attention=output_attention)\n            for _ in range(n_layers)\n        ])\n\n    def forward(self, x, enc_out):\n        B, Ld, _ = x.shape\n        self_mask = triangular_causal_mask(B, Ld, x.device)\n\n        self_infos = []\n        cross_infos = []\n\n        for layer in self.layers:\n            x, si, ci = layer(x, enc_out, self_mask)\n            self_infos.append(si)\n            cross_infos.append(ci)\n\n        return x, self_infos, cross_infos\n\n# ============================================================\n# 6) INFORMER CLASSIFIER\n# ============================================================\nclass InformerClassifier(nn.Module):\n    \"\"\"\n    Informer encoder-decoder classifier:\n    - ProbSparse attention\n    - distillation\n    - decoder\n    - binary classification head\n    \"\"\"\n    def __init__(\n        self,\n        n_features,\n        d_model=128,\n        n_heads=4,\n        enc_layers=3,\n        dec_layers=2,\n        d_ff=256,\n        dropout=0.1,\n        factor=5,\n        distil=True,\n        output_attention=False\n    ):\n        super().__init__()\n\n        self.enc_input = nn.Linear(n_features, d_model)\n        self.dec_input = nn.Linear(n_features, d_model)\n\n        self.enc_pos = PositionalEncoding(d_model, max_len=SEQ_LEN + 16)\n        self.dec_pos = PositionalEncoding(d_model, max_len=LABEL_LEN + 8)\n\n        self.encoder = Encoder(\n            d_model=d_model,\n            n_heads=n_heads,\n            n_layers=enc_layers,\n            d_ff=d_ff,\n            dropout=dropout,\n            factor=factor,\n            distil=distil,\n            output_attention=output_attention\n        )\n\n        self.decoder = Decoder(\n            d_model=d_model,\n            n_heads=n_heads,\n            n_layers=dec_layers,\n            d_ff=d_ff,\n            dropout=dropout,\n            factor=factor,\n            output_attention=output_attention\n        )\n\n        self.dropout = nn.Dropout(dropout)\n        self.head = nn.Linear(d_model, 1)\n\n    def forward(self, x_enc, x_dec, return_info=False):\n        enc = self.enc_pos(self.enc_input(x_enc))\n        dec = self.dec_pos(self.dec_input(x_dec))\n\n        enc_out, enc_infos, enc_lengths = self.encoder(enc)\n        dec_out, dec_self_infos, dec_cross_infos = self.decoder(dec, enc_out)\n\n        h = self.dropout(dec_out[:, -1, :])\n        logits = self.head(h).squeeze(-1)\n\n        if not return_info:\n            return logits\n\n        info = {\n            \"encoder\": enc_infos,\n            \"encoder_lengths\": enc_lengths,\n            \"decoder_self\": dec_self_infos,\n            \"decoder_cross\": dec_cross_infos\n        }\n        return logits, info\n\n# ============================================================\n# 7) TRAINING UTILITIES\n# ============================================================\ndef compute_pos_weight(y_binary):\n    pos = float((y_binary == 1).sum())\n    neg = float((y_binary == 0).sum())\n    return torch.tensor([neg / max(1.0, pos)], dtype=torch.float32, device=DEVICE)\n\n@torch.no_grad()\ndef predict_proba(model, loader):\n    model.eval()\n    preds = []\n    for batch in loader:\n        if len(batch) == 3:\n            x_enc, x_dec, _ = batch\n        else:\n            x_enc, x_dec = batch\n\n        x_enc = x_enc.to(DEVICE, non_blocking=True)\n        x_dec = x_dec.to(DEVICE, non_blocking=True)\n\n        with torch.cuda.amp.autocast(enabled=use_amp):\n            logits = model(x_enc, x_dec)\n            p = torch.sigmoid(logits)\n\n        preds.append(p.detach().cpu().numpy())\n    return np.concatenate(preds)\n\n# ============================================================\n# 8) BUILD MODEL\n# ============================================================\nmodel = InformerClassifier(\n    n_features=N_FEATURES,\n    d_model=128,\n    n_heads=4,\n    enc_layers=3,\n    dec_layers=2,\n    d_ff=256,\n    dropout=0.1,\n    factor=5,\n    distil=True,\n    output_attention=False\n).to(DEVICE)\n\npos_weight = compute_pos_weight(y_tr)\ncriterion = nn.BCEWithLogitsLoss(pos_weight=pos_weight)\noptimizer = torch.optim.AdamW(model.parameters(), lr=2e-4, weight_decay=1e-2)\nscaler = torch.cuda.amp.GradScaler(enabled=use_amp)\n\n# ============================================================\n# 9) TRAIN LOOP\n# ============================================================\nEPOCHS = 5\nPATIENCE = 2\n\nbest_score = -1e18\nbest_state = None\npat_count = 0\n\nfor epoch in range(1, EPOCHS + 1):\n    model.train()\n    running = 0.0\n    seen = 0\n\n    for x_enc, x_dec, yb in train_loader:\n        x_enc = x_enc.to(DEVICE, non_blocking=True)\n        x_dec = x_dec.to(DEVICE, non_blocking=True)\n        yb = yb.to(DEVICE, non_blocking=True)\n\n        optimizer.zero_grad(set_to_none=True)\n\n        with torch.cuda.amp.autocast(enabled=use_amp):\n            logits = model(x_enc, x_dec)\n            loss = criterion(logits, yb)\n\n        scaler.scale(loss).backward()\n        scaler.unscale_(optimizer)\n        torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)\n        scaler.step(optimizer)\n        scaler.update()\n\n        running += loss.item() * x_enc.size(0)\n        seen += x_enc.size(0)\n\n    train_loss = running / max(1, seen)\n\n    valid_pred = predict_proba(model, valid_loader)\n    score = amex_metric(y_va, valid_pred)\n\n    print(f\"Epoch {epoch:02d} | train_loss={train_loss:.5f} | AMEX_val={score:.6f}\")\n\n    if score > best_score + 1e-6:\n        best_score = score\n        best_state = {k: v.detach().cpu().clone() for k, v in model.state_dict().items()}\n        pat_count = 0\n    else:\n        pat_count += 1\n        if pat_count >= PATIENCE:\n            print(\"Early stopping triggered.\")\n            break\n\n# restore best\nif best_state is not None:\n    model.load_state_dict(best_state)\n\n# ============================================================\n# 10) VALIDATION METRICS\n# ============================================================\ninf_valid_pred = predict_proba(model, valid_loader)\n\nth, f1_best = best_threshold_by_f1(y_va, inf_valid_pred)\ninf_metrics = eval_classification(y_va, inf_valid_pred, threshold=th, name=\"Informer-Full\")\nplot_roc_pr(y_va, inf_valid_pred, \"Informer-Full\")\n\nprint(\"\\n[Informer-Full Metrics]\")\nprint(\"AMEX:\", round(float(best_score), 5))\n\nnp.save(OUT_DIR / \"inf_full_valid_pred.npy\", inf_valid_pred.astype(np.float32))\n\n# ============================================================\n# 11) TEST PREDICTION\n# ============================================================\ninf_test_pred = predict_proba(model, test_loader)\nnp.save(OUT_DIR / \"inf_full_test_pred.npy\", inf_test_pred.astype(np.float32))\n\n# ============================================================\n# 12) BUILT-IN INFORMER INTERPRETABILITY\n# ============================================================\nprint(\"\\n[INTERP] Extracting built-in Informer signals...\")\n\ninterp_model = InformerClassifier(\n    n_features=N_FEATURES,\n    d_model=128,\n    n_heads=4,\n    enc_layers=3,\n    dec_layers=2,\n    d_ff=256,\n    dropout=0.1,\n    factor=5,\n    distil=True,\n    output_attention=True\n).to(DEVICE)\n\ninterp_model.load_state_dict(model.state_dict(), strict=True)\ninterp_model.eval()\n\nx_enc_b, x_dec_b, yb_b = next(iter(valid_loader))\nx_enc_b = x_enc_b.to(DEVICE)\nx_dec_b = x_dec_b.to(DEVICE)\n\nwith torch.no_grad():\n    logits_b, info = interp_model(x_enc_b, x_dec_b, return_info=True)\n\nprint(\"Encoder distillation lengths:\", info[\"encoder_lengths\"])\n\n# Active ratios\nenc_active = [d[\"active_ratio\"] for d in info[\"encoder\"]]\ndec_self_active = [d[\"active_ratio\"] for d in info[\"decoder_self\"]]\ndec_cross_active = [d[\"active_ratio\"] for d in info[\"decoder_cross\"]]\n\nprint(\"Encoder active ratios:\", enc_active)\nprint(\"Decoder self active ratios:\", dec_self_active)\nprint(\"Decoder cross active ratios:\", dec_cross_active)\n\nnp.save(OUT_DIR / \"inf_encoder_active_ratios.npy\", np.array(enc_active, dtype=np.float32))\nnp.save(OUT_DIR / \"inf_decoder_self_active_ratios.npy\", np.array(dec_self_active, dtype=np.float32))\nnp.save(OUT_DIR / \"inf_decoder_cross_active_ratios.npy\", np.array(dec_cross_active, dtype=np.float32))\n\n# Distillation trace\nplt.figure(figsize=(6, 3))\nplt.plot(info[\"encoder_lengths\"], marker=\"o\")\nplt.title(\"Informer Distillation Trace\")\nplt.xlabel(\"Stage\")\nplt.ylabel(\"Sequence length\")\nplt.grid(alpha=0.3)\nplt.tight_layout()\nplt.show()\n\nnp.save(OUT_DIR / \"inf_encoder_lengths.npy\", np.array(info[\"encoder_lengths\"], dtype=np.int32))\n\n# Active query histogram\naq = info[\"encoder\"][0][\"active_queries\"]\naq0 = aq[0, 0].detach().cpu().numpy()\n\nplt.figure(figsize=(6, 3))\nplt.hist(aq0, bins=min(20, len(aq0)))\nplt.title(\"Informer Active Query Indices (Encoder L0)\")\nplt.xlabel(\"Time position\")\nplt.ylabel(\"Count\")\nplt.tight_layout()\nplt.show()\n\n# Decoder cross-attention mass\ncross_sparse = info[\"decoder_cross\"][0][\"attn_sparse\"]\nif cross_sparse is not None:\n    cross_mass = cross_sparse[0].mean(dim=0).mean(dim=0).detach().cpu().numpy()\n\n    plt.figure(figsize=(7, 3))\n    plt.plot(cross_mass, marker=\"o\")\n    plt.title(\"Informer Decoder Cross-Attention Mass\")\n    plt.xlabel(\"Distilled encoder position\")\n    plt.ylabel(\"Average attention mass\")\n    plt.grid(alpha=0.3)\n    plt.tight_layout()\n    plt.show()\n\n    np.save(OUT_DIR / \"inf_decoder_cross_attention_mass.npy\", cross_mass.astype(np.float32))\n\n    p = np.abs(cross_mass)\n    p = p / (p.sum() + 1e-8)\n\n    entropy = -np.sum(p * np.log(p + 1e-8))\n    top3 = float(np.sort(p)[-3:].sum())\n    top5 = float(np.sort(p)[-5:].sum())\n\n    cum = np.cumsum(np.sort(p)[::-1])\n    k80 = int(np.argmax(cum >= 0.80) + 1)\n    k90 = int(np.argmax(cum >= 0.90) + 1)\n\n    print(\"\\n[Informer Built-in Temporal Metrics]\")\n    print(\"Entropy:\", round(float(entropy), 4))\n    print(\"Top-3 mass:\", round(top3, 4))\n    print(\"Top-5 mass:\", round(top5, 4))\n    print(\"k80:\", k80, \"| k90:\", k90)\n\n    np.save(OUT_DIR / \"inf_builtin_entropy.npy\", np.array([entropy], dtype=np.float32))\n    np.save(OUT_DIR / \"inf_builtin_topk_mass.npy\", np.array([top3, top5], dtype=np.float32))\n    np.save(OUT_DIR / \"inf_builtin_k_coverage.npy\", np.array([k80, k90], dtype=np.int32))\n\n# ============================================================\n# 13)INTERPRETABILITY\n# ============================================================\nprint(\"\\n[INTERP] Running gradient-based attribution...\")\n\ndef grad_importance(model_for_grad, x_enc, x_dec):\n    model_for_grad.eval()\n\n    x_enc = torch.tensor(x_enc, dtype=torch.float32).to(DEVICE)\n    x_dec = torch.tensor(x_dec, dtype=torch.float32).to(DEVICE)\n    x_enc.requires_grad_()\n\n    logits = model_for_grad(x_enc, x_dec)\n    out = torch.sigmoid(logits).mean()\n    out.backward()\n\n    grad = (x_enc.grad.abs() * x_enc.abs()).mean(dim=0).detach().cpu().numpy()\n    return grad\n\nsample_idx = idx_va[:min(64, len(idx_va))]\nx_sample = X_train_seq[sample_idx]\nx_sample_t = torch.tensor(x_sample, dtype=torch.float32)\nhist_tail = x_sample_t[:, -LABEL_LEN:, :]\nzero_token = torch.zeros(x_sample_t.size(0), 1, x_sample_t.size(2))\nx_dec_sample = torch.cat([hist_tail, zero_token], dim=1)\n\nimp = grad_importance(model, x_sample, x_dec_sample.numpy())\nimp = imp / (np.max(np.abs(imp)) + 1e-8)\n\nfeat_imp = imp.mean(axis=0)\ntime_imp = imp.mean(axis=1)\n\n# Feature importance\ntopk = min(15, len(feat_imp))\ntop_idx = np.argsort(-feat_imp)[:topk]\n\nplt.figure(figsize=(10, 5))\nplt.bar(range(topk), feat_imp[top_idx])\nplt.xticks(range(topk), [FEATURE_NAMES[i] for i in top_idx], rotation=45, ha=\"right\")\nplt.title(\"Informer Feature Importance (Grad×Input)\")\nplt.tight_layout()\nplt.show()\n\n# Grouped feature importance\ngroups = [f.split(\"_\")[0] if \"_\" in f else \"OTHER\" for f in FEATURE_NAMES]\ngroup_imp = {}\nfor g in set(groups):\n    group_imp[g] = sum(\n        feat_imp[i] for i in range(len(feat_imp))\n        if groups[i] == g\n    )\n\ntotal = sum(group_imp.values()) + 1e-8\ngroup_imp = {k: v / total for k, v in group_imp.items()}\n\nprint(\"\\n[Informer Grouped Feature Importance]\")\nfor k, v in sorted(group_imp.items(), key=lambda x: -x[1]):\n    print(f\"{k}: {round(v, 4)}\")\n\nplt.figure(figsize=(6, 4))\nplt.bar(group_imp.keys(), group_imp.values())\nplt.title(\"Informer Grouped Feature Importance\")\nplt.tight_layout()\nplt.show()\n\n# Temporal importance\nplt.figure(figsize=(8, 4))\nplt.plot(time_imp, marker=\"o\")\nplt.title(\"Informer Time-step Importance (Grad×Input)\")\nplt.xlabel(\"Time step\")\nplt.ylabel(\"Importance\")\nplt.grid(alpha=0.3)\nplt.tight_layout()\nplt.show()\n\np_time = np.abs(time_imp)\np_time = p_time / (p_time.sum() + 1e-8)\n\nentropy_time = -np.sum(p_time * np.log(p_time + 1e-8))\ntop3_time = float(np.sort(p_time)[-3:].sum())\ntop5_time = float(np.sort(p_time)[-5:].sum())\n\ncum_time = np.cumsum(np.sort(p_time)[::-1])\nk80_time = int(np.argmax(cum_time >= 0.80) + 1)\nk90_time = int(np.argmax(cum_time >= 0.90) + 1)\n\nprint(\"\\n[Informer Gradient Temporal Metrics]\")\nprint(\"Entropy:\", round(float(entropy_time), 4))\nprint(\"Top-3 mass:\", round(top3_time, 4))\nprint(\"Top-5 mass:\", round(top5_time, 4))\nprint(\"k80:\", k80_time, \"| k90:\", k90_time)\n\n# Save complementary artifacts\nnp.save(OUT_DIR / \"inf_feature_importance.npy\", feat_imp.astype(np.float32))\nnp.save(OUT_DIR / \"inf_time_importance.npy\", time_imp.astype(np.float32))\nnp.save(OUT_DIR / \"inf_time_entropy.npy\", np.array([entropy_time], dtype=np.float32))\nnp.save(OUT_DIR / \"inf_time_topk_mass.npy\", np.array([top3_time, top5_time], dtype=np.float32))\nnp.save(OUT_DIR / \"inf_time_k_coverage.npy\", np.array([k80_time, k90_time], dtype=np.int32))\n\nprint(\"\\n Informer pipeline complete\")\nprint(\"Saved:\")\nprint(\"-\", OUT_DIR / \"inf_full_valid_pred.npy\")\nprint(\"-\", OUT_DIR / \"inf_full_test_pred.npy\")\nprint(\"-\", OUT_DIR / \"inf_feature_importance.npy\")\nprint(\"-\", OUT_DIR / \"inf_time_importance.npy\")\nprint(\"-\", OUT_DIR / \"inf_encoder_active_ratios.npy\")\nprint(\"-\", OUT_DIR / \"inf_decoder_cross_attention_mass.npy\")","metadata":{"_uuid":"4c1a7bd0-c1a7-45c5-9f27-6bb2c5bff835","_cell_guid":"bd56b967-ee90-4907-b306-988a6dc49088","trusted":true,"collapsed":false,"execution":{"execution_failed":"2026-06-05T18:26:24.623Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# MULTI-MODEL COMPARISON\n# ============================================================\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom sklearn.metrics import (\n    roc_curve, roc_auc_score,\n    precision_recall_curve, average_precision_score,\n    brier_score_loss\n)\nfrom sklearn.calibration import calibration_curve\n\n# ============================================================\n# 1) LABELS (STRICT)\n# ============================================================\ny_true = globals().get(\"y_true\", None)\n\nif y_true is None:\n    y_true = globals().get(\"y_va\", globals().get(\"y_valid\"))\n\nif y_true is None:\n    raise ValueError(\"Missing validation labels (y_true / y_va / y_valid)\")\n\ny_true = np.asarray(y_true).astype(int).reshape(-1)\n\nprint(f\"[INFO] y_true shape: {y_true.shape}\")\n\n\n# ============================================================\n# 2) ECE (EXPECTED CALIBRATION ERROR)\n# ============================================================\ndef expected_calibration_error(y_true, y_prob, n_bins=15):\n    bins = np.linspace(0, 1, n_bins + 1)\n    bin_ids = np.clip(np.digitize(y_prob, bins) - 1, 0, n_bins - 1)\n\n    bin_total = np.bincount(bin_ids, minlength=n_bins)\n    nonzero = bin_total > 0\n\n    bin_true = np.bincount(bin_ids, weights=y_true, minlength=n_bins)\n    bin_prob = np.bincount(bin_ids, weights=y_prob, minlength=n_bins)\n\n    avg_true = np.zeros(n_bins)\n    avg_prob = np.zeros(n_bins)\n\n    avg_true[nonzero] = bin_true[nonzero] / bin_total[nonzero]\n    avg_prob[nonzero] = bin_prob[nonzero] / bin_total[nonzero]\n\n    return float(np.sum(\n        np.abs(avg_true - avg_prob) * (bin_total / len(y_true))\n    ))\n\n\n# ============================================================\n# 3) COLLECT MODELS\n# ============================================================\nmodel_preds = []\n\ndef _add_model(name, keys):\n    for k in keys:\n        if k in globals() and globals()[k] is not None:\n            y_prob = np.asarray(globals()[k]).reshape(-1)\n\n            if len(y_prob) != len(y_true):\n                raise ValueError(\n                    f\"[ERROR] Length mismatch for {name}: \"\n                    f\"{len(y_prob)} vs {len(y_true)}\"\n                )\n\n            y_prob = np.clip(y_prob, 1e-7, 1 - 1e-7)\n\n            model_preds.append((name, y_prob))\n            print(f\"[OK] Loaded {name} from '{k}'\")\n            return\n\n    print(f\"[SKIP] {name} not found\")\n\n_add_model(\"LightGBM\",   [\"lgb_val_prob\", \"lgb_valid_pred\"])\n_add_model(\"XGBoost\",    [\"xgb_val_prob\", \"xgb_valid_pred\"])\n_add_model(\"Transformer\",[\"trf_valid_pred\"])\n_add_model(\"Informer\",   [\"inf_full_valid_pred\", \"inf_val_prob\"])\n_add_model(\"TFT\",        [\"tft_val_prob\", \"tft_valid_pred\"])\n\nif len(model_preds) == 0:\n    raise ValueError(\"No valid model predictions found\")\n\nprint(\"\\n Models used:\", [m[0] for m in model_preds])\n\n\n# ============================================================\n# 4) ROC CURVE\n# ============================================================\nplt.figure(figsize=(7, 6))\n\nfor name, y_prob in model_preds:\n    fpr, tpr, _ = roc_curve(y_true, y_prob)\n    auc = roc_auc_score(y_true, y_prob)\n    plt.plot(fpr, tpr, lw=2, label=f\"{name} ({auc:.3f})\")\n\nplt.plot([0,1],[0,1],'--', color='black')\nplt.title(\"ROC Curve\")\nplt.xlabel(\"False Positive Rate\")\nplt.ylabel(\"True Positive Rate\")\nplt.legend()\nplt.grid(alpha=0.3)\nplt.show()\n\n\n# ============================================================\n# 5) PRECISION-RECALL CURVE\n# ============================================================\nplt.figure(figsize=(7, 6))\n\nfor name, y_prob in model_preds:\n    precision, recall, _ = precision_recall_curve(y_true, y_prob)\n    ap = average_precision_score(y_true, y_prob)\n    plt.plot(recall, precision, lw=2, label=f\"{name} ({ap:.3f})\")\n\nplt.axhline(y_true.mean(), linestyle=\"--\", color=\"black\")\nplt.title(\"Precision-Recall Curve\")\nplt.xlabel(\"Recall\")\nplt.ylabel(\"Precision\")\nplt.legend()\nplt.grid(alpha=0.3)\nplt.show()\n\n\n# ============================================================\n# 6) CALIBRATION\n# ============================================================\nplt.figure(figsize=(7, 6))\n\nfor name, y_prob in model_preds:\n    pt, pp = calibration_curve(y_true, y_prob, n_bins=15)\n\n    ece = expected_calibration_error(y_true, y_prob)\n    brier = brier_score_loss(y_true, y_prob)\n\n    plt.plot(pp, pt, marker='o',\n             label=f\"{name} (ECE={ece:.3f}, Brier={brier:.3f})\")\n\nplt.plot([0,1],[0,1],'--', color='black')\nplt.title(\"Calibration Curve\")\nplt.xlabel(\"Predicted Probability\")\nplt.ylabel(\"Observed Probability\")\nplt.legend()\nplt.grid(alpha=0.3)\nplt.show()\n\n\n# ============================================================\n# 7) SUMMARY TABLE\n# ============================================================\nrows = []\n\nfor name, y_prob in model_preds:\n    row = {\n        \"Model\": name,\n        \"ROC-AUC\": roc_auc_score(y_true, y_prob),\n        \"PR-AUC\": average_precision_score(y_true, y_prob),\n        \"ECE\": expected_calibration_error(y_true, y_prob),\n        \"Brier\": brier_score_loss(y_true, y_prob),\n        \"AMEX\": amex_metric(y_true, y_prob)\n    }\n    rows.append(row)\n\nsummary = pd.DataFrame(rows).sort_values(\"AMEX\", ascending=False)\nsummary = summary.round(5)\n\nprint(\"\\n=== MODEL COMPARISON (FINAL) ===\")\nprint(summary)\n\n# ============================================================\n# 8) SAVE OUTPUTS\n# ============================================================\nCSV_PATH = \"final_model_comparison_summary.csv\"\nsummary.to_csv(CSV_PATH, index=False)\n\nprint(f\"\\n Saved CSV: {CSV_PATH}\")","metadata":{"_uuid":"2b38f178-9882-4b2d-8591-ee09c0b8894a","_cell_guid":"92260f16-3e4c-4c85-97fb-d92dc8d3205e","trusted":true,"collapsed":false,"execution":{"execution_failed":"2026-06-05T18:26:24.623Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# RESULTS SUMMARY TABLE\n# ============================================================\n\nimport pandas as pd\nimport numpy as np\nfrom pathlib import Path\n\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\n\n# ============================================================\n# 1) COLLECT ALL MODEL METRICS\n# ============================================================\nall_models = []\n\nmetric_keys = [\n    \"lgb_metrics\",\n    \"xgb_metrics\",\n    \"trf_metrics\",\n    \"tft_metrics\",\n    \"inf_metrics\",\n    \"inf_full_metrics\",\n]\n\nfor m in metric_keys:\n    if m in globals() and globals()[m] is not None:\n        all_models.append(globals()[m])\n\nif len(all_models) == 0:\n    raise ValueError(\"No model metrics found\")\n\nresults = pd.DataFrame.from_records(all_models)\n\n# ============================================================\n# 2) CLEAN MODEL NAMES\n# ============================================================\nif \"name\" in results.columns:\n    results = results.rename(columns={\"name\": \"Model\"})\n\nif \"Model\" not in results.columns:\n    raise ValueError(\"No model name column found\")\n\nresults[\"Model\"] = results[\"Model\"].astype(str)\nresults[\"Model\"] = results[\"Model\"].str.replace(\"(tuned)\", \"\", regex=False).str.strip()\nresults[\"Model_clean\"] = results[\"Model\"].str.lower()\n\n# ============================================================\n# 3) NORMALIZE METRIC COLUMN NAMES\n# ============================================================\n# Make AMEX column robust\nif \"amex\" not in results.columns and \"AMEX\" in results.columns:\n    results[\"amex\"] = results[\"AMEX\"]\n\nif \"amex\" not in results.columns:\n    raise ValueError(\"No AMEX metric found in results\")\n\n# Optional normalization for display consistency\nif \"auc\" in results.columns and \"ROC-AUC\" not in results.columns:\n    results[\"ROC-AUC\"] = results[\"auc\"]\n\nif \"pr_auc\" in results.columns and \"PR-AUC\" not in results.columns:\n    results[\"PR-AUC\"] = results[\"pr_auc\"]\n\nif \"brier\" in results.columns and \"Brier\" not in results.columns:\n    results[\"Brier\"] = results[\"brier\"]\n\n# ============================================================\n# 4) INTERPRETABILITY LEVEL\n# ============================================================\ndef get_interp(m):\n    m = m.lower()\n    if \"lightgbm\" in m:\n        return \"Medium\"\n    if \"xgboost\" in m:\n        return \"Medium\"\n    if \"transformer\" in m:\n        return \"High\"\n    if \"tft\" in m:\n        return \"Very High\"\n    if \"informer\" in m:\n        return \"Very High\"\n    return \"Unknown\"\n\nresults[\"Interpretability\"] = results[\"Model\"].apply(get_interp)\n\n# ============================================================\n# 5) SAFE FILE LOADING\n# ============================================================\ndef safe_load(path):\n    path = Path(path)\n    if not path.exists():\n        return None\n    try:\n        return np.load(path, allow_pickle=True)\n    except Exception:\n        return None\n\ndef scalar_from_file(path):\n    arr = safe_load(path)\n    if arr is None:\n        return None\n    arr = np.asarray(arr).reshape(-1)\n    if len(arr) == 0:\n        return None\n    return float(arr[0])\n\ndef pair_from_file(path):\n    arr = safe_load(path)\n    if arr is None:\n        return [None, None]\n    arr = np.asarray(arr).reshape(-1)\n    if len(arr) < 2:\n        return [None, None]\n    return [float(arr[0]), float(arr[1])]\n\ndef match_model(m, key):\n    return key in m.lower()\n\n# ============================================================\n# 6) LOAD INTERPRETABILITY METRICS\n# ============================================================\nentropy_map = {}\ntopk_map = {}\nkcov_map = {}\n\nfor m in results[\"Model\"]:\n\n    if match_model(m, \"lightgbm\"):\n        e = scalar_from_file(OUT_DIR / \"lgb_entropy.npy\")\n        t = pair_from_file(OUT_DIR / \"lgb_topk_mass.npy\")\n        k = pair_from_file(OUT_DIR / \"lgb_k_coverage.npy\")\n\n    elif match_model(m, \"xgboost\"):\n        e = scalar_from_file(OUT_DIR / \"xgb_entropy.npy\")\n        t = pair_from_file(OUT_DIR / \"xgb_topk_mass.npy\")\n        k = pair_from_file(OUT_DIR / \"xgb_k_coverage.npy\")\n\n    elif match_model(m, \"transformer\"):\n        e = scalar_from_file(OUT_DIR / \"trf_time_entropy.npy\")\n        t = pair_from_file(OUT_DIR / \"trf_time_topk_mass.npy\")\n        k = pair_from_file(OUT_DIR / \"trf_time_k_coverage.npy\")\n\n    elif match_model(m, \"tft\"):\n        e = scalar_from_file(OUT_DIR / \"tft_attention_entropy.npy\")\n        t = pair_from_file(OUT_DIR / \"tft_attention_topk_mass.npy\")\n        k = pair_from_file(OUT_DIR / \"tft_attention_k_coverage.npy\")\n\n    elif match_model(m, \"informer\"):\n        # Support both built-in and gradient-based variants\n        e = scalar_from_file(OUT_DIR / \"inf_builtin_entropy.npy\")\n        if e is None:\n            e = scalar_from_file(OUT_DIR / \"inf_time_entropy.npy\")\n\n        t = pair_from_file(OUT_DIR / \"inf_builtin_topk_mass.npy\")\n        if t == [None, None]:\n            t = pair_from_file(OUT_DIR / \"inf_time_topk_mass.npy\")\n\n        k = pair_from_file(OUT_DIR / \"inf_builtin_k_coverage.npy\")\n        if k == [None, None]:\n            k = pair_from_file(OUT_DIR / \"inf_time_k_coverage.npy\")\n\n    else:\n        e, t, k = None, [None, None], [None, None]\n\n    entropy_map[m] = e\n    topk_map[m] = t\n    kcov_map[m] = k\n\n# attach\nresults[\"Entropy\"] = results[\"Model\"].map(entropy_map)\nresults[\"Top-3 Mass\"] = results[\"Model\"].apply(lambda m: topk_map[m][0] if m in topk_map else None)\nresults[\"Top-5 Mass\"] = results[\"Model\"].apply(lambda m: topk_map[m][1] if m in topk_map else None)\nresults[\"k80\"] = results[\"Model\"].apply(lambda m: kcov_map[m][0] if m in kcov_map else None)\nresults[\"k90\"] = results[\"Model\"].apply(lambda m: kcov_map[m][1] if m in kcov_map else None)\n\n# ============================================================\n# 7) SORT + RANK\n# ============================================================\nresults = results.sort_values(\"amex\", ascending=False).reset_index(drop=True)\n\nresults[\"Rank\"] = results[\"amex\"].rank(ascending=False, method=\"dense\").astype(int)\n\nbest_amex = results[\"amex\"].max()\nresults[\"Δ AMEX vs Best\"] = (best_amex - results[\"amex\"]).round(5)\n\n# ============================================================\n# 8) DISPLAY\n# ============================================================\nresults_display = results.copy()\n\nfor c in results_display.select_dtypes(include=[np.number]).columns:\n    results_display[c] = results_display[c].round(5)\n\nprint(\"\\n=== FINAL MODEL SUMMARY ===\")\nprint(results_display)\n\n# ============================================================\n# 9) SAVE\n# ============================================================\nsave_path = OUT_DIR / \"final_model_comparison_summary.csv\"\nresults.to_csv(save_path, index=False)\n\nprint(f\"\\n Saved: {save_path}\")","metadata":{"_uuid":"2607a167-dd21-4fa3-a4dd-8ed692bdeef6","_cell_guid":"2fd8b9a2-a737-4992-bdda-47444a65368a","trusted":true,"collapsed":false,"execution":{"execution_failed":"2026-06-05T18:26:24.623Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 10) COMBINED ROC + PR + CALIBRATION PLOTS (ALL MODELS)","metadata":{"_uuid":"a860274e-873f-4a0f-833f-eb3f80b4092c","_cell_guid":"998ad2af-d69e-404a-849b-3ceb3e45e88f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# COMBINED ROC + PR + CALIBRATION PLOTS\n# ============================================================\n\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nfrom sklearn.metrics import (\n    roc_curve, roc_auc_score,\n    precision_recall_curve,\n    average_precision_score,\n    brier_score_loss\n)\nfrom sklearn.calibration import calibration_curve\n\n\n# ============================================================\n# 1) Resolve y_true (STRICT)\n# ============================================================\ny_true = globals().get(\"y_true\", None)\n\nif y_true is None:\n    y_true = globals().get(\"y_va\", globals().get(\"y_valid\"))\n\nif y_true is None:\n    raise ValueError(\n        \"Could not find validation labels. Expected one of: y_true, y_va, y_valid\"\n    )\n\ny_true = np.asarray(y_true).astype(int).reshape(-1)\n\nprint(f\"[INFO] y_true shape: {y_true.shape}\")\n\n\n# ============================================================\n# 2) Collect model probabilities\n# ============================================================\nmodel_preds = []\n\ndef _add_model(name, candidate_keys):\n\n    for k in candidate_keys:\n        if k in globals() and globals()[k] is not None:\n\n            y_prob = np.asarray(globals()[k]).reshape(-1)\n\n            # VALIDATION\n            if len(y_prob) != len(y_true):\n                raise ValueError(\n                    f\"[ERROR] {name} length mismatch: \"\n                    f\"{len(y_prob)} vs {len(y_true)}\"\n                )\n\n            y_prob = np.clip(y_prob, 1e-7, 1 - 1e-7)\n\n            model_preds.append((name, y_prob))\n            print(f\"[OK] Loaded {name} from '{k}'\")\n            return\n\n    print(f\"[SKIP] {name}: no valid predictions found\")\n\n\n_add_model(\"LightGBM\",   [\"lgb_val_prob\", \"lgb_valid_pred\", \"best_lgb_pred\"])\n_add_model(\"XGBoost\",    [\"xgb_val_prob\", \"xgb_valid_pred\", \"best_xgb_pred\"])\n_add_model(\"Transformer\",[\"trf_valid_pred\", \"trf_val_prob\", \"best_pred\"])\n_add_model(\"TFT\",        [\"tft_val_prob\", \"tft_valid_pred\", \"val_prob\"])\n_add_model(\"Informer\",   [\"inf_full_valid_pred\", \"inf_val_prob\"])\n\n\nif len(model_preds) == 0:\n    raise ValueError(\"No model prediction arrays found for plotting.\")\n\nprint(\"\\n Models included:\", [m[0] for m in model_preds])\n\n\n# ============================================================\n# 3) Expected Calibration Error (ECE)\n# ============================================================\ndef expected_calibration_error(y_true, y_prob, n_bins=15):\n\n    bins = np.linspace(0.0, 1.0, n_bins + 1)\n    bin_ids = np.clip(np.digitize(y_prob, bins) - 1, 0, n_bins - 1)\n\n    bin_total = np.bincount(bin_ids, minlength=n_bins)\n    nonzero = bin_total > 0\n\n    bin_true = np.bincount(bin_ids, weights=y_true, minlength=n_bins)\n    bin_prob = np.bincount(bin_ids, weights=y_prob, minlength=n_bins)\n\n    avg_true = np.zeros(n_bins)\n    avg_prob = np.zeros(n_bins)\n\n    avg_true[nonzero] = bin_true[nonzero] / bin_total[nonzero]\n    avg_prob[nonzero] = bin_prob[nonzero] / bin_total[nonzero]\n\n    ece = np.sum(np.abs(avg_true - avg_prob) * (bin_total / len(y_true)))\n    return float(ece)\n\n\n# ============================================================\n# 4) ROC CURVE\n# ============================================================\nplt.figure(figsize=(8, 6))\n\nfor name, y_prob in model_preds:\n    fpr, tpr, _ = roc_curve(y_true, y_prob)\n    auc = roc_auc_score(y_true, y_prob)\n\n    plt.plot(fpr, tpr, lw=2, label=f\"{name} (AUC={auc:.3f})\")\n\nplt.plot([0, 1], [0, 1], linestyle=\"--\", color=\"black\", alpha=0.7, label=\"Random\")\n\nplt.title(\"Combined ROC Curve\")\nplt.xlabel(\"False Positive Rate\")\nplt.ylabel(\"True Positive Rate\")\nplt.legend(loc=\"lower right\")\nplt.grid(alpha=0.3)\nplt.tight_layout()\nplt.show()\n\n\n# ============================================================\n# 5) PRECISION-RECALL CURVE\n# ============================================================\nplt.figure(figsize=(8, 6))\n\nfor name, y_prob in model_preds:\n    precision, recall, _ = precision_recall_curve(y_true, y_prob)\n    ap = average_precision_score(y_true, y_prob)\n\n    plt.plot(recall, precision, lw=2, label=f\"{name} (AP={ap:.3f})\")\n\nbaseline = y_true.mean()\nplt.axhline(baseline, linestyle=\"--\", color=\"black\", alpha=0.7,\n            label=f\"Baseline={baseline:.3f}\")\n\nplt.title(\"Combined Precision-Recall Curve\")\nplt.xlabel(\"Recall\")\nplt.ylabel(\"Precision\")\nplt.legend(loc=\"lower left\")\nplt.grid(alpha=0.3)\nplt.tight_layout()\nplt.show()\n\n\n# ============================================================\n# 6) CALIBRATION CURVE\n# ============================================================\nplt.figure(figsize=(8, 6))\n\nfor name, y_prob in model_preds:\n\n    prob_true, prob_pred = calibration_curve(y_true, y_prob, n_bins=15)\n\n    ece = expected_calibration_error(y_true, y_prob)\n    brier = brier_score_loss(y_true, y_prob)\n\n    plt.plot(\n        prob_pred,\n        prob_true,\n        marker=\"o\",\n        lw=2,\n        label=f\"{name} (ECE={ece:.3f}, Brier={brier:.3f})\"\n    )\n\nplt.plot([0, 1], [0, 1], linestyle=\"--\", color=\"black\", alpha=0.7,\n         label=\"Perfect Calibration\")\n\nplt.title(\"Combined Calibration Curve\")\nplt.xlabel(\"Predicted Probability\")\nplt.ylabel(\"Observed Default Frequency\")\nplt.legend(bbox_to_anchor=(1.02, 1), loc=\"upper left\")\nplt.grid(alpha=0.3)\nplt.tight_layout()\nplt.show()\n\nprint(\"\\n Combined evaluation plots complete\")","metadata":{"_uuid":"045abf34-d234-4a3d-b6af-b317d3efd266","_cell_guid":"cef9183e-7086-491b-a6f1-53e81f406f50","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 11) COMBINED ROC-AUC + PR-AUC + ECE + BRIER TABLE (ALL MODELS)","metadata":{"_uuid":"ea2a5329-d77e-4d57-82f0-9f73e6e222bc","_cell_guid":"4b146e81-1ad6-42fb-95af-9f038c236d04","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# FINAL METRICS TABLE\n# ============================================================\n\nimport numpy as np\nimport pandas as pd\n\nfrom sklearn.metrics import (\n    roc_auc_score,\n    average_precision_score,\n    brier_score_loss\n)\n\n# ============================================================\n# 1) LABELS\n# ============================================================\ny_true = globals().get(\"y_true\", None)\nif y_true is None:\n    y_true = globals().get(\"y_va\", globals().get(\"y_valid\"))\n\nif y_true is None:\n    raise ValueError(\"Missing validation labels (expected y_true, y_va, or y_valid)\")\n\ny_true = np.asarray(y_true).astype(int).reshape(-1)\n\nprint(f\"[INFO] y_true shape: {y_true.shape}\")\n\n# ============================================================\n# 2) ECE\n# ============================================================\ndef expected_calibration_error(y_true, y_prob, n_bins=15):\n    bins = np.linspace(0.0, 1.0, n_bins + 1)\n    bin_ids = np.clip(np.digitize(y_prob, bins) - 1, 0, n_bins - 1)\n\n    bin_total = np.bincount(bin_ids, minlength=n_bins)\n    nonzero = bin_total > 0\n\n    bin_true = np.bincount(bin_ids, weights=y_true, minlength=n_bins)\n    bin_prob = np.bincount(bin_ids, weights=y_prob, minlength=n_bins)\n\n    avg_true = np.zeros(n_bins)\n    avg_prob = np.zeros(n_bins)\n\n    avg_true[nonzero] = bin_true[nonzero] / bin_total[nonzero]\n    avg_prob[nonzero] = bin_prob[nonzero] / bin_total[nonzero]\n\n    return float(np.sum(np.abs(avg_true - avg_prob) * (bin_total / len(y_true))))\n\n# ============================================================\n# 3) COLLECT MODELS (ROBUST + STRICT)\n# ============================================================\nmodel_preds = []\n\ndef _add_model(name, keys):\n    for k in keys:\n        if k in globals() and globals()[k] is not None:\n            y_prob = np.asarray(globals()[k]).reshape(-1)\n\n            if len(y_prob) != len(y_true):\n                raise ValueError(\n                    f\"[ERROR] {name} length mismatch: {len(y_prob)} vs {len(y_true)}\"\n                )\n\n            y_prob = np.clip(y_prob, 1e-7, 1 - 1e-7)\n            model_preds.append((name, y_prob))\n            print(f\"[OK] Loaded {name} from '{k}'\")\n            return\n\n    print(f\"[SKIP] {name}: no matching validation prediction vector found\")\n\n_add_model(\"LightGBM\",   [\"lgb_val_prob\", \"lgb_valid_pred\", \"best_lgb_pred\"])\n_add_model(\"XGBoost\",    [\"xgb_val_prob\", \"xgb_valid_pred\", \"best_xgb_pred\"])\n_add_model(\"Transformer\",[\"trf_valid_pred\", \"trf_val_prob\", \"best_pred\"])\n_add_model(\"Informer\",   [\"inf_full_valid_pred\", \"inf_val_prob\"])\n_add_model(\"TFT\",        [\"tft_val_prob\", \"tft_valid_pred\", \"val_prob\"])\n\nif len(model_preds) == 0:\n    raise ValueError(\"No model prediction arrays found.\")\n\nprint(\"\\n Models included:\", [m[0] for m in model_preds])\n\n# ============================================================\n# 4) BUILD TABLE\n# ============================================================\nrows = []\n\nfor name, y_prob in model_preds:\n    row = {\n        \"Model\": name,\n        \"ROC-AUC\": roc_auc_score(y_true, y_prob),\n        \"PR-AUC\": average_precision_score(y_true, y_prob),\n        \"ECE\": expected_calibration_error(y_true, y_prob),\n        \"Brier\": brier_score_loss(y_true, y_prob)\n    }\n\n    if \"amex_metric\" in globals() and amex_metric is not None:\n        try:\n            row[\"AMEX\"] = float(amex_metric(y_true, y_prob))\n        except Exception:\n            row[\"AMEX\"] = np.nan\n\n    rows.append(row)\n\nsummary = pd.DataFrame(rows)\n\n# ============================================================\n# 5) SORT\n# ============================================================\nif \"AMEX\" in summary.columns:\n    summary = summary.sort_values(\n        [\"AMEX\", \"PR-AUC\", \"ROC-AUC\"],\n        ascending=False\n    ).reset_index(drop=True)\nelse:\n    summary = summary.sort_values(\n        [\"PR-AUC\", \"ROC-AUC\"],\n        ascending=False\n    ).reset_index(drop=True)\n\nsummary = summary.round(5)\n\n# ============================================================\n# 6) DISPLAY\n# ============================================================\nprint(\"\\n=== FINAL METRICS TABLE ===\")\nprint(summary)\n\nsummary.to_csv(\"combined_discrimination_calibration_summary.csv\", index=False)\n\nprint(\"\\n Saved: combined_discrimination_calibration_summary.csv\")","metadata":{"_uuid":"298f6944-73ad-414a-a2dd-1cc532c8095e","_cell_guid":"21d36e44-bab5-4cce-baed-f7d7378909e3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 12) Calibration & Cost-Sensitive Evaluation\n\nThis section evaluates model performance beyond accuracy by assessing both **probability calibration** and **business decision impact**.\n\n\n### 1. Calibration Analysis\n\nCalibration analysis evaluates how well predicted probabilities reflect actual default outcomes.\n\n- Calibration curves (reliability diagrams) are plotted for each model\n- Two key metrics are computed:\n  - **ECE (Expected Calibration Error)**  \n    → Measures the gap between predicted probabilities and observed outcomes  \n  - **Brier Score**  \n    → Measures overall probability accuracy (lower is better)\n\n\n### 2. Cost-Sensitive Threshold Optimisation\n\nInstead of using a fixed classification threshold (e.g. 0.5), an optimal threshold is determined by minimising:","metadata":{"_uuid":"140120fb-dfd5-498a-8b2e-8057308c061c","_cell_guid":"8fc90246-1bff-45f6-9322-d36627d076a3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# CALIBRATION + COST PIPELINE\n# ============================================================\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nfrom sklearn.metrics import confusion_matrix, brier_score_loss\nfrom sklearn.calibration import calibration_curve\n\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\n\n# ============================================================\n# 1) LABELS\n# ============================================================\ny_true = globals().get(\"y_true\", None)\n\nif y_true is None:\n    y_true = globals().get(\"y_va\", globals().get(\"y_valid\"))\n\nif y_true is None:\n    raise ValueError(\"Missing validation labels\")\n\ny_true = np.asarray(y_true).astype(int).reshape(-1)\n\n# ============================================================\n# 2) ECE\n# ============================================================\ndef expected_calibration_error(y_true, y_prob, n_bins=15):\n\n    bins = np.linspace(0.0, 1.0, n_bins + 1)\n    bin_ids = np.clip(np.digitize(y_prob, bins) - 1, 0, n_bins - 1)\n\n    bin_total = np.bincount(bin_ids, minlength=n_bins)\n    nonzero = bin_total > 0\n\n    bin_true = np.bincount(bin_ids, weights=y_true, minlength=n_bins)\n    bin_prob = np.bincount(bin_ids, weights=y_prob, minlength=n_bins)\n\n    avg_true = np.zeros(n_bins)\n    avg_prob = np.zeros(n_bins)\n\n    avg_true[nonzero] = bin_true[nonzero] / bin_total[nonzero]\n    avg_prob[nonzero] = bin_prob[nonzero] / bin_total[nonzero]\n\n    return float(np.sum(np.abs(avg_true - avg_prob) * (bin_total / len(y_true))))\n\n# ============================================================\n# 3) COST OPTIMIZATION\n# ============================================================\ndef best_threshold_cost_sensitive(y_true, y_prob, fn_cost=5.0, fp_cost=1.0):\n\n    order = np.argsort(-y_prob)\n    y_sorted = y_true[order]\n    p_sorted = y_prob[order]\n\n    tp = np.cumsum(y_sorted)\n    fp = np.cumsum(1 - y_sorted)\n\n    total_pos = tp[-1]\n    fn = total_pos - tp\n\n    cost = fn_cost * fn + fp_cost * fp\n    idx = np.argmin(cost)\n\n    thr = (\n        p_sorted[idx] if idx == len(p_sorted) - 1\n        else (p_sorted[idx] + p_sorted[idx+1]) / 2\n    )\n\n    yhat = (y_prob >= thr).astype(int)\n    tn, fp_val, fn_val, tp_val = confusion_matrix(y_true, yhat).ravel()\n\n    return float(thr), float(cost[idx]), (tn, fp_val, fn_val, tp_val)\n\n# ============================================================\n# 4) COLLECT MODELS\n# ============================================================\nmodel_preds = []\n\ndef _add(name, keys):\n    for k in keys:\n        if k in globals() and globals()[k] is not None:\n\n            yprob = np.asarray(globals()[k]).reshape(-1)\n\n            if len(yprob) != len(y_true):\n                raise ValueError(\n                    f\"{name} mismatch: {len(yprob)} vs {len(y_true)}\"\n                )\n\n            yprob = np.clip(yprob, 1e-7, 1 - 1e-7)\n            model_preds.append((name, yprob))\n            print(f\"[OK] {name}\")\n            return\n\n    print(f\"[SKIP] {name}\")\n\n_add(\"LightGBM\",   [\"lgb_val_prob\", \"lgb_valid_pred\"])\n_add(\"XGBoost\",    [\"xgb_val_prob\", \"xgb_valid_pred\"])\n_add(\"Transformer\",[\"trf_valid_pred\"])\n_add(\"Informer\",   [\"inf_full_valid_pred\", \"inf_val_prob\"])\n_add(\"TFT\",        [\"tft_val_prob\"])\n\nif len(model_preds) == 0:\n    raise ValueError(\"No model predictions found.\")\n\n# ============================================================\n# 5) RUN PIPELINE\n# ============================================================\nFN_COST = 5.0\nFP_COST = 1.0\n\ncalib_rows = []\ncost_rows = []\n\nplt.figure(figsize=(8,6))\nplt.plot([0,1],[0,1],'--', label=\"Perfect\")\n\nfor name, yprob in model_preds:\n\n    # --- calibration ---\n    pt, pp = calibration_curve(y_true, yprob, n_bins=15)\n    ece = expected_calibration_error(y_true, yprob)\n    brier = brier_score_loss(y_true, yprob)\n\n    plt.plot(pp, pt, marker=\"o\", label=f\"{name} (ECE={ece:.3f})\")\n\n    calib_rows.append({\n        \"Model\": name,\n        \"ECE\": ece,\n        \"Brier\": brier\n    })\n\n    # --- optimal cost ---\n    thr, min_cost, (tn, fp, fn, tp) = best_threshold_cost_sensitive(\n        y_true, yprob, FN_COST, FP_COST\n    )\n\n    # --- cost at 0.5 ---\n    yhat_05 = (yprob >= 0.5).astype(int)\n    tn05, fp05, fn05, tp05 = confusion_matrix(y_true, yhat_05).ravel()\n\n    cost_05 = FN_COST * fn05 + FP_COST * fp05\n\n    cost_rows.append({\n        \"Model\": name,\n        \"thr_opt\": thr,\n        \"min_cost\": min_cost,\n        \"cost@0.5\": cost_05,\n        \"saving\": cost_05 - min_cost,\n        \"TP\": tp, \"FP\": fp, \"FN\": fn, \"TN\": tn\n    })\n\nplt.title(\"Calibration Curves\")\nplt.xlabel(\"Predicted Probability\")\nplt.ylabel(\"Observed Frequency\")\nplt.legend()\nplt.grid(alpha=0.3)\nplt.show()\n\n# ============================================================\n# 6) TABLES\n# ============================================================\ncalibration_table = pd.DataFrame(calib_rows).sort_values(\"ECE\")\ncost_table = pd.DataFrame(cost_rows).sort_values(\"min_cost\")\n\nprint(\"\\n=== Calibration Summary ===\")\nprint(calibration_table.round(5))\n\nprint(\"\\n=== Cost Optimization Summary ===\")\nprint(cost_table.round(5))\n\n# ============================================================\n# SAVE\n# ============================================================\ncalibration_table.to_csv(OUT_DIR / \"calibration_summary.csv\", index=False)\ncost_table.to_csv(OUT_DIR / \"cost_summary.csv\", index=False)\n\nprint(\"\\n Calibration + Cost pipeline complete\")","metadata":{"_uuid":"450359c6-ccb8-498b-a71c-7fdbc7ecb6de","_cell_guid":"e3105168-e25e-4d4e-9cf2-b72ef67f5b19","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 13) Chapter 4: Tables & Export Pipeline\n\nThis section generates and exports all key results tables that will be used in Chapter 4, ensuring **reproducibility and publication-ready formatting**.\n\n---\n\n### Export Helper\n\nA reusable function is used to export tables to:\n\n- **CSV (full precision)** → for analysis and reproducibility  \n- **LaTeX (longtable format)** → for direct inclusion in the thesis  \n\n**Key features:**\n- Optional timestamping for experiment tracking  \n- Automatic directory creation  \n- Safe handling of missing and infinite values  \n- Consistent float formatting for publication  \n\n---\n\n### Generated Tables\n\nThe pipeline produces the following Chapter 4 tables:\n\n#### **Table 4.1 — Experimental Setup**\nSummarises:\n- Dataset and preprocessing approach  \n- Feature engineering strategy  \n- Validation methodology  \n- Metrics and interpretability techniques  \n\n---\n\n#### **Table 4.2 — Performance Comparison**\nCompares all models across:\n- AMEX (primary metric)  \n- AUC, PR-AUC, Brier, F1, Accuracy  \n- Model type (Tree vs Deep Learning)  \n- Interpretability level  \n\nAdditional features:\n- Model ranking (based on AMEX)  \n- Performance gap vs best model  \n\n---\n\n#### **Table 4.3 — Confusion Matrix Summary**\nReports:\n- TN, FP, FN, TP  \n- Decision threshold used  \n- **Cost metric (FN weighted 5:1 vs FP)**  \n\n---\n\n#### **Table 4.4 — Interpretability Comparison**\nSummarises interpretability across models:\n- Method used (e.g., Feature Importance, Gradient × Input, TFT attention)  \n- Whether temporal dynamics are captured  \n- Overall interpretability level  \n\n---\n\n#### **Table 4.5 — Efficiency Comparison**\nTracks:\n- Training time  \n- Inference time  \n- AMEX score  \n\nSupports computational performance analysis.\n\n---\n\n#### **Table 4.6 — Model Selection**\nIdentifies the final selected model based on:\n- Best AMEX score  \n- Performance gap vs second-best model  \n- Interpretability considerations  \n\n---\n\n### Outputs\n\nEach table is exported as:\n\n- `.csv` → for reproducibility  \n- `.tex` → for direct inclusion in LaTeX thesis  \n\n---\n\n### Why this matters\n\nThis pipeline ensures:\n\n- **Consistency** → all results follow the same structure  \n- **Reproducibility** → tables can be regenerated automatically  \n- **Traceability** → experimental outputs are version-controlled (timestamps optional)  \n- **Publication readiness** → minimal manual formatting required  \n\nEnables a structured and defensible presentation of results in Chapter 4.","metadata":{"_uuid":"0484f9dc-b5d9-464d-8e7e-5836f7948a02","_cell_guid":"af58a49a-80d5-4c2f-88ce-1b4169b7ad79","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ============================================================\n# CALIBRATION + COST PIPELINE \n# ============================================================\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nfrom sklearn.metrics import confusion_matrix, brier_score_loss\nfrom sklearn.calibration import calibration_curve\n\n# ============================================================\n# CONFIG\n# ============================================================\nOUT_DIR = Path(\"/kaggle/working/seq_artifacts\")\n\nFN_COST = 5.0   # cost of missing default (FN)\nFP_COST = 1.0   # cost of false alarm (FP)\n\n# ============================================================\n# 1) LABELS \n# ============================================================\ny_true = globals().get(\"y_true\", None)\n\nif y_true is None:\n    y_true = globals().get(\"y_va\", globals().get(\"y_valid\"))\n\nif y_true is None:\n    raise ValueError(\"Missing validation labels (expected y_true / y_va / y_valid)\")\n\ny_true = np.asarray(y_true).astype(int).reshape(-1)\n\nprint(\"[INFO] y_true shape:\", y_true.shape)\n\n# ============================================================\n# 2) ECE\n# ============================================================\ndef expected_calibration_error(y_true, y_prob, n_bins=15):\n\n    bins = np.linspace(0.0, 1.0, n_bins + 1)\n    bin_ids = np.clip(np.digitize(y_prob, bins) - 1, 0, n_bins - 1)\n\n    bin_total = np.bincount(bin_ids, minlength=n_bins)\n    nonzero = bin_total > 0\n\n    bin_true = np.bincount(bin_ids, weights=y_true, minlength=n_bins)\n    bin_prob = np.bincount(bin_ids, weights=y_prob, minlength=n_bins)\n\n    avg_true = np.zeros(n_bins)\n    avg_prob = np.zeros(n_bins)\n\n    avg_true[nonzero] = bin_true[nonzero] / bin_total[nonzero]\n    avg_prob[nonzero] = bin_prob[nonzero] / bin_total[nonzero]\n\n    return float(np.sum(np.abs(avg_true - avg_prob) * (bin_total / len(y_true)))\n\n# ============================================================\n# 3) COST FUNCTIONS\n# ============================================================\ndef best_threshold_cost_sensitive(y_true, y_prob, fn_cost=5.0, fp_cost=1.0):\n\n    order = np.argsort(-y_prob)\n    y_sorted = y_true[order]\n    p_sorted = y_prob[order]\n\n    tp = np.cumsum(y_sorted)\n    fp = np.cumsum(1 - y_sorted)\n\n    total_pos = tp[-1]\n    fn = total_pos - tp\n\n    cost = fn_cost * fn + fp_cost * fp\n    idx = np.argmin(cost)\n\n    thr = (\n        p_sorted[idx] if idx == len(p_sorted) - 1\n        else (p_sorted[idx] + p_sorted[idx + 1]) / 2\n    )\n\n    yhat = (y_prob >= thr).astype(int)\n    tn, fp_val, fn_val, tp_val = confusion_matrix(y_true, yhat).ravel()\n\n    return float(thr), float(cost[idx]), (tn, fp_val, fn_val, tp_val)\n\ndef cost_at_threshold(y_true, y_prob, thr, fn_cost=5.0, fp_cost=1.0):\n\n    yhat = (y_prob >= thr).astype(int)\n    tn, fp, fn, tp = confusion_matrix(y_true, yhat).ravel()\n\n    return fn_cost * fn + fp_cost * fp\n\n# ============================================================\n# 4) COLLECT MODELS\n# ============================================================\nmodel_preds = []\n\ndef _add(name, keys):\n    for k in keys:\n        if k in globals() and globals()[k] is not None:\n\n            yprob = np.asarray(globals()[k]).reshape(-1)\n\n            # STRICT CHECK (no silent skipping)\n            if len(yprob) != len(y_true):\n                raise ValueError(\n                    f\"[ERROR] {name} length mismatch: {len(yprob)} vs {len(y_true)}\"\n                )\n\n            yprob = np.clip(yprob, 1e-7, 1 - 1e-7)\n\n            model_preds.append((name, yprob))\n            print(f\"[OK] Loaded {name} from '{k}'\")\n            return\n\n    print(f\"[SKIP] {name} not found\")\n\n_add(\"LightGBM\",   [\"lgb_val_prob\", \"lgb_valid_pred\"])\n_add(\"XGBoost\",    [\"xgb_val_prob\", \"xgb_valid_pred\"])\n_add(\"Transformer\",[\"trf_valid_pred\", \"trf_val_prob\"])\n_add(\"Informer\",   [\"inf_full_valid_pred\", \"inf_val_prob\"])\n_add(\"TFT\",        [\"tft_val_prob\", \"tft_valid_pred\"])\n\nif len(model_preds) == 0:\n    raise ValueError(\"No valid model predictions found\")\n\nprint(\"\\n Models included:\", [m[0] for m in model_preds])\n\n# ============================================================\n# 5) RUN PIPELINE\n# ============================================================\n\ncalib_rows = []\ncost_rows = []\n\nplt.figure(figsize=(8, 6))\nplt.plot([0, 1], [0, 1], '--', label=\"Perfect Calibration\")\n\nfor name, yprob in model_preds:\n\n    # --------------------------------------------------------\n    # CALIBRATION\n    # --------------------------------------------------------\n    prob_true, prob_pred = calibration_curve(y_true, yprob, n_bins=15)\n\n    ece = expected_calibration_error(y_true, yprob)\n    brier = brier_score_loss(y_true, yprob)\n\n    plt.plot(prob_pred, prob_true, marker=\"o\", label=f\"{name} (ECE={ece:.3f})\")\n\n    calib_rows.append({\n        \"Model\": name,\n        \"ECE\": ece,\n        \"Brier\": brier\n    })\n\n    # --------------------------------------------------------\n    # COST-SENSITIVE OPTIMIZATION\n    # --------------------------------------------------------\n    thr, min_cost, (tn, fp, fn, tp) = best_threshold_cost_sensitive(\n        y_true, yprob, FN_COST, FP_COST\n    )\n\n    cost_05 = cost_at_threshold(y_true, yprob, 0.5, FN_COST, FP_COST)\n\n    cost_rows.append({\n        \"Model\": name,\n        \"thr_opt\": thr,\n        \"min_cost\": min_cost,\n        \"cost@0.5\": cost_05,\n        \"saving\": cost_05 - min_cost,\n        \"TP\": tp,\n        \"FP\": fp,\n        \"FN\": fn,\n        \"TN\": tn\n    })\n\n# ============================================================\n# 6) PLOT\n# ============================================================\n\nplt.title(\"Calibration Curves\")\nplt.xlabel(\"Predicted Probability\")\nplt.ylabel(\"Observed Default Frequency\")\nplt.legend()\nplt.grid(alpha=0.3)\nplt.tight_layout()\nplt.show()\n\n# ============================================================\n# 7) TABLES\n# ============================================================\n\ncalibration_table = pd.DataFrame(calib_rows).sort_values(\"ECE\")\ncost_table = pd.DataFrame(cost_rows).sort_values(\"min_cost\")\n\nprint(\"\\n=== Calibration Summary ===\")\nprint(calibration_table.round(5))\n\nprint(\"\\n=== Cost Optimization Summary ===\")\nprint(cost_table.round(5))\n\n# ============================================================\n# 8) SAVE OUTPUTS\n# ============================================================\n\ncalibration_table.to_csv(OUT_DIR / \"calibration_summary.csv\", index=False)\ncost_table.to_csv(OUT_DIR / \"cost_summary.csv\", index=False)\n\nprint(\"\\n Calibration + Cost pipeline complete\")","metadata":{"_uuid":"7aa937fd-b942-4b28-a92f-e0282616310d","_cell_guid":"18cc0c08-77d2-4f51-b072-93d1a96d6e8f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Notes for my Dissertation (where this fits)\n\n### Chapter 3 (Methodology)\n- Data preprocessing: renaming, datetime parsing, train-fitted imputation\n- Two modeling tracks:\n  1) **Customer-level aggregation** (tree baselines)\n  2) **Sequence modelling** (transformer)\n- Imbalance handling:\n  - `scale_pos_weight` for trees\n  - `WeightedRandomSampler` for transformer\n- Metrics: AMEX + AUC + PR-AUC + Brier + confusion-matrix metrics\n- Interpretability:\n  - Trees: gain/split importance\n  - Transformer: gradient attribution (feature + time-step)\n\n### Chapter 4 (Experiments & Results)\n- Report the `results` dataframe\n- Include:\n  - ROC/PR curves for each model\n  - Feature importance plots\n  - Transformer attribution plots","metadata":{"_uuid":"2ff45c31-ae7e-4258-9586-9bc34714ffed","_cell_guid":"82cca216-3d93-49e1-a0e4-618239123a8b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"","metadata":{"_uuid":"2b7fd14f-135f-4b38-a12c-1d21818656a0","_cell_guid":"56d1712b-a4e8-4bbb-92ae-0c0285d37573","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null}]}