{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":38760,"databundleVersionId":4493939},{"sourceType":"datasetVersion","sourceId":4436180,"datasetId":2597726,"databundleVersionId":4495742},{"sourceType":"datasetVersion","sourceId":4809111,"datasetId":2784704,"databundleVersionId":4872635},{"sourceType":"datasetVersion","sourceId":4819222,"datasetId":2787166,"databundleVersionId":4882874},{"sourceType":"datasetVersion","sourceId":4483558,"datasetId":2623568,"databundleVersionId":4543798}],"dockerImageVersionId":31328,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n# Introduction of Notebook\n\nNotebook này triển khai kiến trúc Two-Stage Recommender System, được chia thành 3 bước xử lý chính như sau:\n\n### Bước 1: Model Training\n*   **Bước 1.1 Loading Training Data:**\n    Để tối ưu thời gian chạy nháp và gỡ lỗi , dữ liệu huấn luyện được trích mẫu ngẫu nhiên bằng phép chia lấy dư: `train_df = train_df[train_df['session'] % 10 == 1]`. Nhãn thực tế được lấy từ tệp `test_labels.parquet` tệp này chứa nhãn cho cả tập train và test cho local validation\n*   **Bước 1.2  Feature Engineering :**\n    Nhằm tiết kiệm bộ nhớ, toàn bộ các đặc trưng đã được **tính toán từ trước (pre-calculated)** và lưu trữ dưới dạng tệp Parquet tại dataset `/kaggle/input/otto-validation`. Tại bước này, chúng ta chỉ cần thực hiện thao tác `JOIN` (kết hợp) đơn giản giữa dữ liệu huấn luyện và bảng đặc trưng này.\n*   **Bước 1.3 Model Training :**\n    Đưa tập dữ liệu đã tích hợp đầy đủ đặc trưng vào huấn luyện với thuật toán **LightGBM Ranker**.\n\n### Bước 2: Model Inference\n*   **Bước 2.1 Loading Testing Data :**\n    Tương tự bước 1, tập kiểm thử được trích xuất bằng logic: `test_df = test_df[test_df['session'] % 10 == 0]` nhằm đảm bảo phân tách tuyệt đối dữ liệu, tránh data leakage.\n*   **Bước 2.2 Generate Candidates :**\n    Áp dụng logic từ lõi thuật toán của Chris Deotte để sinh ra **40 ứng viên tiềm năng nhất** cho mỗi phiên mua sắm. Co-visitation dictionary dùng cho bước này đã được tính toán sẵn và nạp từ `/kaggle/input/otto-covisitation-matrix-parquet-files`. *(Xem giải thích chi tiết logic tại [notebook của Chris Deotte](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575))*\n*   **Bước 2.3 Ranker Model :**\n    *   Ghép danh sách 40 ứng viên với dữ liệu của tập kiểm thử.\n    *   Feature preprocessing bằng luồng xử lý giống hệt như tập huấn luyện.\n    *   Chạy mô hình LGBM để dự đoán và chấm điểm  cho từng ứng viên.\n    *   Sắp xếp điểm số từ cao xuống thấp cho từng phiên (`session`).\n    *   Sử dụng hàm `groupby('session').last(20)` để cắt lấy **Top 20 sản phẩm xuất sắc nhất**.\n*   **Bước 2.4 Export to CSV :**\n    Lưu kết quả cuối cùng ra tệp CSV theo đúng định dạng yêu cầu của hệ thống chấm điểm.\n\n### Bước 3: Model Evaluation\nÁp dụng cơ chế tính điểm Validation nội bộ (CV Score) tương tự như hệ thống đánh giá của Chris Deotte tại: [Compute Validation Score CV](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565)\n\n---\n\n### Credit \nNotebook này được xây dựng dựa trên việc tích hợp các giải pháp tối ưu nhất từ cộng đồng Kaggle. Xin gửi lời cảm ơn tới:\n*   **Chiến lược Validation:** Tham khảo từ notebook của Chris Deotte [tại đây](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565).\n*   **Logic tính toán Ma trận Đồng xuất hiện :** Tham khảo từ notebook của Chris Deotte [tại đây](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575).\n*   **Cốt lõi mã nguồn LGBM Ranker:** Lấy cảm hứng từ cấu trúc PoC (Proof of Concept) của Radek Osmulski [tại đây](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n","metadata":{}},{"cell_type":"markdown","source":"### Về ma trận đồng xuất hiện \nTrong bài toán OTTO, hành vi của người dùng có các mức độ \"cam kết\"  khác nhau. Do đó, ta không thể gộp chung mọi thứ vào một ma trận. \n\nDựa theo thuật toán của Chris Deotte, ta có cơ sở chia ra làm 3 loại ma trận chuyên biệt:\n\n* *Ma trận 1: top_20_clicks (Click $\\rightarrow$ Bất kỳ hành động nào)*\n  \nÝ nghĩa: Trả lời câu hỏi \"Sau khi người dùng xem (click) món A, họ thường tương tác với những món nào tiếp theo?\n\n\"Mục đích: Bắt được luồng duyệt web (Browsing flow). Rất tốt để gợi ý các sản phẩm mang tính chất \"Khám phá\" hoặc \"So sánh\" (Ví dụ: Xem iPhone 13 $\\rightarrow$ Xem iPhone 14).\n* *Ma trận 2: carts_orders (Bất kỳ hành động nào $\\rightarrow$ Cart / Order)*\n\nÝ nghĩa: Trả lời câu hỏi \"Sau khi tương tác với món A, người dùng cuối cùng chốt hạ bỏ vào giỏ hoặc mua món nào?\"\n\nMục đích: Lọc bỏ tín hiệu nhiễu (những món xem cho vui) để tìm ra những món đồ có tỷ lệ chuyển đổi (Conversion Rate) cao.\n\n* *Ma trận 3: buy2buy (Cart / Order $\\rightarrow$ Cart / Order)*\n  \nÝ nghĩa: Trả lời câu hỏi \"Khi người ta quyết định mua món A, họ thường mua kèm món nào?\"Mục đích: Bắt được tính chất Hàng hóa bổ sung (Complementary goods). \n\nVí dụ: Mua Điện thoại $\\rightarrow$ Mua Ốp lưng; Mua Giày $\\rightarrow$ Mua Tất.","metadata":{}},{"cell_type":"code","source":"!pip install polars\n!pip -q install lightgbm pandarallel -U","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:37:59.679446Z","iopub.execute_input":"2026-04-05T04:37:59.679756Z","iopub.status.idle":"2026-04-05T04:38:07.634783Z","shell.execute_reply.started":"2026-04-05T04:37:59.679728Z","shell.execute_reply":"2026-04-05T04:38:07.633271Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os, gc, glob, itertools, warnings\nfrom collections import Counter\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nfrom lightgbm import LGBMRanker\nfrom pandarallel import pandarallel\n\npandarallel.initialize(nb_workers=4, progress_bar=False)\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:07.638269Z","iopub.execute_input":"2026-04-05T04:38:07.638628Z","iopub.status.idle":"2026-04-05T04:38:07.649070Z","shell.execute_reply.started":"2026-04-05T04:38:07.638591Z","shell.execute_reply":"2026-04-05T04:38:07.647711Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"VER = 7\n\nTYPE_LABELS = {'clicks': 0, 'carts': 1, 'orders': 2}\nID2TYPE = {0: 'clicks', 1: 'carts', 2: 'orders'}\n\ntype_weight_multipliers = {0: 0.5, 1: 9, 2: 0.5}\n\n# Candidate cap (generation stage)\nCANDIDATE_CAP = {'clicks': 40, 'carts': 20, 'orders': 20}\n\n# Final output top-k (submission/eval stage)\nFINAL_TOP_K = 20\n\nNEG_SAMPLE_FRAC = {'clicks': 0.35, 'carts': 0.35, 'orders': 0.35}\n\nLGBM_PARAMS = dict(\n    objective=\"lambdarank\",\n    metric=\"ndcg\",\n    boosting_type=\"gbdt\",\n    n_estimators=120,\n    learning_rate=0.08,\n    num_leaves=63,\n    min_data_in_leaf=50,\n    max_depth=-1,\n    importance_type='gain',\n    n_jobs=4\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:07.650988Z","iopub.execute_input":"2026-04-05T04:38:07.651385Z","iopub.status.idle":"2026-04-05T04:38:07.667176Z","shell.execute_reply.started":"2026-04-05T04:38:07.651343Z","shell.execute_reply":"2026-04-05T04:38:07.666280Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2) Utilities\n\nCác hàm hỗ trợ:\n- `reduce_pd_mem`: ép dtype giảm RAM\n- `pqt_to_dict`: load co-vis parquet thành dict `{aid_x: [aid_y,...]}`\n- `load_events`: load split theo session `%10`\n- `load_labels`: load ground truth labels","metadata":{}},{"cell_type":"code","source":"def reduce_pd_mem(df: pd.DataFrame):\n    for c in df.columns:\n        if pd.api.types.is_integer_dtype(df[c]):\n            if df[c].min() >= 0:\n                if df[c].max() < np.iinfo(np.uint8).max:\n                    df[c] = df[c].astype(np.uint8)\n                elif df[c].max() < np.iinfo(np.uint16).max:\n                    df[c] = df[c].astype(np.uint16)\n                elif df[c].max() < np.iinfo(np.uint32).max:\n                    df[c] = df[c].astype(np.uint32)\n                else:\n                    df[c] = df[c].astype(np.uint64)\n            else:\n                if df[c].min() > np.iinfo(np.int8).min and df[c].max() < np.iinfo(np.int8).max:\n                    df[c] = df[c].astype(np.int8)\n                elif df[c].min() > np.iinfo(np.int16).min and df[c].max() < np.iinfo(np.int16).max:\n                    df[c] = df[c].astype(np.int16)\n                elif df[c].min() > np.iinfo(np.int32).min and df[c].max() < np.iinfo(np.int32).max:\n                    df[c] = df[c].astype(np.int32)\n                else:\n                    df[c] = df[c].astype(np.int64)\n        elif pd.api.types.is_float_dtype(df[c]):\n            df[c] = df[c].astype(np.float32)\n    return df\n\ndef pqt_to_dict(path):\n    tmp = pl.read_parquet(path).group_by('aid_x').agg(pl.col('aid_y'))\n    return (\n        tmp.to_pandas()\n           .set_index('aid_x')['aid_y']\n           .apply(lambda x: list(x) if not isinstance(x, list) else x)\n           .to_dict()\n    )\n\ndef load_events(sample_mod_value):\n    dfs = []\n    for f in glob.glob('/kaggle/input/datasets/cdeotte/otto-validation/test_parquet/*'):\n        x = pd.read_parquet(f)\n        x.ts = (x.ts / 1000).astype('int32')\n        x['type'] = x['type'].map(TYPE_LABELS).astype('int8')\n        dfs.append(x)\n    df = pd.concat(dfs, ignore_index=True)\n    df = df[df.session % 10 == sample_mod_value].reset_index(drop=True)\n    return reduce_pd_mem(df)\n\ndef load_labels():\n    y = pd.read_parquet('/kaggle/input/datasets/cdeotte/otto-validation/test_labels.parquet')\n    y['type'] = y['type'].map(TYPE_LABELS).astype('int8')\n    return y","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:07.668704Z","iopub.execute_input":"2026-04-05T04:38:07.668997Z","iopub.status.idle":"2026-04-05T04:38:07.686101Z","shell.execute_reply.started":"2026-04-05T04:38:07.668968Z","shell.execute_reply":"2026-04-05T04:38:07.684932Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3) Load precomputed dense features\n\nLoad 3 bảng feature đã chuẩn bị sẵn từ dataset:\n- global counter\n- user counter\n- time-weighted counter\n\nEmbedding để `OFF` mặc định cho an toàn RAM.","metadata":{}},{"cell_type":"code","source":"aid_global_counter_all_types = pl.read_parquet('/kaggle/input/datasets/alexlods/ottoprecalculatedfeatureparquet/aid_global_counter_all_types.pqt')\naid_global_user_counter_all_types = pl.read_parquet('/kaggle/input/datasets/alexlods/ottoprecalculatedfeatureparquet/aid_global_user_counter_all_types.pqt')\naid_global_user_counter_all_types_time_weighted = pl.read_parquet('/kaggle/input/datasets/alexlods/ottoprecalculatedfeatureparquet/aid_global_user_counter_all_types_time_weighted (1).pqt')\n\nUSE_EMBED = True \nif USE_EMBED:\n    emb = np.load('/kaggle/input/datasets/alexlods/ottoprecalculatedfeatureparquet/word2vec.model.wv.vectors.npy')\n    emb_df = pl.from_pandas(\n        pd.DataFrame(emb, columns=[f'Embedding_{i}' for i in range(emb.shape[1])])\n          .reset_index()\n          .rename(columns={'index': 'aid'})\n    ).with_columns(pl.col('aid').cast(pl.Int32))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:07.688508Z","iopub.execute_input":"2026-04-05T04:38:07.689100Z","iopub.status.idle":"2026-04-05T04:38:08.921949Z","shell.execute_reply.started":"2026-04-05T04:38:07.689066Z","shell.execute_reply":"2026-04-05T04:38:08.921255Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4) Load co-visitation matrices\n\n- `top_20_clicks`\n- `top_20_buys`\n- `top_20_buy2buy`\n","metadata":{}},{"cell_type":"code","source":"DISK_PIECES = 4\n\ntop_20_clicks = pqt_to_dict(f'/kaggle/input/datasets/alexlods/otto-covisitation-matrix-parquet-files/top_20_clicks_v{VER}_0.pqt')\nfor k in range(1, DISK_PIECES):\n    top_20_clicks.update(\n        pqt_to_dict(f'/kaggle/input/datasets/alexlods/otto-covisitation-matrix-parquet-files/top_20_clicks_v{VER}_{k}.pqt')\n    )\n\ntop_20_buys = pqt_to_dict(f'/kaggle/input/datasets/alexlods/otto-covisitation-matrix-parquet-files/top_15_carts_orders_v{VER}_0.pqt')\nfor k in range(1, DISK_PIECES):\n    top_20_buys.update(\n        pqt_to_dict(f'/kaggle/input/datasets/alexlods/otto-covisitation-matrix-parquet-files/top_15_carts_orders_v{VER}_{k}.pqt')\n    )\n\ntop_20_buy2buy = pqt_to_dict(f'/kaggle/input/datasets/alexlods/otto-covisitation-matrix-parquet-files/top_15_buy2buy_v{VER}_0.pqt')\n\nprint(len(top_20_clicks), len(top_20_buys), len(top_20_buy2buy))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:08.923156Z","iopub.execute_input":"2026-04-05T04:38:08.923474Z","iopub.status.idle":"2026-04-05T04:38:37.711895Z","shell.execute_reply.started":"2026-04-05T04:38:08.923427Z","shell.execute_reply":"2026-04-05T04:38:37.711114Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5) Candidate Generation functions\n\nCác hàm generate candidates theo từng target:\n- `suggest_clicks` (đã fix giữ `unique_aids` để không mất user history)\n- `suggest_carts`\n- `suggest_orders`","metadata":{}},{"cell_type":"code","source":"def suggest_clicks(df, top_clicks, cap=40):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n\n    if len(unique_aids) >= cap:\n        weights = np.logspace(0.1, 1, len(aids), base=2, endpoint=True) - 1\n        scores = Counter()\n        for aid, w, t in zip(aids, weights, types):\n            scores[aid] += w * type_weight_multipliers[t]\n        return [k for k, _ in scores.most_common(cap)]\n\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(cap) if aid2 not in unique_aids]\n\n    result = unique_aids + top_aids2[:max(0, cap - len(unique_aids))]\n    if len(result) < cap:\n        result += list(top_clicks)[:cap - len(result)]\n    return result[:cap]\n\ndef suggest_carts(df, top_carts, cap=20):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    dff = df.loc[(df['type'] == 0) | (df['type'] == 1)]\n    unique_buys = list(dict.fromkeys(dff.aid.tolist()[::-1]))\n\n    if len(unique_aids) >= cap:\n        weights = np.logspace(0.5, 1, len(aids), base=2, endpoint=True) - 1\n        scores = Counter()\n        for aid, w, t in zip(aids, weights, types):\n            scores[aid] += w * type_weight_multipliers[t]\n        aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_buys if aid in top_20_buys]))\n        for aid in aids2:\n            scores[aid] += 0.1\n        return [k for k, _ in scores.most_common(cap)]\n\n    aids1 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    aids2 = list(itertools.chain(*[top_20_buys[aid] * 2 for aid in unique_aids if aid in top_20_buys]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids1 + aids2).most_common(cap) if aid2 not in unique_aids]\n    result = unique_aids + top_aids2[:max(0, cap - len(unique_aids))]\n    if len(result) < cap:\n        result += list(top_carts)[:cap - len(result)]\n    return result[:cap]\n\ndef suggest_orders(df, top_orders, cap=20):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    dff = df.loc[(df['type'] == 1) | (df['type'] == 2)]\n    unique_buys = list(dict.fromkeys(dff.aid.tolist()[::-1]))\n\n    if len(unique_aids) >= cap:\n        weights = np.logspace(0.5, 1, len(aids), base=2, endpoint=True) - 1\n        scores = Counter()\n        for aid, w, t in zip(aids, weights, types):\n            scores[aid] += w * type_weight_multipliers[t]\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3:\n            scores[aid] += 0.1\n        return [k for k, _ in scores.most_common(cap)]\n\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2 + aids3).most_common(cap) if aid2 not in unique_aids]\n    result = unique_aids + top_aids2[:max(0, cap - len(unique_aids))]\n    if len(result) < cap:\n        result += list(top_orders)[:cap - len(result)]\n    return result[:cap]\n\ndef build_candidates(df):\n    top_clicks = df.loc[df['type'] == 0, 'aid'].value_counts().index.values[:40]\n    top_carts = df.loc[df['type'] == 1, 'aid'].value_counts().index.values[:20]\n    top_orders = df.loc[df['type'] == 2, 'aid'].value_counts().index.values[:20]\n\n    g = df.sort_values([\"session\", \"ts\"]).groupby([\"session\"])\n\n    pred_clicks = g.parallel_apply(lambda x: suggest_clicks(x, top_clicks, CANDIDATE_CAP['clicks']))\n    pred_carts  = g.parallel_apply(lambda x: suggest_carts(x, top_carts, CANDIDATE_CAP['carts']))\n    pred_orders = g.parallel_apply(lambda x: suggest_orders(x, top_orders, CANDIDATE_CAP['orders']))\n\n    return pred_clicks, pred_carts, pred_orders","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:37.713221Z","iopub.execute_input":"2026-04-05T04:38:37.713588Z","iopub.status.idle":"2026-04-05T04:38:37.783131Z","shell.execute_reply.started":"2026-04-05T04:38:37.713547Z","shell.execute_reply":"2026-04-05T04:38:37.782323Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6) Build training/inference table cho từng target\n\n`make_target_frame()`:\n- Gộp `real actions` + `generated candidates`\n- Dedup theo `(session, aid)` ưu tiên real action\n- Join dense features\n\n`attach_labels()`:\n- Join ground truth label theo target\n\n`downsample_negatives()`:\n- Giảm negatives để tiết kiệm RAM\n\n`get_group_lengths()`:\n- Tạo group lengths cho LGBMRanker","metadata":{}},{"cell_type":"code","source":"def make_target_frame(base_df: pd.DataFrame, pred_series: pd.Series, target_type: str):\n    t = TYPE_LABELS[target_type]\n\n    real_df = base_df[base_df['type'] == t][['session', 'aid', 'ts', 'type']].copy()\n    real_df['real_action'] = 1\n    real_df['CG_ranking'] = 0\n\n    cand = pd.DataFrame(pred_series, columns=['aid']).reset_index().explode('aid')\n    cand['aid'] = cand['aid'].astype(np.int32)\n    cand['session'] = cand['session'].astype(np.int32)\n    cand['ts'] = 0\n    cand['type'] = t\n    cand['real_action'] = 0\n    cand['CG_ranking'] = cand.groupby('session').cumcount().astype(np.int32)\n\n    all_df = pd.concat([real_df, cand], ignore_index=True)\n    all_df = all_df.sort_values(['session', 'aid', 'real_action'], ascending=[True, True, False])\n    all_df = all_df.drop_duplicates(['session', 'aid'], keep='first')\n    all_df = reduce_pd_mem(all_df)\n\n    p = pl.from_pandas(all_df)\n\n    p = (\n        p.join(aid_global_counter_all_types, on='aid', how='left', suffix='_global_counter')\n         .join(aid_global_user_counter_all_types, on='aid', how='left', suffix='_user_counter')\n         .join(aid_global_user_counter_all_types_time_weighted, on='aid', how='left', suffix='_timed_global_counter')\n    )\n\n    if USE_EMBED:\n        p = p.join(emb_df, on='aid', how='left')\n\n    return p\n\ndef attach_labels(train_pl: pl.DataFrame, labels_pd: pd.DataFrame, target_type: str):\n    t = TYPE_LABELS[target_type]\n    y = labels_pd[labels_pd['type'] == t][['session', 'ground_truth']].copy()\n    y = y.explode('ground_truth')\n    y['aid'] = y['ground_truth'].astype(np.int32)\n    y['label'] = 1\n    y = y[['session', 'aid', 'label']]\n    y = reduce_pd_mem(y)\n\n    ypl = pl.from_pandas(y)\n    out = train_pl.join(ypl, on=['session', 'aid'], how='left').with_columns(\n        pl.col('label').fill_null(0).cast(pl.Int8)\n    )\n    return out\n\ndef downsample_negatives(df_pl: pl.DataFrame, frac=0.35):\n    pos = df_pl.filter(pl.col('label') == 1)\n    neg = df_pl.filter(pl.col('label') == 0).sample(fraction=frac, with_replacement=False, seed=42)\n    out = pl.concat([pos, neg], how='vertical').sort(['session'])\n    return out\n\ndef get_group_lengths(df_pl: pl.DataFrame):\n    return (\n        df_pl.group_by('session', maintain_order=True)\n             .agg(pl.len().alias('n'))\n             .select('n')\n             .to_series()\n             .to_numpy()\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:37.784412Z","iopub.execute_input":"2026-04-05T04:38:37.784763Z","iopub.status.idle":"2026-04-05T04:38:37.807980Z","shell.execute_reply.started":"2026-04-05T04:38:37.784726Z","shell.execute_reply":"2026-04-05T04:38:37.807040Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7) Train + Predict cho từng target\n\n`train_and_predict_target()`:\n- Build train frame\n- Attach labels + downsample\n- Train LGBMRanker\n- Free memory\n- Build valid frame và predict\n- Lấy **top 20** cuối cùng cho mỗi session","metadata":{}},{"cell_type":"code","source":"def train_and_predict_target(train_events, valid_events, labels_pd, pred_train, pred_valid, target_type):\n    print(f\"\\n===== TARGET: {target_type} =====\")\n\n    tr = make_target_frame(train_events, pred_train, target_type)\n    tr = attach_labels(tr, labels_pd, target_type)\n    tr = downsample_negatives(tr, frac=NEG_SAMPLE_FRAC[target_type])\n\n    feature_cols = [c for c in tr.columns if c not in ['session', 'aid', 'label']]\n    X_train = tr.select(feature_cols).to_pandas()\n    y_train = tr.select('label').to_pandas().values.ravel()\n    grp_train = get_group_lengths(tr)\n    X_train = reduce_pd_mem(X_train)\n\n    model = LGBMRanker(**LGBM_PARAMS)\n    model.fit(X_train, y_train, group=grp_train)\n\n    del tr, X_train, y_train, grp_train\n    gc.collect()\n\n    va = make_target_frame(valid_events, pred_valid, target_type)\n    X_valid = va.select(feature_cols).to_pandas()\n    X_valid = reduce_pd_mem(X_valid)\n\n    scores = model.predict(X_valid)\n    va = va.with_columns(pl.Series('score', scores.astype(np.float32)))\n\n    out = (\n    va.sort(['session', 'score'], descending=[False, True])\n      .group_by('session', maintain_order=True)\n      .agg(pl.col('aid').head(FINAL_TOP_K).alias('labels'))\n      .to_pandas()\n    )\n    out['session_type'] = out['session'].astype(str) + f'_{target_type}'\n    out = out[['session_type', 'labels']]\n\n    del va, X_valid, scores, feature_cols, model\n    gc.collect()\n\n    return out","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:38:37.809123Z","iopub.execute_input":"2026-04-05T04:38:37.809996Z","iopub.status.idle":"2026-04-05T04:38:37.829621Z","shell.execute_reply.started":"2026-04-05T04:38:37.809962Z","shell.execute_reply":"2026-04-05T04:38:37.828653Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8) Hàm đánh giá Recall local\n\nTính Recall theo trọng số official:\n- clicks 0.10\n- carts 0.30\n- orders 0.60","metadata":{}},{"cell_type":"code","source":"import ast\nimport re\ndef _parse_labels_safe(x):\n    # 1) đã là list/array\n    if isinstance(x, list):\n        return [int(i) for i in x]\n    if isinstance(x, np.ndarray):\n        return [int(i) for i in x.tolist()]\n\n    # 2) null\n    if pd.isna(x):\n        return []\n\n    s = str(x).strip()\n\n    # 3) dạng \"[1, 2, 3]\"\n    if s.startswith('[') and s.endswith(']'):\n        try:\n            v = ast.literal_eval(s)\n            if isinstance(v, (list, tuple, np.ndarray)):\n                return [int(i) for i in v]\n        except Exception:\n            pass\n\n    # 4) dạng \"1 2 3\" hoặc \"1,2,3\"\n    parts = re.split(r'[\\s,]+', s.strip('[]'))\n    out = []\n    for p in parts:\n        if p == '':\n            continue\n        out.append(int(p))\n    return out\n\ndef eval_recall(pred_df, labels_path='/kaggle/input/datasets/cdeotte/otto-validation/test_labels.parquet'):\n    test_labels = pd.read_parquet(labels_path)\n    score = 0.0\n    weights = {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}\n\n    for t in ['clicks', 'carts', 'orders']:\n        sub = pred_df[pred_df.session_type.str.contains(t)].copy()\n        sub['session'] = sub.session_type.apply(lambda x: int(str(x).split('_')[0]))\n        sub['labels'] = sub['labels'].apply(_parse_labels_safe)\n\n        gt = test_labels[test_labels['type'] == t].merge(\n            sub[['session', 'labels']], on='session', how='left'\n        ).dropna(subset=['labels'])\n\n        gt['hits'] = gt.apply(lambda r: len(set(r['ground_truth']).intersection(set(r['labels']))), axis=1)\n        gt['gt_count'] = gt['ground_truth'].str.len().clip(0, 20)\n\n        recall = gt['hits'].sum() / gt['gt_count'].sum()\n        score += weights[t] * recall\n        print(f'{t} recall = {recall:.6f}')\n\n    print('==============')\n    print(f'Overall Recall = {score:.6f}')\n    print('==============')\n    return score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T05:13:40.995860Z","iopub.execute_input":"2026-04-05T05:13:40.996226Z","iopub.status.idle":"2026-04-05T05:13:41.008184Z","shell.execute_reply.started":"2026-04-05T05:13:40.996169Z","shell.execute_reply":"2026-04-05T05:13:41.007103Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9) Chạy pipeline end-to-end\n\n- Load split train/valid\n- Build candidates cho từng split\n- Train 3 rankers riêng\n- Gộp kết quả, export CSV\n- Evaluate local recall","metadata":{}},{"cell_type":"code","source":"# Load data splits\ntrain_events = load_events(1)   # train-like split\nvalid_events = load_events(0)   # valid-like split\nlabels_pd = load_labels()\n\nprint(\"train_events:\", train_events.shape)\nprint(\"valid_events:\", valid_events.shape)\n\n# Candidate generation\npred_train_clicks, pred_train_carts, pred_train_orders = build_candidates(train_events)\npred_valid_clicks, pred_valid_carts, pred_valid_orders = build_candidates(valid_events)\n\n# Train/predict clicks\nclicks_pred_df = train_and_predict_target(\n    train_events, valid_events, labels_pd,\n    pred_train_clicks, pred_valid_clicks, 'clicks'\n)\ngc.collect()\n\n# Train/predict carts\ncarts_pred_df = train_and_predict_target(\n    train_events, valid_events, labels_pd,\n    pred_train_carts, pred_valid_carts, 'carts'\n)\ngc.collect()\n\n# Train/predict orders\norders_pred_df = train_and_predict_target(\n    train_events, valid_events, labels_pd,\n    pred_train_orders, pred_valid_orders, 'orders'\n)\ngc.collect()\n\n# Merge output\npred_df_list = pd.concat([clicks_pred_df, carts_pred_df, orders_pred_df], ignore_index=True)\n\n# Save submission-like file\npred_df_save = pred_df_list.copy()\npred_df_save['labels'] = pred_df_save['labels'].apply(lambda x: ' '.join(map(str, x)))\npred_df_save.to_csv('validation_preds.csv', index=False)\n\nprint(\"Saved validation_preds.csv:\", pred_df_save.shape)\npred_df_save.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T04:53:33.458133Z","iopub.execute_input":"2026-04-05T04:53:33.460547Z","iopub.status.idle":"2026-04-05T05:06:51.106937Z","shell.execute_reply.started":"2026-04-05T04:53:33.460489Z","shell.execute_reply":"2026-04-05T05:06:51.105619Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_ = eval_recall(pred_df_list)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-05T05:13:54.144339Z","iopub.execute_input":"2026-04-05T05:13:54.144715Z","iopub.status.idle":"2026-04-05T05:14:02.980122Z","shell.execute_reply.started":"2026-04-05T05:13:54.144682Z","shell.execute_reply":"2026-04-05T05:14:02.978942Z"}},"outputs":[],"execution_count":null}]}