{
  "id": 379631,
  "title": "Training GBDT Ranker based on Chris's discussion",
  "url": "/competitions/otto-recommender-system/discussion/379631",
  "author_name": "",
  "post_date": "2023-01-20T12:06:08.166055100Z",
  "votes": 23,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Based on this self-explanatory and awesome <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">discussion</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I would like to share the complete code for training ranker models in this competition. I just put everything together that Chris have mentioned in the discussion. </p>\n<p>The full code for the ranking is <a href=\"https://www.kaggle.com/code/snnclsr/gbdt-ranking\" target=\"_blank\">here</a>.</p>\n<pre><code> polars  pl\n\ntrain = pl.read_parquet()\nvalid = pl.read_parquet()\nvalid_labels = pl.read_parquet()\n</code></pre>\n<h4>Generating the item features</h4>\n<pre><code>\nitem_features = pl.concat([train, valid]).groupby().agg([\n    pl.count().alias(), \n    pl.n_unique().alias(), \n    pl.mean().alias().cast(pl.Float32)\n])\nitem_features.write_parquet()\n</code></pre>\n<h4>Generating the user features</h4>\n<pre><code>\nuser_features = valid.groupby().agg([\n    pl.count().alias(),\n    pl.n_unique().alias(),\n    pl.mean().alias().cast(pl.Float32)\n])\nuser_features.write_parquet()\n</code></pre>\n<p>Candidates generated from this <a href=\"https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\" target=\"_blank\">notebook</a>. I just extracted the aids from the co-occurence matrices. </p>\n<pre><code>event_paths = {\n    : ,\n    : ,\n    : \n}\n</code></pre>\n<h4>Training the Ranker</h4>\n<p>As Chris mentioned, we are also training XGBoost. The changes I made here are:</p>\n<ul>\n<li>group parameter: Since we have more than 50 candidates generated by the co-occurrence matrices, we update the group parameter as well to reflect that. We tell xgboost how many aids we have in each user (session) to rank. </li>\n<li>New hyperparameters</li>\n</ul>\n<pre><code> xgboost  xgb\n sklearn.model_selection  GroupKFold\n\n ():\n\n    skf = GroupKFold(n_splits=n_splits)\n    FEATURES = [\n        , , , , , , \n    ]\n    TARGET = \n     fold,(train_idx, valid_idx)  (skf.split(df_cands, df_cands[], groups=df_cands[])):\n\n        X_train = df_cands.loc[train_idx, FEATURES]\n        y_train = df_cands.loc[train_idx, TARGET]\n        X_valid = df_cands.loc[valid_idx, FEATURES]\n        y_valid = df_cands.loc[valid_idx, TARGET]\n\n        X_train = X_train.sort_values().reset_index(drop=)\n        X_valid = X_valid.sort_values().reset_index(drop=)\n\n        train_group = X_train.groupby().session.agg().values\n        valid_group = X_valid.groupby().session.agg().values\n\n        X_train = X_train.drop([], axis=)\n        X_valid = X_valid.drop([], axis=)\n\n        dtrain = xgb.DMatrix(X_train, y_train, group=train_group) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=valid_group) \n        xgb_parms = {\n            :, \n            :,\n            : , \n            : ,\n            : , \n            : , \n            : ,\n            : ,\n            \n        }\n        model = xgb.train(\n            xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,), (dvalid,)],\n            num_boost_round=,\n            verbose_eval=\n        )\n        model.save_model()\n        gc.collect()\n</code></pre>\n<p>Lastly we merge the features with the candidate dataframe. It's important that the sessions are sorted otherwise grouping logic above does not work.</p>\n<pre><code>NEGATIVE_FRAC = \n event, path  event_paths.items():\n    ()\n    \n    df_cands = pl.read_parquet(path)\n    df_cands = df_cands.explode().with_columns([\n        pl.col().cast(pl.Int32),\n        pl.col().cast(pl.Int32).alias()\n    ]).drop().unique(subset=[, ])\n    \n    df_cands = df_cands.join(item_features, on=, how=).fill_nan(-)\n    \n    df_cands = df_cands.join(user_features, on=, how=).fill_nan(-)\n    cand_labels = valid_labels.(valid_labels[] == event).explode().with_columns([\n        pl.col().cast(pl.Int32),\n        pl.col().cast(pl.Int32)\n    ]).rename({: })\n    cand_labels = cand_labels.with_column(pl.lit().alias().cast(pl.Int8)).drop()\n    \n    df_cands = df_cands.join(cand_labels, on=[, ], how=).fill_null()\n    \n    df_cands = pl.concat([\n        df_cands.(df_cands[] == ).sample(frac=NEGATIVE_FRAC, seed=),\n        df_cands.(df_cands[] == )\n    ])\n    (df_cands.groupby().agg(pl.count()))\n    df_cands = df_cands.to_pandas()\n    df_cands = df_cands.sort_values().reset_index(drop=)\n    ()\n\n    train_ranker(event, df_cands)\n     df_cands, cand_labels\n    gc.collect()\n</code></pre>",
  "messages": [
    {
      "id": "2108300",
      "postDate": "01/20/2023 12:06:08",
      "content": "<p>Hi everyone,</p>\n<p>Based on this self-explanatory and awesome <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">discussion</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I would like to share the complete code for training ranker models in this competition. I just put everything together that Chris have mentioned in the discussion. </p>\n<p>The full code for the ranking is <a href=\"https://www.kaggle.com/code/snnclsr/gbdt-ranking\" target=\"_blank\">here</a>.</p>\n<pre><code> polars  pl\n\ntrain = pl.read_parquet()\nvalid = pl.read_parquet()\nvalid_labels = pl.read_parquet()\n</code></pre>\n<h4>Generating the item features</h4>\n<pre><code>\nitem_features = pl.concat([train, valid]).groupby().agg([\n    pl.count().alias(), \n    pl.n_unique().alias(), \n    pl.mean().alias().cast(pl.Float32)\n])\nitem_features.write_parquet()\n</code></pre>\n<h4>Generating the user features</h4>\n<pre><code>\nuser_features = valid.groupby().agg([\n    pl.count().alias(),\n    pl.n_unique().alias(),\n    pl.mean().alias().cast(pl.Float32)\n])\nuser_features.write_parquet()\n</code></pre>\n<p>Candidates generated from this <a href=\"https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\" target=\"_blank\">notebook</a>. I just extracted the aids from the co-occurence matrices. </p>\n<pre><code>event_paths = {\n    : ,\n    : ,\n    : \n}\n</code></pre>\n<h4>Training the Ranker</h4>\n<p>As Chris mentioned, we are also training XGBoost. The changes I made here are:</p>\n<ul>\n<li>group parameter: Since we have more than 50 candidates generated by the co-occurrence matrices, we update the group parameter as well to reflect that. We tell xgboost how many aids we have in each user (session) to rank. </li>\n<li>New hyperparameters</li>\n</ul>\n<pre><code> xgboost  xgb\n sklearn.model_selection  GroupKFold\n\n ():\n\n    skf = GroupKFold(n_splits=n_splits)\n    FEATURES = [\n        , , , , , , \n    ]\n    TARGET = \n     fold,(train_idx, valid_idx)  (skf.split(df_cands, df_cands[], groups=df_cands[])):\n\n        X_train = df_cands.loc[train_idx, FEATURES]\n        y_train = df_cands.loc[train_idx, TARGET]\n        X_valid = df_cands.loc[valid_idx, FEATURES]\n        y_valid = df_cands.loc[valid_idx, TARGET]\n\n        X_train = X_train.sort_values().reset_index(drop=)\n        X_valid = X_valid.sort_values().reset_index(drop=)\n\n        train_group = X_train.groupby().session.agg().values\n        valid_group = X_valid.groupby().session.agg().values\n\n        X_train = X_train.drop([], axis=)\n        X_valid = X_valid.drop([], axis=)\n\n        dtrain = xgb.DMatrix(X_train, y_train, group=train_group) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=valid_group) \n        xgb_parms = {\n            :, \n            :,\n            : , \n            : ,\n            : , \n            : , \n            : ,\n            : ,\n            \n        }\n        model = xgb.train(\n            xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,), (dvalid,)],\n            num_boost_round=,\n            verbose_eval=\n        )\n        model.save_model()\n        gc.collect()\n</code></pre>\n<p>Lastly we merge the features with the candidate dataframe. It's important that the sessions are sorted otherwise grouping logic above does not work.</p>\n<pre><code>NEGATIVE_FRAC = \n event, path  event_paths.items():\n    ()\n    \n    df_cands = pl.read_parquet(path)\n    df_cands = df_cands.explode().with_columns([\n        pl.col().cast(pl.Int32),\n        pl.col().cast(pl.Int32).alias()\n    ]).drop().unique(subset=[, ])\n    \n    df_cands = df_cands.join(item_features, on=, how=).fill_nan(-)\n    \n    df_cands = df_cands.join(user_features, on=, how=).fill_nan(-)\n    cand_labels = valid_labels.(valid_labels[] == event).explode().with_columns([\n        pl.col().cast(pl.Int32),\n        pl.col().cast(pl.Int32)\n    ]).rename({: })\n    cand_labels = cand_labels.with_column(pl.lit().alias().cast(pl.Int8)).drop()\n    \n    df_cands = df_cands.join(cand_labels, on=[, ], how=).fill_null()\n    \n    df_cands = pl.concat([\n        df_cands.(df_cands[] == ).sample(frac=NEGATIVE_FRAC, seed=),\n        df_cands.(df_cands[] == )\n    ])\n    (df_cands.groupby().agg(pl.count()))\n    df_cands = df_cands.to_pandas()\n    df_cands = df_cands.sort_values().reset_index(drop=)\n    ()\n\n    train_ranker(event, df_cands)\n     df_cands, cand_labels\n    gc.collect()\n</code></pre>",
      "rawMarkdown": "Hi everyone,\n\nBased on this self-explanatory and awesome [discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) by [@cdeotte](https://www.kaggle.com/cdeotte), I would like to share the complete code for training ranker models in this competition. I just put everything together that Chris have mentioned in the discussion. \n\nThe full code for the ranking is [here](https://www.kaggle.com/code/snnclsr/gbdt-ranking).\n\n```python\nimport polars as pl\n# we are reading Radek's validation files.\ntrain = pl.read_parquet(\"../input/otto-train-and-test-data-for-local-validation/train.parquet\")\nvalid = pl.read_parquet(\"../input/otto-train-and-test-data-for-local-validation/test.parquet\")\nvalid_labels = pl.read_parquet(\"../input/otto-train-and-test-data-for-local-validation/test_labels.parquet\")\n```\n\n#### Generating the item features\n```python\n\n# {'aid':'count', 'session':'nunique', 'type': 'mean'}\nitem_features = pl.concat([train, valid]).groupby('aid').agg([\n    pl.count(\"aid\").alias(\"item_item_count\"), \n    pl.n_unique(\"session\").alias(\"item_user_count\"), \n    pl.mean(\"type\").alias(\"item_buy_ratio\").cast(pl.Float32)\n])\nitem_features.write_parquet('item_features.parquet')\n```\n#### Generating the user features\n```python\n# {'session':'count','aid':'nunique','type':'mean'}\nuser_features = valid.groupby('session').agg([\n    pl.count(\"session\").alias(\"user_user_count\"),\n    pl.n_unique(\"aid\").alias(\"user_item_count\"),\n    pl.mean(\"type\").alias(\"user_buy_ratio\").cast(pl.Float32)\n])\nuser_features.write_parquet('user_features.parquet')\n```\nCandidates generated from this [notebook](https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577). I just extracted the aids from the co-occurence matrices. \n```python\nevent_paths = {\n    \"clicks\": \"../input/otto-577-validation-candidates/valid_click_candidates.parquet\",\n    \"carts\": \"../input/otto-577-validation-candidates/valid_carts_candidates.parquet\",\n    \"buys\": \"../input/otto-577-validation-candidates/valid_buys_candidates.parquet\"\n}\n```\n#### Training the Ranker\n\nAs Chris mentioned, we are also training XGBoost. The changes I made here are:\n- group parameter: Since we have more than 50 candidates generated by the co-occurrence matrices, we update the group parameter as well to reflect that. We tell xgboost how many aids we have in each user (session) to rank. \n- New hyperparameters\n\n```python\nimport xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\ndef train_ranker(event, df_cands, n_splits=5):\n    \n    skf = GroupKFold(n_splits=n_splits)\n    FEATURES = [\n        'session', 'item_item_count', 'item_user_count', 'item_buy_ratio', 'user_user_count', 'user_item_count', 'user_buy_ratio'\n    ]\n    TARGET = \"target\"\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(df_cands, df_cands['target'], groups=df_cands['session'])):\n\n        X_train = df_cands.loc[train_idx, FEATURES]\n        y_train = df_cands.loc[train_idx, TARGET]\n        X_valid = df_cands.loc[valid_idx, FEATURES]\n        y_valid = df_cands.loc[valid_idx, TARGET]\n\n        X_train = X_train.sort_values(\"session\").reset_index(drop=True)\n        X_valid = X_valid.sort_values(\"session\").reset_index(drop=True)\n\n        train_group = X_train.groupby('session').session.agg('count').values\n        valid_group = X_valid.groupby('session').session.agg('count').values\n\n        X_train = X_train.drop([\"session\"], axis=1)\n        X_valid = X_valid.drop([\"session\"], axis=1)\n\n        dtrain = xgb.DMatrix(X_train, y_train, group=train_group) # [50] * (len(train_idx)//50) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=valid_group) # [50] * (len(valid_idx)//50)\n        xgb_parms = {\n            'objective':'rank:pairwise', \n            'tree_method':'hist',\n            'random_state': 42, \n            'learning_rate': 0.1,\n            \"colsample_bytree\": 0.8, \n            'eta': 0.05, \n            'max_depth': 6,\n            'subsample': 0.75,\n            # n_estimators=110,\n        }\n        model = xgb.train(\n            xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'), (dvalid,'valid')],\n            num_boost_round=100,\n            verbose_eval=20\n        )\n        model.save_model(f'XGB_fold{fold}_{event}.xgb')\n        gc.collect()\n\n```\nLastly we merge the features with the candidate dataframe. It's important that the sessions are sorted otherwise grouping logic above does not work.\n```python\nNEGATIVE_FRAC = 0.15\nfor event, path in event_paths.items():\n    print(f\"Started ranking model for: {event}\")\n    # Reading the candidates\n    df_cands = pl.read_parquet(path)\n    df_cands = df_cands.explode(\"candidates\").with_columns([\n        pl.col(\"session\").cast(pl.Int32),\n        pl.col(\"candidates\").cast(pl.Int32).alias(\"aid\")\n    ]).drop(\"candidates\").unique(subset=[\"session\", \"aid\"])\n    # Joining the item features\n    df_cands = df_cands.join(item_features, on='aid', how='left').fill_nan(-1)\n    # Joining the user features\n    df_cands = df_cands.join(user_features, on='session', how='left').fill_nan(-1)\n    cand_labels = valid_labels.filter(valid_labels[\"type\"] == event).explode(\"ground_truth\").with_columns([\n        pl.col(\"session\").cast(pl.Int32),\n        pl.col(\"ground_truth\").cast(pl.Int32)# .alias(\"aid\")\n    ]).rename({\"ground_truth\": \"aid\"})\n    cand_labels = cand_labels.with_column(pl.lit(1).alias(\"target\").cast(pl.Int8)).drop(\"type\")\n    # Joining the labels\n    df_cands = df_cands.join(cand_labels, on=[\"session\", \"aid\"], how=\"left\").fill_null(0)\n    # Negative sampling\n    df_cands = pl.concat([\n        df_cands.filter(df_cands[\"target\"] == 0).sample(frac=NEGATIVE_FRAC, seed=42),\n        df_cands.filter(df_cands[\"target\"] == 1)\n    ])\n    print(df_cands.groupby(\"target\").agg(pl.count()))\n    df_cands = df_cands.to_pandas()\n    df_cands = df_cands.sort_values(\"session\").reset_index(drop=True)\n    print(f\"Event: {event} - started training...\")\n\n    train_ranker(event, df_cands)\n    del df_cands, cand_labels\n    gc.collect()\n```",
      "votes": null
    },
    {
      "id": "2108580",
      "postDate": "01/20/2023 16:26:25",
      "content": "<p>Great ! however I don't understand the goal of doing<code>pl.mean(\"type\")</code>. The column is categorical, even if you can average it, you have more than 2 modalities, so It doesn't make sense</p>",
      "rawMarkdown": "Great ! however I don't understand the goal of doing` pl.mean(\"type\")`. The column is categorical, even if you can average it, you have more than 2 modalities, so It doesn't make sense",
      "votes": null
    },
    {
      "id": "2108598",
      "postDate": "01/20/2023 16:53:55",
      "content": "<p>Thank you for sharing.<br>\nHow do you infer from this code?</p>",
      "rawMarkdown": "Thank you for sharing.\nHow do you infer from this code?",
      "votes": null
    },
    {
      "id": "2108619",
      "postDate": "01/20/2023 17:08:15",
      "content": "<p>don't underestimate the mean :)<br>\nIn most notebooks type is encoded in this way:<code>{'click':0, 'cart':1, 'order':2}</code>.<br>\nSo  type is more ordinal than categorical.</p>",
      "rawMarkdown": "don't underestimate the mean :)\nIn most notebooks type is encoded in this way:`{'click':0, 'cart':1, 'order':2}`.\nSo  type is more ordinal than categorical.",
      "votes": null
    },
    {
      "id": "2108626",
      "postDate": "01/20/2023 17:15:33",
      "content": "<p>I totally agree ! I didn't see it from this point of view. Thanks :)</p>",
      "rawMarkdown": "I totally agree ! I didn't see it from this point of view. Thanks :)",
      "votes": null
    },
    {
      "id": "2108764",
      "postDate": "01/20/2023 19:28:54",
      "content": "<p>great notebook. just one point: where there is \"buys\" in your code it should be \"orders\" otherwise we dont have any positives for this category.</p>",
      "rawMarkdown": "great notebook. just one point: where there is \"buys\" in your code it should be \"orders\" otherwise we dont have any positives for this category.",
      "votes": null
    },
    {
      "id": "2108797",
      "postDate": "01/20/2023 20:21:13",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2108980",
      "postDate": "01/21/2023 02:19:53",
      "content": "<p>Thank you for sharing!!<br>\nIn my experience, splitting training data and validation data after negative sampling resulted in abnormally high CV.</p>",
      "rawMarkdown": "Thank you for sharing!!\nIn my experience, splitting training data and validation data after negative sampling resulted in abnormally high CV.",
      "votes": null
    },
    {
      "id": "2110492",
      "postDate": "01/22/2023 08:29:16",
      "content": "<p>Thank you for your sharing great notebook.<br>\nYou mentioned that \"Candidates generated from this notebook. I just extracted the aids from the co-occurence matrices\".<br>\nCan I ask how you made \"valid_click_candidates.parquet\"?<br>\nThis file was made by pred_df in the public notebook you mentioned?</p>",
      "rawMarkdown": "Thank you for your sharing great notebook.\nYou mentioned that \"Candidates generated from this notebook. I just extracted the aids from the co-occurence matrices\".\nCan I ask how you made \"valid_click_candidates.parquet\"?\nThis file was made by pred_df in the public notebook you mentioned?",
      "votes": null
    },
    {
      "id": "2110548",
      "postDate": "01/22/2023 09:11:47",
      "content": "<p>Considerate！！</p>",
      "rawMarkdown": "Considerate！！",
      "votes": null
    },
    {
      "id": "2111811",
      "postDate": "01/23/2023 08:05:33",
      "content": "<p>Interesting, what was your negative sampling ratio? and how many zeros and ones do you have?</p>",
      "rawMarkdown": "Interesting, what was your negative sampling ratio? and how many zeros and ones do you have?",
      "votes": null
    },
    {
      "id": "2111813",
      "postDate": "01/23/2023 08:10:32",
      "content": "<p>It's not but what I did is:</p>\n<pre><code>...\naids2 = (itertools.chain(*[top_20_clicks[aid]  aid  unique_aids  aid  top_20_clicks]))\n\ntop_aids2 = [aid2  aid2, cnt  Counter(aids2).most_common()  aid2   unique_aids]\n...\n</code></pre>\n<p>this is the code from the original notebook for suggesting clicks. I returned the aids2 list from this function and took the unique values. I did the same for carts and buys as well.</p>",
      "rawMarkdown": "It's not but what I did is:\n\n```python\n...\naids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n# RERANK CANDIDATES\ntop_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]\n...\n```\nthis is the code from the original notebook for suggesting clicks. I returned the aids2 list from this function and took the unique values. I did the same for carts and buys as well.",
      "votes": null
    },
    {
      "id": "2111816",
      "postDate": "01/23/2023 08:12:17",
      "content": "<p>One important side note is that I generated those candidates also using Radek's validation files.</p>",
      "rawMarkdown": "One important side note is that I generated those candidates also using Radek's validation files.",
      "votes": null
    },
    {
      "id": "2112915",
      "postDate": "01/24/2023 00:08:45",
      "content": "<p>How would you do it ?</p>",
      "rawMarkdown": "How would you do it ?",
      "votes": null
    },
    {
      "id": "2114285",
      "postDate": "01/24/2023 20:37:50",
      "content": "<p>It worked! Thanks!</p>",
      "rawMarkdown": "It worked! Thanks!",
      "votes": null
    },
    {
      "id": "2118489",
      "postDate": "01/28/2023 03:40:05",
      "content": "<p>hi, thanks for share this notebook👍<br>\nyet I wonder if this part is necessary here? </p>\n<pre><code>X_train = X_train.sort_values(\"session\").reset_index(drop=True)\nX_valid = X_valid.sort_values(\"session\").reset_index(drop=True)\n</code></pre>\n<p>The index already keeps all session together. Correct me if I'm wrong :)</p>",
      "rawMarkdown": "hi, thanks for share this notebook👍\nyet I wonder if this part is necessary here? \n```\nX_train = X_train.sort_values(\"session\").reset_index(drop=True)\nX_valid = X_valid.sort_values(\"session\").reset_index(drop=True)\n```\nThe index already keeps all session together. Correct me if I'm wrong :)",
      "votes": null
    },
    {
      "id": "2118827",
      "postDate": "01/28/2023 10:00:54",
      "content": "<p>Yes since we do it here as well:</p>\n<pre><code>df_cands = df_cands.sort_values().reset_index(drop=)\n</code></pre>\n<p>it becomes unnecessary. I didn't remove it since the first edition because I didn't had this part earlier then it got stuck 🙃</p>",
      "rawMarkdown": "Yes since we do it here as well:\n\n```python\ndf_cands = df_cands.sort_values(\"session\").reset_index(drop=True)\n```\nit becomes unnecessary. I didn't remove it since the first edition because I didn't had this part earlier then it got stuck 🙃",
      "votes": null
    },
    {
      "id": "2122814",
      "postDate": "01/31/2023 06:36:31",
      "content": "<p>Thx for sharing! </p>\n<p>yet, I hope to know how to infer! </p>",
      "rawMarkdown": "Thx for sharing! \n\nyet, I hope to know how to infer!",
      "votes": null
    },
    {
      "id": "2123212",
      "postDate": "01/31/2023 11:42:26",
      "content": "<p>Following Chris's implementation - after generating the candidates for the inference, we make predictions with each of these models (5 for clicks, 5 for buys, etc.) and take the average. Then we sort the scores in descending order (ranking). Lastly, we choose the final 20. Here is the example for clicks candidates: </p>\n<pre><code>preds = np.zeros((df_test_cand_clicks))\n fold  ():\n    model = xgb.Booster()\n    model.load_model()\n    dtest = xgb.DMatrix(data=df_test_cand_clicks[FEATURES].drop(, axis=))\n    preds += model.predict(dtest)/\npredictions = df_test_cand_clicks[[, ]].copy()\npredictions[] = preds\n\npredictions = predictions.sort_values([, ], ascending=[,]).reset_index(drop=)\npredictions[] = predictions.groupby().aid.cumcount().astype()\npredictions = predictions.loc[predictions.n&lt;]\n</code></pre>",
      "rawMarkdown": "Following Chris's implementation - after generating the candidates for the inference, we make predictions with each of these models (5 for clicks, 5 for buys, etc.) and take the average. Then we sort the scores in descending order (ranking). Lastly, we choose the final 20. Here is the example for clicks candidates: \n```python\npreds = np.zeros(len(df_test_cand_clicks))\nfor fold in range(5):\n    model = xgb.Booster()\n    model.load_model(f'XGB_fold{fold}_click.xgb')\n    dtest = xgb.DMatrix(data=df_test_cand_clicks[FEATURES].drop(\"session\", axis=\"columns\"))\n    preds += model.predict(dtest)/5\npredictions = df_test_cand_clicks[['session', 'aid']].copy()\npredictions['pred'] = preds\n\npredictions = predictions.sort_values(['session', 'aid'], ascending=[True,False]).reset_index(drop=True)\npredictions['n'] = predictions.groupby('session').aid.cumcount().astype('int8')\npredictions = predictions.loc[predictions.n<20]\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2108580,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/20/2023 16:26:25",
      "content": "<p>Great ! however I don't understand the goal of doing<code>pl.mean(\"type\")</code>. The column is categorical, even if you can average it, you have more than 2 modalities, so It doesn't make sense</p>",
      "votes": null,
      "replies": [
        {
          "id": 2108619,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "01/20/2023 17:08:15",
          "content": "<p>don't underestimate the mean :)<br>\nIn most notebooks type is encoded in this way:<code>{'click':0, 'cart':1, 'order':2}</code>.<br>\nSo  type is more ordinal than categorical.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2108626,
              "author_name": "rayanaay",
              "author_url": "",
              "post_date": "01/20/2023 17:15:33",
              "content": "<p>I totally agree ! I didn't see it from this point of view. Thanks :)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2108598,
      "author_name": "ayaanjang",
      "author_url": "",
      "post_date": "01/20/2023 16:53:55",
      "content": "<p>Thank you for sharing.<br>\nHow do you infer from this code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2108764,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "01/20/2023 19:28:54",
      "content": "<p>great notebook. just one point: where there is \"buys\" in your code it should be \"orders\" otherwise we dont have any positives for this category.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2108797,
      "author_name": "mtrofficus",
      "author_url": "",
      "post_date": "01/20/2023 20:21:13",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2108980,
      "author_name": "yurimaeda",
      "author_url": "",
      "post_date": "01/21/2023 02:19:53",
      "content": "<p>Thank you for sharing!!<br>\nIn my experience, splitting training data and validation data after negative sampling resulted in abnormally high CV.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2111811,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/23/2023 08:05:33",
          "content": "<p>Interesting, what was your negative sampling ratio? and how many zeros and ones do you have?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2112915,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/24/2023 00:08:45",
          "content": "<p>How would you do it ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2110492,
      "author_name": "aesoptacit",
      "author_url": "",
      "post_date": "01/22/2023 08:29:16",
      "content": "<p>Thank you for your sharing great notebook.<br>\nYou mentioned that \"Candidates generated from this notebook. I just extracted the aids from the co-occurence matrices\".<br>\nCan I ask how you made \"valid_click_candidates.parquet\"?<br>\nThis file was made by pred_df in the public notebook you mentioned?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2111813,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/23/2023 08:10:32",
          "content": "<p>It's not but what I did is:</p>\n<pre><code>...\naids2 = (itertools.chain(*[top_20_clicks[aid]  aid  unique_aids  aid  top_20_clicks]))\n\ntop_aids2 = [aid2  aid2, cnt  Counter(aids2).most_common()  aid2   unique_aids]\n...\n</code></pre>\n<p>this is the code from the original notebook for suggesting clicks. I returned the aids2 list from this function and took the unique values. I did the same for carts and buys as well.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2111816,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "01/23/2023 08:12:17",
              "content": "<p>One important side note is that I generated those candidates also using Radek's validation files.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2110548,
      "author_name": "yidalin",
      "author_url": "",
      "post_date": "01/22/2023 09:11:47",
      "content": "<p>Considerate！！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2114285,
      "author_name": "butterfly915",
      "author_url": "",
      "post_date": "01/24/2023 20:37:50",
      "content": "<p>It worked! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2118489,
      "author_name": "yuzhang0422",
      "author_url": "",
      "post_date": "01/28/2023 03:40:05",
      "content": "<p>hi, thanks for share this notebook👍<br>\nyet I wonder if this part is necessary here? </p>\n<pre><code>X_train = X_train.sort_values(\"session\").reset_index(drop=True)\nX_valid = X_valid.sort_values(\"session\").reset_index(drop=True)\n</code></pre>\n<p>The index already keeps all session together. Correct me if I'm wrong :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2118827,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/28/2023 10:00:54",
          "content": "<p>Yes since we do it here as well:</p>\n<pre><code>df_cands = df_cands.sort_values().reset_index(drop=)\n</code></pre>\n<p>it becomes unnecessary. I didn't remove it since the first edition because I didn't had this part earlier then it got stuck 🙃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2122814,
      "author_name": "sangwupark",
      "author_url": "",
      "post_date": "01/31/2023 06:36:31",
      "content": "<p>Thx for sharing! </p>\n<p>yet, I hope to know how to infer! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2123212,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/31/2023 11:42:26",
          "content": "<p>Following Chris's implementation - after generating the candidates for the inference, we make predictions with each of these models (5 for clicks, 5 for buys, etc.) and take the average. Then we sort the scores in descending order (ranking). Lastly, we choose the final 20. Here is the example for clicks candidates: </p>\n<pre><code>preds = np.zeros((df_test_cand_clicks))\n fold  ():\n    model = xgb.Booster()\n    model.load_model()\n    dtest = xgb.DMatrix(data=df_test_cand_clicks[FEATURES].drop(, axis=))\n    preds += model.predict(dtest)/\npredictions = df_test_cand_clicks[[, ]].copy()\npredictions[] = preds\n\npredictions = predictions.sort_values([, ], ascending=[,]).reset_index(drop=)\npredictions[] = predictions.groupby().aid.cumcount().astype()\npredictions = predictions.loc[predictions.n&lt;]\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2108300": "Hi everyone,\n\nBased on this self-explanatory and awesome [discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) by [@cdeotte](https://www.kaggle.com/cdeotte), I would like to share the complete code for training ranker models in this competition. I just put everything together that Chris have mentioned in the discussion. \n\nThe full code for the ranking is [here](https://www.kaggle.com/code/snnclsr/gbdt-ranking).\n\n```python\nimport polars as pl\n# we are reading Radek's validation files.\ntrain = pl.read_parquet(\"../input/otto-train-and-test-data-for-local-validation/train.parquet\")\nvalid = pl.read_parquet(\"../input/otto-train-and-test-data-for-local-validation/test.parquet\")\nvalid_labels = pl.read_parquet(\"../input/otto-train-and-test-data-for-local-validation/test_labels.parquet\")\n```\n\n#### Generating the item features\n```python\n\n# {'aid':'count', 'session':'nunique', 'type': 'mean'}\nitem_features = pl.concat([train, valid]).groupby('aid').agg([\n    pl.count(\"aid\").alias(\"item_item_count\"), \n    pl.n_unique(\"session\").alias(\"item_user_count\"), \n    pl.mean(\"type\").alias(\"item_buy_ratio\").cast(pl.Float32)\n])\nitem_features.write_parquet('item_features.parquet')\n```\n#### Generating the user features\n```python\n# {'session':'count','aid':'nunique','type':'mean'}\nuser_features = valid.groupby('session').agg([\n    pl.count(\"session\").alias(\"user_user_count\"),\n    pl.n_unique(\"aid\").alias(\"user_item_count\"),\n    pl.mean(\"type\").alias(\"user_buy_ratio\").cast(pl.Float32)\n])\nuser_features.write_parquet('user_features.parquet')\n```\nCandidates generated from this [notebook](https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577). I just extracted the aids from the co-occurence matrices. \n```python\nevent_paths = {\n    \"clicks\": \"../input/otto-577-validation-candidates/valid_click_candidates.parquet\",\n    \"carts\": \"../input/otto-577-validation-candidates/valid_carts_candidates.parquet\",\n    \"buys\": \"../input/otto-577-validation-candidates/valid_buys_candidates.parquet\"\n}\n```\n#### Training the Ranker\n\nAs Chris mentioned, we are also training XGBoost. The changes I made here are:\n- group parameter: Since we have more than 50 candidates generated by the co-occurrence matrices, we update the group parameter as well to reflect that. We tell xgboost how many aids we have in each user (session) to rank. \n- New hyperparameters\n\n```python\nimport xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\ndef train_ranker(event, df_cands, n_splits=5):\n    \n    skf = GroupKFold(n_splits=n_splits)\n    FEATURES = [\n        'session', 'item_item_count', 'item_user_count', 'item_buy_ratio', 'user_user_count', 'user_item_count', 'user_buy_ratio'\n    ]\n    TARGET = \"target\"\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(df_cands, df_cands['target'], groups=df_cands['session'])):\n\n        X_train = df_cands.loc[train_idx, FEATURES]\n        y_train = df_cands.loc[train_idx, TARGET]\n        X_valid = df_cands.loc[valid_idx, FEATURES]\n        y_valid = df_cands.loc[valid_idx, TARGET]\n\n        X_train = X_train.sort_values(\"session\").reset_index(drop=True)\n        X_valid = X_valid.sort_values(\"session\").reset_index(drop=True)\n\n        train_group = X_train.groupby('session').session.agg('count').values\n        valid_group = X_valid.groupby('session').session.agg('count').values\n\n        X_train = X_train.drop([\"session\"], axis=1)\n        X_valid = X_valid.drop([\"session\"], axis=1)\n\n        dtrain = xgb.DMatrix(X_train, y_train, group=train_group) # [50] * (len(train_idx)//50) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=valid_group) # [50] * (len(valid_idx)//50)\n        xgb_parms = {\n            'objective':'rank:pairwise', \n            'tree_method':'hist',\n            'random_state': 42, \n            'learning_rate': 0.1,\n            \"colsample_bytree\": 0.8, \n            'eta': 0.05, \n            'max_depth': 6,\n            'subsample': 0.75,\n            # n_estimators=110,\n        }\n        model = xgb.train(\n            xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'), (dvalid,'valid')],\n            num_boost_round=100,\n            verbose_eval=20\n        )\n        model.save_model(f'XGB_fold{fold}_{event}.xgb')\n        gc.collect()\n\n```\nLastly we merge the features with the candidate dataframe. It's important that the sessions are sorted otherwise grouping logic above does not work.\n```python\nNEGATIVE_FRAC = 0.15\nfor event, path in event_paths.items():\n    print(f\"Started ranking model for: {event}\")\n    # Reading the candidates\n    df_cands = pl.read_parquet(path)\n    df_cands = df_cands.explode(\"candidates\").with_columns([\n        pl.col(\"session\").cast(pl.Int32),\n        pl.col(\"candidates\").cast(pl.Int32).alias(\"aid\")\n    ]).drop(\"candidates\").unique(subset=[\"session\", \"aid\"])\n    # Joining the item features\n    df_cands = df_cands.join(item_features, on='aid', how='left').fill_nan(-1)\n    # Joining the user features\n    df_cands = df_cands.join(user_features, on='session', how='left').fill_nan(-1)\n    cand_labels = valid_labels.filter(valid_labels[\"type\"] == event).explode(\"ground_truth\").with_columns([\n        pl.col(\"session\").cast(pl.Int32),\n        pl.col(\"ground_truth\").cast(pl.Int32)# .alias(\"aid\")\n    ]).rename({\"ground_truth\": \"aid\"})\n    cand_labels = cand_labels.with_column(pl.lit(1).alias(\"target\").cast(pl.Int8)).drop(\"type\")\n    # Joining the labels\n    df_cands = df_cands.join(cand_labels, on=[\"session\", \"aid\"], how=\"left\").fill_null(0)\n    # Negative sampling\n    df_cands = pl.concat([\n        df_cands.filter(df_cands[\"target\"] == 0).sample(frac=NEGATIVE_FRAC, seed=42),\n        df_cands.filter(df_cands[\"target\"] == 1)\n    ])\n    print(df_cands.groupby(\"target\").agg(pl.count()))\n    df_cands = df_cands.to_pandas()\n    df_cands = df_cands.sort_values(\"session\").reset_index(drop=True)\n    print(f\"Event: {event} - started training...\")\n\n    train_ranker(event, df_cands)\n    del df_cands, cand_labels\n    gc.collect()\n```",
    "2108580": "Great ! however I don't understand the goal of doing` pl.mean(\"type\")`. The column is categorical, even if you can average it, you have more than 2 modalities, so It doesn't make sense",
    "2108598": "Thank you for sharing.\nHow do you infer from this code?",
    "2108619": "don't underestimate the mean :)\nIn most notebooks type is encoded in this way:`{'click':0, 'cart':1, 'order':2}`.\nSo  type is more ordinal than categorical.",
    "2108626": "I totally agree ! I didn't see it from this point of view. Thanks :)",
    "2108764": "great notebook. just one point: where there is \"buys\" in your code it should be \"orders\" otherwise we dont have any positives for this category.",
    "2108797": "Thanks for sharing!",
    "2108980": "Thank you for sharing!!\nIn my experience, splitting training data and validation data after negative sampling resulted in abnormally high CV.",
    "2110492": "Thank you for your sharing great notebook.\nYou mentioned that \"Candidates generated from this notebook. I just extracted the aids from the co-occurence matrices\".\nCan I ask how you made \"valid_click_candidates.parquet\"?\nThis file was made by pred_df in the public notebook you mentioned?",
    "2110548": "Considerate！！",
    "2111811": "Interesting, what was your negative sampling ratio? and how many zeros and ones do you have?",
    "2111813": "It's not but what I did is:\n\n```python\n...\naids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n# RERANK CANDIDATES\ntop_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]\n...\n```\nthis is the code from the original notebook for suggesting clicks. I returned the aids2 list from this function and took the unique values. I did the same for carts and buys as well.",
    "2111816": "One important side note is that I generated those candidates also using Radek's validation files.",
    "2112915": "How would you do it ?",
    "2114285": "It worked! Thanks!",
    "2118489": "hi, thanks for share this notebook👍\nyet I wonder if this part is necessary here? \n```\nX_train = X_train.sort_values(\"session\").reset_index(drop=True)\nX_valid = X_valid.sort_values(\"session\").reset_index(drop=True)\n```\nThe index already keeps all session together. Correct me if I'm wrong :)",
    "2118827": "Yes since we do it here as well:\n\n```python\ndf_cands = df_cands.sort_values(\"session\").reset_index(drop=True)\n```\nit becomes unnecessary. I didn't remove it since the first edition because I didn't had this part earlier then it got stuck 🙃",
    "2122814": "Thx for sharing! \n\nyet, I hope to know how to infer!",
    "2123212": "Following Chris's implementation - after generating the candidates for the inference, we make predictions with each of these models (5 for clicks, 5 for buys, etc.) and take the average. Then we sort the scores in descending order (ranking). Lastly, we choose the final 20. Here is the example for clicks candidates: \n```python\npreds = np.zeros(len(df_test_cand_clicks))\nfor fold in range(5):\n    model = xgb.Booster()\n    model.load_model(f'XGB_fold{fold}_click.xgb')\n    dtest = xgb.DMatrix(data=df_test_cand_clicks[FEATURES].drop(\"session\", axis=\"columns\"))\n    preds += model.predict(dtest)/5\npredictions = df_test_cand_clicks[['session', 'aid']].copy()\npredictions['pred'] = preds\n\npredictions = predictions.sort_values(['session', 'aid'], ascending=[True,False]).reset_index(drop=True)\npredictions['n'] = predictions.groupby('session').aid.cumcount().astype('int8')\npredictions = predictions.loc[predictions.n<20]\n```"
  },
  "source": "meta"
}