{
  "id": 379952,
  "title": "Some optimization tips for training ranker model",
  "url": "/competitions/otto-recommender-system/discussion/379952",
  "author_name": "Gunes Evitan",
  "post_date": "2023-01-21T18:41:13.880000",
  "votes": 31,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I love this competition so much because I evolved my training code into a hyper optimized abomination. I replaced pandas with polars in most of the cases and I'm using a single polars dataframe in my training loop. I'm sharing ideas as bullet points and some code snippets with comments so you can optimize your code further.</p>\n<ul>\n<li>Use polars dataframes for faster merge operations while merging features to candidates </li>\n</ul>\n<pre><code>\ndf_aid_features = pl.from_pandas(pd.read_pickle(settings.DATA /  / ))\n\naid_merge_columns = [column  column  df_aid_features.columns  column  config[][event_type][]]\n (aid_merge_columns) &gt; :\n    df_candidate = df_candidate.join(df_aid_features.rename({: })[[] + aid_merge_columns], how=, on=)\n\n df_aid_features\n</code></pre>\n<ul>\n<li>Create folds as columns so you can select training/validation sets later without creating additional dataframes</li>\n</ul>\n<pre><code>n_splits = \ngkf = GroupKFold(n_splits=n_splits)\n fold, (train_idx, val_idx)  (gkf.split(X=df_candidate, groups=df_candidate[]), ):\n    df_candidate = df_candidate.with_column(pl.lit().alias().cast(pl.UInt8))\n    df_candidate[val_idx, ] = \nfolds = [column  column  df_candidate.columns  column.startswith()]\n</code></pre>\n<ul>\n<li>Index candidate dataframe only twice in a fold. One for undersampling negatives in training set and two for indexing training/validation sets for LightGBM datasets.</li>\n</ul>\n<pre><code> fold  folds:\n    \n    train_idx = df_candidate[fold].to_pandas() == \n    \n    val_idx = np.where(df_candidate[fold] == )[]\n\n    \n    train_negative_labels = df_candidate[target].to_pandas().loc[train_idx &amp; (df_candidate[target].to_pandas() == )]\n    negative_idx = train_negative_labels.sample(frac=, random_state=).index.to_numpy()\n     train_negative_labels\n    \n    train_idx = np.hstack((\n        np.where(train_idx &amp; (df_candidate[target].to_pandas() == ))[],\n        negative_idx\n    ))\n    \n    train_idx.sort()\n\n    \n    query_train = np.unique(df_candidate[train_idx, ], return_counts=)[].astype(np.uint16)\n    query_val = np.unique(df_candidate[val_idx, ], return_counts=)[].astype(np.uint16)\n\n    \n    train_dataset = lgb.Dataset(\n        data=df_candidate[train_idx, features].to_pandas(),\n        label=df_candidate[train_idx, target].to_pandas(),\n        group=query_train,\n        categorical_feature=categorical_features\n    )\n    val_dataset = lgb.Dataset(\n        data=df_candidate[val_idx, features].to_pandas(),\n        label=df_candidate[val_idx, target].to_pandas(),\n        group=query_val,\n        categorical_feature=categorical_features\n    )\n     query_train, query_val\n</code></pre>\n<ul>\n<li>Cast score predictions of model to 32-bit floats. Sort predictions by sessions and scores (descending) and take top 20 candidates with highest scores</li>\n</ul>\n<pre><code>\ndf_candidate[val_idx, ] = np.float32(model.predict(df_candidate[val_idx, features].to_pandas()))\ndf_val_predictions = df_candidate[\n    val_idx, [, , ]\n].sort(by=[, ], reverse=[, ])[[\n     , \n]].groupby().head()\n</code></pre>",
  "messages": [
    {
      "id": 2109900,
      "postDate": "2023-01-21T18:41:13.880Z",
      "content": "<p>I love this competition so much because I evolved my training code into a hyper optimized abomination. I replaced pandas with polars in most of the cases and I'm using a single polars dataframe in my training loop. I'm sharing ideas as bullet points and some code snippets with comments so you can optimize your code further.</p>\n<ul>\n<li>Use polars dataframes for faster merge operations while merging features to candidates </li>\n</ul>\n<pre><code>\ndf_aid_features = pl.from_pandas(pd.read_pickle(settings.DATA /  / ))\n\naid_merge_columns = [column  column  df_aid_features.columns  column  config[][event_type][]]\n (aid_merge_columns) &gt; :\n    df_candidate = df_candidate.join(df_aid_features.rename({: })[[] + aid_merge_columns], how=, on=)\n\n df_aid_features\n</code></pre>\n<ul>\n<li>Create folds as columns so you can select training/validation sets later without creating additional dataframes</li>\n</ul>\n<pre><code>n_splits = \ngkf = GroupKFold(n_splits=n_splits)\n fold, (train_idx, val_idx)  (gkf.split(X=df_candidate, groups=df_candidate[]), ):\n    df_candidate = df_candidate.with_column(pl.lit().alias().cast(pl.UInt8))\n    df_candidate[val_idx, ] = \nfolds = [column  column  df_candidate.columns  column.startswith()]\n</code></pre>\n<ul>\n<li>Index candidate dataframe only twice in a fold. One for undersampling negatives in training set and two for indexing training/validation sets for LightGBM datasets.</li>\n</ul>\n<pre><code> fold  folds:\n    \n    train_idx = df_candidate[fold].to_pandas() == \n    \n    val_idx = np.where(df_candidate[fold] == )[]\n\n    \n    train_negative_labels = df_candidate[target].to_pandas().loc[train_idx &amp; (df_candidate[target].to_pandas() == )]\n    negative_idx = train_negative_labels.sample(frac=, random_state=).index.to_numpy()\n     train_negative_labels\n    \n    train_idx = np.hstack((\n        np.where(train_idx &amp; (df_candidate[target].to_pandas() == ))[],\n        negative_idx\n    ))\n    \n    train_idx.sort()\n\n    \n    query_train = np.unique(df_candidate[train_idx, ], return_counts=)[].astype(np.uint16)\n    query_val = np.unique(df_candidate[val_idx, ], return_counts=)[].astype(np.uint16)\n\n    \n    train_dataset = lgb.Dataset(\n        data=df_candidate[train_idx, features].to_pandas(),\n        label=df_candidate[train_idx, target].to_pandas(),\n        group=query_train,\n        categorical_feature=categorical_features\n    )\n    val_dataset = lgb.Dataset(\n        data=df_candidate[val_idx, features].to_pandas(),\n        label=df_candidate[val_idx, target].to_pandas(),\n        group=query_val,\n        categorical_feature=categorical_features\n    )\n     query_train, query_val\n</code></pre>\n<ul>\n<li>Cast score predictions of model to 32-bit floats. Sort predictions by sessions and scores (descending) and take top 20 candidates with highest scores</li>\n</ul>\n<pre><code>\ndf_candidate[val_idx, ] = np.float32(model.predict(df_candidate[val_idx, features].to_pandas()))\ndf_val_predictions = df_candidate[\n    val_idx, [, , ]\n].sort(by=[, ], reverse=[, ])[[\n     , \n]].groupby().head()\n</code></pre>",
      "rawMarkdown": "I love this competition so much because I evolved my training code into a hyper optimized abomination. I replaced pandas with polars in most of the cases and I'm using a single polars dataframe in my training loop. I'm sharing ideas as bullet points and some code snippets with comments so you can optimize your code further.\n\n* Use polars dataframes for faster merge operations while merging features to candidates \n\n```python\n# Use polars dataframe for fast merge\ndf_aid_features = pl.from_pandas(pd.read_pickle(settings.DATA / 'feature_engineering' / 'train_aid_features.pkl'))\n# Merge columns that you use in training\naid_merge_columns = [column for column in df_aid_features.columns if column in config['dataset'][event_type]['features']]\nif len(aid_merge_columns) > 0:\n    df_candidate = df_candidate.join(df_aid_features.rename({'aid': 'candidates'})[['candidates'] + aid_merge_columns], how='left', on='candidates')\n# Delete dataframe objects after you are done with them\ndel df_aid_features\n```\n* Create folds as columns so you can select training/validation sets later without creating additional dataframes\n\n```python\nn_splits = 5\ngkf = GroupKFold(n_splits=n_splits)\nfor fold, (train_idx, val_idx) in enumerate(gkf.split(X=df_candidate, groups=df_candidate['session']), 1):\n    df_candidate = df_candidate.with_column(pl.lit(0).alias(f'fold{fold}').cast(pl.UInt8))\n    df_candidate[val_idx, f'fold{fold}'] = 1\nfolds = [column for column in df_candidate.columns if column.startswith('fold')]\n```\n* Index candidate dataframe only twice in a fold. One for undersampling negatives in training set and two for indexing training/validation sets for LightGBM datasets.\n\n```python\nfor fold in folds:\n    # train_idx has to be pandas because of negative sampling\n    train_idx = df_candidate[fold].to_pandas() == 0\n    # polars boolean masks can be converted into index with np.where\n    val_idx = np.where(df_candidate[fold] == 1)[0]\n\n    # Index negatives in training set and sample from them\n    train_negative_labels = df_candidate[target].to_pandas().loc[train_idx & (df_candidate[target].to_pandas() == 0)]\n    negative_idx = train_negative_labels.sample(frac=0.1, random_state=42).index.to_numpy()\n    del train_negative_labels\n    # Combine train positive index and sampled negative index\n    train_idx = np.hstack((\n        np.where(train_idx & (df_candidate[target].to_pandas() == 1))[0],\n        negative_idx\n    ))\n    # Sort training index for retaining session order\n    train_idx.sort()\n\n    # Extract occurrence counts of groups in training and validation sets\n    query_train = np.unique(df_candidate[train_idx, 'session'], return_counts=True)[1].astype(np.uint16)\n    query_val = np.unique(df_candidate[val_idx, 'session'], return_counts=True)[1].astype(np.uint16)\n\n    # LightGBM doesn't work with polars dataframes for some reason\n    train_dataset = lgb.Dataset(\n        data=df_candidate[train_idx, features].to_pandas(),\n        label=df_candidate[train_idx, target].to_pandas(),\n        group=query_train,\n        categorical_feature=categorical_features\n    )\n    val_dataset = lgb.Dataset(\n        data=df_candidate[val_idx, features].to_pandas(),\n        label=df_candidate[val_idx, target].to_pandas(),\n        group=query_val,\n        categorical_feature=categorical_features\n    )\n    del query_train, query_val\n```\n\n* Cast score predictions of model to 32-bit floats. Sort predictions by sessions and scores (descending) and take top 20 candidates with highest scores\n```python\n# Predict validation dataset and retrieve top 20 predictions with the highest score for every session\ndf_candidate[val_idx, 'predictions'] = np.float32(model.predict(df_candidate[val_idx, features].to_pandas()))\ndf_val_predictions = df_candidate[\n    val_idx, ['session', 'candidates', 'predictions']\n].sort(by=['session', 'predictions'], reverse=[False, True])[[\n     'session', 'candidates'\n]].groupby('session').head(20)\n```",
      "votes": 31
    },
    {
      "id": 2110140,
      "postDate": "2023-01-22T02:13:16.253Z",
      "content": "<p>How much of a speed up did you see going from pandas to polars?</p>",
      "rawMarkdown": "How much of a speed up did you see going from pandas to polars?",
      "votes": 1,
      "replies": [
        {
          "id": 2110408,
          "postDate": "2023-01-22T07:22:24.623Z",
          "content": "<p>About 6x on AMD Ryzen 9 5950X 16 core, 32 thread CPU.</p>",
          "rawMarkdown": "About 6x on AMD Ryzen 9 5950X 16 core, 32 thread CPU."
        },
        {
          "id": 2110601,
          "postDate": "2023-01-22T09:57:33.500Z",
          "content": "<p>Transferring your pipe from pandas to polars also introduces memory efficiency most of the time. (especially if you use the benefits of lazy mode)</p>",
          "rawMarkdown": "Transferring your pipe from pandas to polars also introduces memory efficiency most of the time. (especially if you use the benefits of lazy mode)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2112694,
      "postDate": "2023-01-23T18:57:50.850Z",
      "content": "<p>Nice note, thanks! Cudf also not bad for merging/joining/applying</p>",
      "rawMarkdown": "Nice note, thanks! Cudf also not bad for merging/joining/applying"
    },
    {
      "id": 2111518,
      "postDate": "2023-01-22T23:56:51.803Z",
      "content": "<p>This is a great post! Thank you for sharing!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "This is a great post! Thank you for sharing!\n\nThe Devastator.\n"
    },
    {
      "id": 2110543,
      "postDate": "2023-01-22T09:10:42.620Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2111582,
      "postDate": "2023-01-23T02:07:51.710Z",
      "content": "<p>Thanks!!!!</p>",
      "rawMarkdown": "Thanks!!!!"
    }
  ],
  "comments": [
    {
      "id": 2110140,
      "author_name": "datadote",
      "author_url": "",
      "post_date": "2023-01-22T02:13:16.253000",
      "content": "<p>How much of a speed up did you see going from pandas to polars?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2110408,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-01-22T07:22:24.623000",
          "content": "<p>About 6x on AMD Ryzen 9 5950X 16 core, 32 thread CPU.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2110601,
          "author_name": "Anil Ozturk",
          "author_url": "",
          "post_date": "2023-01-22T09:57:33.500000",
          "content": "<p>Transferring your pipe from pandas to polars also introduces memory efficiency most of the time. (especially if you use the benefits of lazy mode)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2112694,
      "author_name": "Mike Mazurov",
      "author_url": "",
      "post_date": "2023-01-23T18:57:50.850000",
      "content": "<p>Nice note, thanks! Cudf also not bad for merging/joining/applying</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2111518,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-01-22T23:56:51.803000",
      "content": "<p>This is a great post! Thank you for sharing!</p>\n<p>The Devastator.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2110543,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-22T09:10:42.620000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2111582,
      "author_name": "Qinfeng He",
      "author_url": "",
      "post_date": "2023-01-23T02:07:51.710000",
      "content": "<p>Thanks!!!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2109900": "I love this competition so much because I evolved my training code into a hyper optimized abomination. I replaced pandas with polars in most of the cases and I'm using a single polars dataframe in my training loop. I'm sharing ideas as bullet points and some code snippets with comments so you can optimize your code further.\n\n* Use polars dataframes for faster merge operations while merging features to candidates \n\n```python\n# Use polars dataframe for fast merge\ndf_aid_features = pl.from_pandas(pd.read_pickle(settings.DATA / 'feature_engineering' / 'train_aid_features.pkl'))\n# Merge columns that you use in training\naid_merge_columns = [column for column in df_aid_features.columns if column in config['dataset'][event_type]['features']]\nif len(aid_merge_columns) > 0:\n    df_candidate = df_candidate.join(df_aid_features.rename({'aid': 'candidates'})[['candidates'] + aid_merge_columns], how='left', on='candidates')\n# Delete dataframe objects after you are done with them\ndel df_aid_features\n```\n* Create folds as columns so you can select training/validation sets later without creating additional dataframes\n\n```python\nn_splits = 5\ngkf = GroupKFold(n_splits=n_splits)\nfor fold, (train_idx, val_idx) in enumerate(gkf.split(X=df_candidate, groups=df_candidate['session']), 1):\n    df_candidate = df_candidate.with_column(pl.lit(0).alias(f'fold{fold}').cast(pl.UInt8))\n    df_candidate[val_idx, f'fold{fold}'] = 1\nfolds = [column for column in df_candidate.columns if column.startswith('fold')]\n```\n* Index candidate dataframe only twice in a fold. One for undersampling negatives in training set and two for indexing training/validation sets for LightGBM datasets.\n\n```python\nfor fold in folds:\n    # train_idx has to be pandas because of negative sampling\n    train_idx = df_candidate[fold].to_pandas() == 0\n    # polars boolean masks can be converted into index with np.where\n    val_idx = np.where(df_candidate[fold] == 1)[0]\n\n    # Index negatives in training set and sample from them\n    train_negative_labels = df_candidate[target].to_pandas().loc[train_idx & (df_candidate[target].to_pandas() == 0)]\n    negative_idx = train_negative_labels.sample(frac=0.1, random_state=42).index.to_numpy()\n    del train_negative_labels\n    # Combine train positive index and sampled negative index\n    train_idx = np.hstack((\n        np.where(train_idx & (df_candidate[target].to_pandas() == 1))[0],\n        negative_idx\n    ))\n    # Sort training index for retaining session order\n    train_idx.sort()\n\n    # Extract occurrence counts of groups in training and validation sets\n    query_train = np.unique(df_candidate[train_idx, 'session'], return_counts=True)[1].astype(np.uint16)\n    query_val = np.unique(df_candidate[val_idx, 'session'], return_counts=True)[1].astype(np.uint16)\n\n    # LightGBM doesn't work with polars dataframes for some reason\n    train_dataset = lgb.Dataset(\n        data=df_candidate[train_idx, features].to_pandas(),\n        label=df_candidate[train_idx, target].to_pandas(),\n        group=query_train,\n        categorical_feature=categorical_features\n    )\n    val_dataset = lgb.Dataset(\n        data=df_candidate[val_idx, features].to_pandas(),\n        label=df_candidate[val_idx, target].to_pandas(),\n        group=query_val,\n        categorical_feature=categorical_features\n    )\n    del query_train, query_val\n```\n\n* Cast score predictions of model to 32-bit floats. Sort predictions by sessions and scores (descending) and take top 20 candidates with highest scores\n```python\n# Predict validation dataset and retrieve top 20 predictions with the highest score for every session\ndf_candidate[val_idx, 'predictions'] = np.float32(model.predict(df_candidate[val_idx, features].to_pandas()))\ndf_val_predictions = df_candidate[\n    val_idx, ['session', 'candidates', 'predictions']\n].sort(by=['session', 'predictions'], reverse=[False, True])[[\n     'session', 'candidates'\n]].groupby('session').head(20)\n```",
    "2110140": "How much of a speed up did you see going from pandas to polars?",
    "2112694": "Nice note, thanks! Cudf also not bad for merging/joining/applying",
    "2111518": "This is a great post! Thank you for sharing!\n\nThe Devastator.\n",
    "2110543": "",
    "2111582": "Thanks!!!!"
  }
}