{
  "id": 383016,
  "title": "46th Place Solution (Code and Journey)",
  "url": "/competitions/otto-recommender-system/discussion/383016",
  "author_name": "",
  "post_date": "2023-02-02T01:23:18.876459700Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong>Leaderboard Journey:</strong><br>\nAs always the Leaderboard was an emotional roller-coaster…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc2a4c91ac3f11f2af39404d4ae408967%2Fleaderboard_journey.PNG?generation=1675298568954378&amp;alt=media\" alt=\"leaderboard journey1\"></p>\n<ol>\n<li>I started with a submission using just items (aids) already in the basket, padded by covisitation matrix candidates to make sure I could replicate the public scores.</li>\n<li>Next I created 30 candidates by allowing more items and used a xgboost model trained on half the data to order the candidates. This was the first time I got close to the best public notebooks. </li>\n<li>My biggest boost came from increasing the number of candidates to 200, and training seperate clicks, carts, and orders booster on all the data using very simple features (e.g count of interactions, test popularity, etc.) .</li>\n<li>Adding a bunch of well thought out features than pushed my score to 0.59. Features such as percent of carts that turn to orders and minutes since a user interacted with the item turned out to be very useful. </li>\n<li>Adding what I thought where even more logical features such as \"minutes since someone else has ordered this product\" and \"percent of people that re-order this product\" did not move the score from 0.59. </li>\n<li>As a last attempt I trained a lightgbm model and ensembled that with the xgboost models using these new clever features, but that still didn't shift the score! </li>\n</ol>\n<p><strong>What the final pipeline looks like</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fe5d4789ea44792c74e86d37d50e00fab%2Fstrategy_overview.PNG?generation=1675300350438259&amp;alt=media\" alt=\"finalpipeline\"></p>\n<p><strong>Competition Metric at 200 candidates</strong><br>\nI prioritised the strategies in the diagram above from left to right, until 200 candidates where generated. For example, if you had 100 basket (session) candidates, and 100 covisitation candidates, you wouldn't get any from the word2vec and ALS models. Using this strategy I was able to get the following CV recall scores for my 200 candidates. </p>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Recall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>clicks recall</td>\n<td>0.6583</td>\n</tr>\n<tr>\n<td>carts recall</td>\n<td>0.5294</td>\n</tr>\n<tr>\n<td>orders recall</td>\n<td>0.7211</td>\n</tr>\n<tr>\n<td>Overall Recall</td>\n<td>0.6573</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Model Importance</strong><br>\nThe best features where generally:</p>\n<ol>\n<li>Those relating to the items position in the session - e.g was this the last click in the session, 3rd last click, or</li>\n<li>What rank the item achieved under the various candidate generation strategies (these all have a n_ prefix in the charts below).</li>\n</ol>\n<p>You can see 2. below - the number 1 feature in the clicks booster is what the item ranked under the covisitation metric, and things with low rankings (meaning the top candidate) under the strategy have a strong positive impact. </p>\n<p>Top 10 features in the clicks model:<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2F98d5fd5b0fd4cc3a3bf9c9aed590c182%2FCapture.PNG?generation=1675298216480638&amp;alt=media\" alt=\"Top 10 features clicks model\"></p>\n<p>Similarly for order predictions we can see the ranking under the basket generation candidate strategy (basically how recently the item occurred in the session) is the number 1 feature</p>\n<p>Top 10 features in the orders model:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc9da8552827a5e821b45d51e11da15dd%2Forder_shap_values.PNG?generation=1675298385004945&amp;alt=media\" alt=\"Top 10 features orders model\"></p>\n<p><strong>What worked</strong></p>\n<ul>\n<li>Being able to test the recall from each of the candidate generation methods felt pretty valuable. I probably spent two weeks building candidates never submitting anything to the leaderboard, but when I eventually did there was a decent score boost, so the candidates must have been okay!</li>\n<li>Moving the feature generation code from Pandas to Polars helped avoid a lot of out of memory issues, and ran a lot quicker. </li>\n<li>Subsampling the data for development, and only using all the data when building final outputs saved a lot of time. </li>\n<li>Investing time in thinking through quality features was very useful. </li>\n<li>I did everything on colab, and being able to link colab to a google cloud instance paired with free google cloud credits for new users was pretty useful!</li>\n</ul>\n<p><strong>What didn't work</strong></p>\n<ul>\n<li>Using different models such as lightgbm, or spending time trying to finetune my hyperparameters didn't have a significant impact on the public leaderboard score.</li>\n<li>At the end of the competition I could increase local CV scores, but this didn't translate to an increase on the leader board. Initially I thought this was because the booster features for training where being calculated across 3 weeks meaning they were materially different from those used in the test set calculated across 4 weeks and throwing out the models, but even after normalising features by doing everything at the per week level the issue remained! If anyone else had an issue with CV gains not translating to the leaderboard would love to hear how you solved them. </li>\n</ul>\n<p><strong>Set Up</strong><br>\nI used google colab for everything. When I started having memory issues (very quickly) I began spinning out google cloud instances and linking colab to them. All the data was always stored on my google drive which I would mount to the colab notebook. I would probably use a different set up in future - those that had a different set up what did you use? Would you recommend it? </p>\n<p><strong>Shout outs</strong></p>\n<ul>\n<li>At the heart of the notebook is the validation dataset uploaded by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>.</li>\n<li>The initial covisitation matrices  from <a href=\"https://www.kaggle.com/cdeottes\" target=\"_blank\">@cdeottes</a> notebook here : <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575</a> helped a lot in generation a large number of candidate. My biggest regret is that I couldn't get rapids to work on colab though and had to generate these using pandas 😪</li>\n<li>The word2vec candidate strategy is heavily inspired by <a href=\"https://www.kaggle.com/radek1s\" target=\"_blank\">@radek1s</a> notebook here: <a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission</a> </li>\n</ul>\n<p><strong>Code</strong><br>\n<a href=\"https://github.com/rhyscook92/kaggle-otto-recommender-2022\" target=\"_blank\">https://github.com/rhyscook92/kaggle-otto-recommender-2022</a></p>\n<p><strong>Last Words</strong><br>\nBig thanks to the team at Kaggle and Otto for organising such a fun competition, and for everyone commenting in the discussions and sharing code for the opportunity to learn a lot!</p>",
  "messages": [
    {
      "id": "2125896",
      "postDate": "02/02/2023 01:23:18",
      "content": "<p><strong>Leaderboard Journey:</strong><br>\nAs always the Leaderboard was an emotional roller-coaster…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc2a4c91ac3f11f2af39404d4ae408967%2Fleaderboard_journey.PNG?generation=1675298568954378&amp;alt=media\" alt=\"leaderboard journey1\"></p>\n<ol>\n<li>I started with a submission using just items (aids) already in the basket, padded by covisitation matrix candidates to make sure I could replicate the public scores.</li>\n<li>Next I created 30 candidates by allowing more items and used a xgboost model trained on half the data to order the candidates. This was the first time I got close to the best public notebooks. </li>\n<li>My biggest boost came from increasing the number of candidates to 200, and training seperate clicks, carts, and orders booster on all the data using very simple features (e.g count of interactions, test popularity, etc.) .</li>\n<li>Adding a bunch of well thought out features than pushed my score to 0.59. Features such as percent of carts that turn to orders and minutes since a user interacted with the item turned out to be very useful. </li>\n<li>Adding what I thought where even more logical features such as \"minutes since someone else has ordered this product\" and \"percent of people that re-order this product\" did not move the score from 0.59. </li>\n<li>As a last attempt I trained a lightgbm model and ensembled that with the xgboost models using these new clever features, but that still didn't shift the score! </li>\n</ol>\n<p><strong>What the final pipeline looks like</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fe5d4789ea44792c74e86d37d50e00fab%2Fstrategy_overview.PNG?generation=1675300350438259&amp;alt=media\" alt=\"finalpipeline\"></p>\n<p><strong>Competition Metric at 200 candidates</strong><br>\nI prioritised the strategies in the diagram above from left to right, until 200 candidates where generated. For example, if you had 100 basket (session) candidates, and 100 covisitation candidates, you wouldn't get any from the word2vec and ALS models. Using this strategy I was able to get the following CV recall scores for my 200 candidates. </p>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Recall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>clicks recall</td>\n<td>0.6583</td>\n</tr>\n<tr>\n<td>carts recall</td>\n<td>0.5294</td>\n</tr>\n<tr>\n<td>orders recall</td>\n<td>0.7211</td>\n</tr>\n<tr>\n<td>Overall Recall</td>\n<td>0.6573</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Model Importance</strong><br>\nThe best features where generally:</p>\n<ol>\n<li>Those relating to the items position in the session - e.g was this the last click in the session, 3rd last click, or</li>\n<li>What rank the item achieved under the various candidate generation strategies (these all have a n_ prefix in the charts below).</li>\n</ol>\n<p>You can see 2. below - the number 1 feature in the clicks booster is what the item ranked under the covisitation metric, and things with low rankings (meaning the top candidate) under the strategy have a strong positive impact. </p>\n<p>Top 10 features in the clicks model:<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2F98d5fd5b0fd4cc3a3bf9c9aed590c182%2FCapture.PNG?generation=1675298216480638&amp;alt=media\" alt=\"Top 10 features clicks model\"></p>\n<p>Similarly for order predictions we can see the ranking under the basket generation candidate strategy (basically how recently the item occurred in the session) is the number 1 feature</p>\n<p>Top 10 features in the orders model:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc9da8552827a5e821b45d51e11da15dd%2Forder_shap_values.PNG?generation=1675298385004945&amp;alt=media\" alt=\"Top 10 features orders model\"></p>\n<p><strong>What worked</strong></p>\n<ul>\n<li>Being able to test the recall from each of the candidate generation methods felt pretty valuable. I probably spent two weeks building candidates never submitting anything to the leaderboard, but when I eventually did there was a decent score boost, so the candidates must have been okay!</li>\n<li>Moving the feature generation code from Pandas to Polars helped avoid a lot of out of memory issues, and ran a lot quicker. </li>\n<li>Subsampling the data for development, and only using all the data when building final outputs saved a lot of time. </li>\n<li>Investing time in thinking through quality features was very useful. </li>\n<li>I did everything on colab, and being able to link colab to a google cloud instance paired with free google cloud credits for new users was pretty useful!</li>\n</ul>\n<p><strong>What didn't work</strong></p>\n<ul>\n<li>Using different models such as lightgbm, or spending time trying to finetune my hyperparameters didn't have a significant impact on the public leaderboard score.</li>\n<li>At the end of the competition I could increase local CV scores, but this didn't translate to an increase on the leader board. Initially I thought this was because the booster features for training where being calculated across 3 weeks meaning they were materially different from those used in the test set calculated across 4 weeks and throwing out the models, but even after normalising features by doing everything at the per week level the issue remained! If anyone else had an issue with CV gains not translating to the leaderboard would love to hear how you solved them. </li>\n</ul>\n<p><strong>Set Up</strong><br>\nI used google colab for everything. When I started having memory issues (very quickly) I began spinning out google cloud instances and linking colab to them. All the data was always stored on my google drive which I would mount to the colab notebook. I would probably use a different set up in future - those that had a different set up what did you use? Would you recommend it? </p>\n<p><strong>Shout outs</strong></p>\n<ul>\n<li>At the heart of the notebook is the validation dataset uploaded by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>.</li>\n<li>The initial covisitation matrices  from <a href=\"https://www.kaggle.com/cdeottes\" target=\"_blank\">@cdeottes</a> notebook here : <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575</a> helped a lot in generation a large number of candidate. My biggest regret is that I couldn't get rapids to work on colab though and had to generate these using pandas 😪</li>\n<li>The word2vec candidate strategy is heavily inspired by <a href=\"https://www.kaggle.com/radek1s\" target=\"_blank\">@radek1s</a> notebook here: <a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission</a> </li>\n</ul>\n<p><strong>Code</strong><br>\n<a href=\"https://github.com/rhyscook92/kaggle-otto-recommender-2022\" target=\"_blank\">https://github.com/rhyscook92/kaggle-otto-recommender-2022</a></p>\n<p><strong>Last Words</strong><br>\nBig thanks to the team at Kaggle and Otto for organising such a fun competition, and for everyone commenting in the discussions and sharing code for the opportunity to learn a lot!</p>",
      "rawMarkdown": "**Leaderboard Journey:**\nAs always the Leaderboard was an emotional roller-coaster...\n\n![leaderboard journey1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc2a4c91ac3f11f2af39404d4ae408967%2Fleaderboard_journey.PNG?generation=1675298568954378&alt=media)\n\n\n1. I started with a submission using just items (aids) already in the basket, padded by covisitation matrix candidates to make sure I could replicate the public scores.\n2. Next I created 30 candidates by allowing more items and used a xgboost model trained on half the data to order the candidates. This was the first time I got close to the best public notebooks. \n3. My biggest boost came from increasing the number of candidates to 200, and training seperate clicks, carts, and orders booster on all the data using very simple features (e.g count of interactions, test popularity, etc.) .\n4. Adding a bunch of well thought out features than pushed my score to 0.59. Features such as percent of carts that turn to orders and minutes since a user interacted with the item turned out to be very useful. \n5. Adding what I thought where even more logical features such as \"minutes since someone else has ordered this product\" and \"percent of people that re-order this product\" did not move the score from 0.59. \n6. As a last attempt I trained a lightgbm model and ensembled that with the xgboost models using these new clever features, but that still didn't shift the score! \n\n**What the final pipeline looks like**\n![finalpipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fe5d4789ea44792c74e86d37d50e00fab%2Fstrategy_overview.PNG?generation=1675300350438259&alt=media)\n\n**Competition Metric at 200 candidates**\nI prioritised the strategies in the diagram above from left to right, until 200 candidates where generated. For example, if you had 100 basket (session) candidates, and 100 covisitation candidates, you wouldn't get any from the word2vec and ALS models. Using this strategy I was able to get the following CV recall scores for my 200 candidates. \n| Type | Recall |\n| --- | --- |\n| clicks recall | 0.6583 |\n| carts recall | 0.5294 |\n| orders recall | 0.7211 |\n| Overall Recall | 0.6573 |\n\n**Model Importance**\nThe best features where generally:\n1. Those relating to the items position in the session - e.g was this the last click in the session, 3rd last click, or\n2. What rank the item achieved under the various candidate generation strategies (these all have a n_ prefix in the charts below).\n\nYou can see 2. below - the number 1 feature in the clicks booster is what the item ranked under the covisitation metric, and things with low rankings (meaning the top candidate) under the strategy have a strong positive impact. \n\nTop 10 features in the clicks model:\n ![Top 10 features clicks model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2F98d5fd5b0fd4cc3a3bf9c9aed590c182%2FCapture.PNG?generation=1675298216480638&alt=media)\n\nSimilarly for order predictions we can see the ranking under the basket generation candidate strategy (basically how recently the item occurred in the session) is the number 1 feature\n\nTop 10 features in the orders model:\n![Top 10 features orders model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc9da8552827a5e821b45d51e11da15dd%2Forder_shap_values.PNG?generation=1675298385004945&alt=media)\n\n**What worked**\n- Being able to test the recall from each of the candidate generation methods felt pretty valuable. I probably spent two weeks building candidates never submitting anything to the leaderboard, but when I eventually did there was a decent score boost, so the candidates must have been okay!\n- Moving the feature generation code from Pandas to Polars helped avoid a lot of out of memory issues, and ran a lot quicker. \n- Subsampling the data for development, and only using all the data when building final outputs saved a lot of time. \n- Investing time in thinking through quality features was very useful. \n- I did everything on colab, and being able to link colab to a google cloud instance paired with free google cloud credits for new users was pretty useful!\n\n**What didn't work**\n- Using different models such as lightgbm, or spending time trying to finetune my hyperparameters didn't have a significant impact on the public leaderboard score.\n- At the end of the competition I could increase local CV scores, but this didn't translate to an increase on the leader board. Initially I thought this was because the booster features for training where being calculated across 3 weeks meaning they were materially different from those used in the test set calculated across 4 weeks and throwing out the models, but even after normalising features by doing everything at the per week level the issue remained! If anyone else had an issue with CV gains not translating to the leaderboard would love to hear how you solved them. \n\n**Set Up**\nI used google colab for everything. When I started having memory issues (very quickly) I began spinning out google cloud instances and linking colab to them. All the data was always stored on my google drive which I would mount to the colab notebook. I would probably use a different set up in future - those that had a different set up what did you use? Would you recommend it? \n\n**Shout outs**\n- At the heart of the notebook is the validation dataset uploaded by @radek1.\n- The initial covisitation matrices  from @cdeottes notebook here : https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575 helped a lot in generation a large number of candidate. My biggest regret is that I couldn't get rapids to work on colab though and had to generate these using pandas 😪\n- The word2vec candidate strategy is heavily inspired by @radek1s notebook here: https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission \n\n**Code**\nhttps://github.com/rhyscook92/kaggle-otto-recommender-2022\n\n**Last Words**\nBig thanks to the team at Kaggle and Otto for organising such a fun competition, and for everyone commenting in the discussions and sharing code for the opportunity to learn a lot!",
      "votes": null
    },
    {
      "id": "2125900",
      "postDate": "02/02/2023 01:29:20",
      "content": "<p>Congratulations on solo 52nd Silver finish. Great work! </p>",
      "rawMarkdown": "Congratulations on solo 52nd Silver finish. Great work!",
      "votes": null
    },
    {
      "id": "2125904",
      "postDate": "02/02/2023 01:38:37",
      "content": "<p>Thanks Chris! I picked up a lot from your sharing, and congratulations on the 4th place!</p>",
      "rawMarkdown": "Thanks Chris! I picked up a lot from your sharing, and congratulations on the 4th place!",
      "votes": null
    },
    {
      "id": "2128118",
      "postDate": "02/03/2023 13:29:39",
      "content": "<p>I love the graph, as I can very well relate to it. Congratulations</p>",
      "rawMarkdown": "I love the graph, as I can very well relate to it. Congratulations",
      "votes": null
    },
    {
      "id": "2129116",
      "postDate": "02/04/2023 10:44:57",
      "content": "<p>Thanks! Good to know the highs and lows are the same all the way at the top of the leaderboard too - congratulations on 6th. </p>",
      "rawMarkdown": "Thanks! Good to know the highs and lows are the same all the way at the top of the leaderboard too - congratulations on 6th.",
      "votes": null
    },
    {
      "id": "2153016",
      "postDate": "02/21/2023 06:26:33",
      "content": "<p>Hi. <a href=\"https://www.kaggle.com/rhysie\" target=\"_blank\">@rhysie</a> </p>\n<p>Thanks for sharing.  Finally there is solution with codes. It will be a good resource to learn for me! Thanks!!</p>",
      "rawMarkdown": "Hi. @rhysie \n\nThanks for sharing.  Finally there is solution with codes. It will be a good resource to learn for me! Thanks!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125900,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/02/2023 01:29:20",
      "content": "<p>Congratulations on solo 52nd Silver finish. Great work! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2125904,
          "author_name": "rhysie",
          "author_url": "",
          "post_date": "02/02/2023 01:38:37",
          "content": "<p>Thanks Chris! I picked up a lot from your sharing, and congratulations on the 4th place!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2128118,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "02/03/2023 13:29:39",
      "content": "<p>I love the graph, as I can very well relate to it. Congratulations</p>",
      "votes": null,
      "replies": [
        {
          "id": 2129116,
          "author_name": "rhysie",
          "author_url": "",
          "post_date": "02/04/2023 10:44:57",
          "content": "<p>Thanks! Good to know the highs and lows are the same all the way at the top of the leaderboard too - congratulations on 6th. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2153016,
      "author_name": "huaguo",
      "author_url": "",
      "post_date": "02/21/2023 06:26:33",
      "content": "<p>Hi. <a href=\"https://www.kaggle.com/rhysie\" target=\"_blank\">@rhysie</a> </p>\n<p>Thanks for sharing.  Finally there is solution with codes. It will be a good resource to learn for me! Thanks!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2125896": "**Leaderboard Journey:**\nAs always the Leaderboard was an emotional roller-coaster...\n\n![leaderboard journey1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc2a4c91ac3f11f2af39404d4ae408967%2Fleaderboard_journey.PNG?generation=1675298568954378&alt=media)\n\n\n1. I started with a submission using just items (aids) already in the basket, padded by covisitation matrix candidates to make sure I could replicate the public scores.\n2. Next I created 30 candidates by allowing more items and used a xgboost model trained on half the data to order the candidates. This was the first time I got close to the best public notebooks. \n3. My biggest boost came from increasing the number of candidates to 200, and training seperate clicks, carts, and orders booster on all the data using very simple features (e.g count of interactions, test popularity, etc.) .\n4. Adding a bunch of well thought out features than pushed my score to 0.59. Features such as percent of carts that turn to orders and minutes since a user interacted with the item turned out to be very useful. \n5. Adding what I thought where even more logical features such as \"minutes since someone else has ordered this product\" and \"percent of people that re-order this product\" did not move the score from 0.59. \n6. As a last attempt I trained a lightgbm model and ensembled that with the xgboost models using these new clever features, but that still didn't shift the score! \n\n**What the final pipeline looks like**\n![finalpipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fe5d4789ea44792c74e86d37d50e00fab%2Fstrategy_overview.PNG?generation=1675300350438259&alt=media)\n\n**Competition Metric at 200 candidates**\nI prioritised the strategies in the diagram above from left to right, until 200 candidates where generated. For example, if you had 100 basket (session) candidates, and 100 covisitation candidates, you wouldn't get any from the word2vec and ALS models. Using this strategy I was able to get the following CV recall scores for my 200 candidates. \n| Type | Recall |\n| --- | --- |\n| clicks recall | 0.6583 |\n| carts recall | 0.5294 |\n| orders recall | 0.7211 |\n| Overall Recall | 0.6573 |\n\n**Model Importance**\nThe best features where generally:\n1. Those relating to the items position in the session - e.g was this the last click in the session, 3rd last click, or\n2. What rank the item achieved under the various candidate generation strategies (these all have a n_ prefix in the charts below).\n\nYou can see 2. below - the number 1 feature in the clicks booster is what the item ranked under the covisitation metric, and things with low rankings (meaning the top candidate) under the strategy have a strong positive impact. \n\nTop 10 features in the clicks model:\n ![Top 10 features clicks model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2F98d5fd5b0fd4cc3a3bf9c9aed590c182%2FCapture.PNG?generation=1675298216480638&alt=media)\n\nSimilarly for order predictions we can see the ranking under the basket generation candidate strategy (basically how recently the item occurred in the session) is the number 1 feature\n\nTop 10 features in the orders model:\n![Top 10 features orders model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F389918%2Fc9da8552827a5e821b45d51e11da15dd%2Forder_shap_values.PNG?generation=1675298385004945&alt=media)\n\n**What worked**\n- Being able to test the recall from each of the candidate generation methods felt pretty valuable. I probably spent two weeks building candidates never submitting anything to the leaderboard, but when I eventually did there was a decent score boost, so the candidates must have been okay!\n- Moving the feature generation code from Pandas to Polars helped avoid a lot of out of memory issues, and ran a lot quicker. \n- Subsampling the data for development, and only using all the data when building final outputs saved a lot of time. \n- Investing time in thinking through quality features was very useful. \n- I did everything on colab, and being able to link colab to a google cloud instance paired with free google cloud credits for new users was pretty useful!\n\n**What didn't work**\n- Using different models such as lightgbm, or spending time trying to finetune my hyperparameters didn't have a significant impact on the public leaderboard score.\n- At the end of the competition I could increase local CV scores, but this didn't translate to an increase on the leader board. Initially I thought this was because the booster features for training where being calculated across 3 weeks meaning they were materially different from those used in the test set calculated across 4 weeks and throwing out the models, but even after normalising features by doing everything at the per week level the issue remained! If anyone else had an issue with CV gains not translating to the leaderboard would love to hear how you solved them. \n\n**Set Up**\nI used google colab for everything. When I started having memory issues (very quickly) I began spinning out google cloud instances and linking colab to them. All the data was always stored on my google drive which I would mount to the colab notebook. I would probably use a different set up in future - those that had a different set up what did you use? Would you recommend it? \n\n**Shout outs**\n- At the heart of the notebook is the validation dataset uploaded by @radek1.\n- The initial covisitation matrices  from @cdeottes notebook here : https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575 helped a lot in generation a large number of candidate. My biggest regret is that I couldn't get rapids to work on colab though and had to generate these using pandas 😪\n- The word2vec candidate strategy is heavily inspired by @radek1s notebook here: https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission \n\n**Code**\nhttps://github.com/rhyscook92/kaggle-otto-recommender-2022\n\n**Last Words**\nBig thanks to the team at Kaggle and Otto for organising such a fun competition, and for everyone commenting in the discussions and sharing code for the opportunity to learn a lot!",
    "2125900": "Congratulations on solo 52nd Silver finish. Great work!",
    "2125904": "Thanks Chris! I picked up a lot from your sharing, and congratulations on the 4th place!",
    "2128118": "I love the graph, as I can very well relate to it. Congratulations",
    "2129116": "Thanks! Good to know the highs and lows are the same all the way at the top of the leaderboard too - congratulations on 6th.",
    "2153016": "Hi. @rhysie \n\nThanks for sharing.  Finally there is solution with codes. It will be a good resource to learn for me! Thanks!!"
  },
  "source": "meta"
}