{
  "id": 382975,
  "title": "3rd Place Solution - Theo's Part",
  "url": "/competitions/otto-recommender-system/discussion/382975",
  "author_name": "",
  "post_date": "2023-02-01T19:02:37.378255900Z",
  "votes": 49,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Update :</strong> Code is publicly available here : <a href=\"https://github.com/TheoViel/kaggle_otto_rs\" target=\"_blank\">https://github.com/TheoViel/kaggle_otto_rs</a></p>\n<p>Thanks everyone for the cool competition. Huge congratz to the handful of people who could reach (or beat!) 0.603, this was a far from easy task. I personnally could not have done it alone, but I had the chance to team-up with amazing people who had a lot more experience in RecSys than me, namely <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> &amp; <a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">@benediktschifferer</a> !</p>\n<h3>Overview</h3>\n<p>My pipeline follows the classical candidates extraction &amp; reranker scheme, I will focus on the interesting stuff.</p>\n<p>CV scores are as follow : </p>\n<blockquote>\n  <p>cv = 0.5917  -  [0.5621, 0.4438, 0.6706]  -&gt;  <strong>LB 0.6028</strong></p>\n</blockquote>\n<p>Clicks is single model, I blend a few XGBs for carts &amp; orders but the boost is small. Blending with models from my teammates gave our Public 0.604 / Private (high) 0.603 LB !</p>\n<p>I use the candidates from Chris (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013\" target=\"_blank\">link</a>), as well as a slightly modified version of the ones from his public kernel. This results in approx. 80 candidates per sessions.</p>\n<h3>Feature engineering</h3>\n<p>Most of my (744) features come from the following process :</p>\n<ul>\n<li>Compute item-item scores (such as w2v similarities, matrix factorization similarity, Chris' covisitation matrices coefficients) between the candidate and items in the session</li>\n<li>Compute a weight adding information about to the item position in the session, timestamp, and type</li>\n<li>Aggregate !<br>\n<a href=\"https://postimg.cc/0rJVdzLf\" target=\"_blank\"><img src=\"https://i.postimg.cc/44WjGtPr/otto-de.png\" alt=\"otto-de\"></a></li>\n</ul>\n<p>Features are computed per batch on a 32Gb V100 using RAPIDS. It's fast :) </p>\n<h3>Extra session trick</h3>\n<p>An idea that gave a nice (0.0005 to 0.001) boost to my carts &amp; orders models is to retrieve sessions that were initially removed from the hosts' train/test split strategy and use them for training.</p>\n<p><a href=\"https://postimg.cc/0KGChBRC\" target=\"_blank\"><img src=\"https://i.postimg.cc/sxq0QkTb/otto-sess.png\" alt=\"otto-sess\"></a></p>\n<h3>Overall pipeline</h3>\n<p>I tune an Optuna for each fold (which is not a good practice, but I had a really reliable CV setup), pipeline can be a bit long to run but actually, the bottleneck is reading huge parquet files. Heavy downsampling makes it possible to have everything in RAM, and to train on GPU using the tricks Chris shared publicly. </p>\n<p><a href=\"https://postimg.cc/XZJ8Zqf9\" target=\"_blank\"><img src=\"https://i.postimg.cc/6QV1mGTb/otto-pipe.png\" alt=\"otto-pipe\"></a> </p>\n<p>Sadly, I could not get a lightgbm reranker to match my XGB performance and did not try Catboost, so no free boost from blending diverse models for me.</p>\n<p>Thanks for reading ! :)</p>",
  "messages": [
    {
      "id": "2125643",
      "postDate": "02/01/2023 19:02:37",
      "content": "<p><strong>Update :</strong> Code is publicly available here : <a href=\"https://github.com/TheoViel/kaggle_otto_rs\" target=\"_blank\">https://github.com/TheoViel/kaggle_otto_rs</a></p>\n<p>Thanks everyone for the cool competition. Huge congratz to the handful of people who could reach (or beat!) 0.603, this was a far from easy task. I personnally could not have done it alone, but I had the chance to team-up with amazing people who had a lot more experience in RecSys than me, namely <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> &amp; <a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">@benediktschifferer</a> !</p>\n<h3>Overview</h3>\n<p>My pipeline follows the classical candidates extraction &amp; reranker scheme, I will focus on the interesting stuff.</p>\n<p>CV scores are as follow : </p>\n<blockquote>\n  <p>cv = 0.5917  -  [0.5621, 0.4438, 0.6706]  -&gt;  <strong>LB 0.6028</strong></p>\n</blockquote>\n<p>Clicks is single model, I blend a few XGBs for carts &amp; orders but the boost is small. Blending with models from my teammates gave our Public 0.604 / Private (high) 0.603 LB !</p>\n<p>I use the candidates from Chris (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013\" target=\"_blank\">link</a>), as well as a slightly modified version of the ones from his public kernel. This results in approx. 80 candidates per sessions.</p>\n<h3>Feature engineering</h3>\n<p>Most of my (744) features come from the following process :</p>\n<ul>\n<li>Compute item-item scores (such as w2v similarities, matrix factorization similarity, Chris' covisitation matrices coefficients) between the candidate and items in the session</li>\n<li>Compute a weight adding information about to the item position in the session, timestamp, and type</li>\n<li>Aggregate !<br>\n<a href=\"https://postimg.cc/0rJVdzLf\" target=\"_blank\"><img src=\"https://i.postimg.cc/44WjGtPr/otto-de.png\" alt=\"otto-de\"></a></li>\n</ul>\n<p>Features are computed per batch on a 32Gb V100 using RAPIDS. It's fast :) </p>\n<h3>Extra session trick</h3>\n<p>An idea that gave a nice (0.0005 to 0.001) boost to my carts &amp; orders models is to retrieve sessions that were initially removed from the hosts' train/test split strategy and use them for training.</p>\n<p><a href=\"https://postimg.cc/0KGChBRC\" target=\"_blank\"><img src=\"https://i.postimg.cc/sxq0QkTb/otto-sess.png\" alt=\"otto-sess\"></a></p>\n<h3>Overall pipeline</h3>\n<p>I tune an Optuna for each fold (which is not a good practice, but I had a really reliable CV setup), pipeline can be a bit long to run but actually, the bottleneck is reading huge parquet files. Heavy downsampling makes it possible to have everything in RAM, and to train on GPU using the tricks Chris shared publicly. </p>\n<p><a href=\"https://postimg.cc/XZJ8Zqf9\" target=\"_blank\"><img src=\"https://i.postimg.cc/6QV1mGTb/otto-pipe.png\" alt=\"otto-pipe\"></a> </p>\n<p>Sadly, I could not get a lightgbm reranker to match my XGB performance and did not try Catboost, so no free boost from blending diverse models for me.</p>\n<p>Thanks for reading ! :)</p>",
      "rawMarkdown": "**Update :** Code is publicly available here : https://github.com/TheoViel/kaggle_otto_rs\n\nThanks everyone for the cool competition. Huge congratz to the handful of people who could reach (or beat!) 0.603, this was a far from easy task. I personnally could not have done it alone, but I had the chance to team-up with amazing people who had a lot more experience in RecSys than me, namely @cdeotte @titericz & @benediktschifferer !\n\n### Overview\n\nMy pipeline follows the classical candidates extraction & reranker scheme, I will focus on the interesting stuff.\n\nCV scores are as follow : \n> cv = 0.5917  -  [0.5621, 0.4438, 0.6706]  ->  **LB 0.6028**\n\nClicks is single model, I blend a few XGBs for carts & orders but the boost is small. Blending with models from my teammates gave our Public 0.604 / Private (high) 0.603 LB !\n\nI use the candidates from Chris ([link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013)), as well as a slightly modified version of the ones from his public kernel. This results in approx. 80 candidates per sessions.\n\n### Feature engineering\n\nMost of my (744) features come from the following process :\n- Compute item-item scores (such as w2v similarities, matrix factorization similarity, Chris' covisitation matrices coefficients) between the candidate and items in the session\n- Compute a weight adding information about to the item position in the session, timestamp, and type\n- Aggregate !\n<a href=\"https://postimg.cc/0rJVdzLf\" target=\"_blank\"><img src=\"https://i.postimg.cc/44WjGtPr/otto-de.png\" alt=\"otto-de\"/></a>\n\nFeatures are computed per batch on a 32Gb V100 using RAPIDS. It's fast :) \n\n### Extra session trick\n\nAn idea that gave a nice (0.0005 to 0.001) boost to my carts & orders models is to retrieve sessions that were initially removed from the hosts' train/test split strategy and use them for training.\n\n<a href='https://postimg.cc/0KGChBRC' target='_blank'><img src='https://i.postimg.cc/sxq0QkTb/otto-sess.png' border='0' alt='otto-sess'/></a>\n\n### Overall pipeline\n\nI tune an Optuna for each fold (which is not a good practice, but I had a really reliable CV setup), pipeline can be a bit long to run but actually, the bottleneck is reading huge parquet files. Heavy downsampling makes it possible to have everything in RAM, and to train on GPU using the tricks Chris shared publicly. \n\n<a href=\"https://postimg.cc/XZJ8Zqf9\" target=\"_blank\"><img src=\"https://i.postimg.cc/6QV1mGTb/otto-pipe.png\" alt=\"otto-pipe\"/></a> \n\nSadly, I could not get a lightgbm reranker to match my XGB performance and did not try Catboost, so no free boost from blending diverse models for me.\n\nThanks for reading ! :)",
      "votes": null
    },
    {
      "id": "2136772",
      "postDate": "02/09/2023 14:36:26",
      "content": "<p>Congratulations, thanks for sharing. Really elegant code. I learned a lot from your code. Thanks</p>",
      "rawMarkdown": "Congratulations, thanks for sharing. Really elegant code. I learned a lot from your code. Thanks",
      "votes": null
    },
    {
      "id": "2166950",
      "postDate": "03/03/2023 06:34:26",
      "content": "<p>Thank you so much for your sharing! <br>\nMay I ask more about why function <code>compute_popularities_new</code> only has <code>day_map</code> for 8 days and not for the whole period?</p>",
      "rawMarkdown": "Thank you so much for your sharing! \nMay I ask more about why function `compute_popularities_new` only has `day_map` for 8 days and not for the whole period?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2136772,
      "author_name": "yasso1",
      "author_url": "",
      "post_date": "02/09/2023 14:36:26",
      "content": "<p>Congratulations, thanks for sharing. Really elegant code. I learned a lot from your code. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2166950,
      "author_name": "ngcuynnguyn",
      "author_url": "",
      "post_date": "03/03/2023 06:34:26",
      "content": "<p>Thank you so much for your sharing! <br>\nMay I ask more about why function <code>compute_popularities_new</code> only has <code>day_map</code> for 8 days and not for the whole period?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2125643": "**Update :** Code is publicly available here : https://github.com/TheoViel/kaggle_otto_rs\n\nThanks everyone for the cool competition. Huge congratz to the handful of people who could reach (or beat!) 0.603, this was a far from easy task. I personnally could not have done it alone, but I had the chance to team-up with amazing people who had a lot more experience in RecSys than me, namely @cdeotte @titericz & @benediktschifferer !\n\n### Overview\n\nMy pipeline follows the classical candidates extraction & reranker scheme, I will focus on the interesting stuff.\n\nCV scores are as follow : \n> cv = 0.5917  -  [0.5621, 0.4438, 0.6706]  ->  **LB 0.6028**\n\nClicks is single model, I blend a few XGBs for carts & orders but the boost is small. Blending with models from my teammates gave our Public 0.604 / Private (high) 0.603 LB !\n\nI use the candidates from Chris ([link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013)), as well as a slightly modified version of the ones from his public kernel. This results in approx. 80 candidates per sessions.\n\n### Feature engineering\n\nMost of my (744) features come from the following process :\n- Compute item-item scores (such as w2v similarities, matrix factorization similarity, Chris' covisitation matrices coefficients) between the candidate and items in the session\n- Compute a weight adding information about to the item position in the session, timestamp, and type\n- Aggregate !\n<a href=\"https://postimg.cc/0rJVdzLf\" target=\"_blank\"><img src=\"https://i.postimg.cc/44WjGtPr/otto-de.png\" alt=\"otto-de\"/></a>\n\nFeatures are computed per batch on a 32Gb V100 using RAPIDS. It's fast :) \n\n### Extra session trick\n\nAn idea that gave a nice (0.0005 to 0.001) boost to my carts & orders models is to retrieve sessions that were initially removed from the hosts' train/test split strategy and use them for training.\n\n<a href='https://postimg.cc/0KGChBRC' target='_blank'><img src='https://i.postimg.cc/sxq0QkTb/otto-sess.png' border='0' alt='otto-sess'/></a>\n\n### Overall pipeline\n\nI tune an Optuna for each fold (which is not a good practice, but I had a really reliable CV setup), pipeline can be a bit long to run but actually, the bottleneck is reading huge parquet files. Heavy downsampling makes it possible to have everything in RAM, and to train on GPU using the tricks Chris shared publicly. \n\n<a href=\"https://postimg.cc/XZJ8Zqf9\" target=\"_blank\"><img src=\"https://i.postimg.cc/6QV1mGTb/otto-pipe.png\" alt=\"otto-pipe\"/></a> \n\nSadly, I could not get a lightgbm reranker to match my XGB performance and did not try Catboost, so no free boost from blending diverse models for me.\n\nThanks for reading ! :)",
    "2136772": "Congratulations, thanks for sharing. Really elegant code. I learned a lot from your code. Thanks",
    "2166950": "Thank you so much for your sharing! \nMay I ask more about why function `compute_popularities_new` only has `day_map` for 8 days and not for the whole period?"
  },
  "source": "meta"
}