{
  "id": 383843,
  "title": "34th (ex 37th) Place Solution (Polars is here to stay !)",
  "url": "/competitions/otto-recommender-system/writeups/steubk-34th-ex-37th-place-solution-polars-is-here-",
  "author_name": "",
  "post_date": "2023-02-08T07:52:57.667Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thank you to Otto and Kaggle for hosting this competition with such a challenging dataset.<br>\nA special thank to kagglers who shared their work making the competition even more stimulating  and - as usual - congratulations to the winners and everyone who enjoied the competition !</p>\n<p>This competition introduced me to the Marlin daloader - a super efficient NVIDIA <a href=\"https://github.com/NVIDIA-Merlin/dataloader\" target=\"_blank\">dataloder </a> for recommender systems (thank you <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> !) and <a href=\"https://www.pola.rs/\" target=\"_blank\">Polars</a>:  a blazingly fast DataFrame library for huge datasets that natively supports multithreading and whose syntax is - for me - more intuitive than pandas. During the competition I was able to easily replace all the pandas code with polars code with a significant gain in performance (RAM and  CPU).</p>\n<p>I think Polars it's a library that will become more and more important in the data ecosystem.</p>\n<h2>My solution</h2>\n<p>This is the visualization of the steps of my solution:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fb0b946b4770066127596d3b62d326f34%2Fjourney.png?generation=1675611218130910&amp;alt=media\" alt=\"\"></p>\n<h3>Validation</h3>\n<p>1_000_000 random truncated sessions out of 1_801_251 sessions from 4th week</p>\n<h3>Candidates Selection (Heuristic Model) (val:0.570, test:0.576)</h3>\n<p>Generated covisitation matrices for all aids, carts and orders and computed probability <strong>P(aid,next-aid)</strong> = <strong>next-aid</strong> follow <strong>aid</strong> in a (window of a) session containing <strong>aid</strong>  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F92de3828cb209e69ad8227cce705dabe%2Fprobabilities.png?generation=1675611468571181&amp;alt=media\" alt=\"\"></p>\n<p>Calculated the probability that an aid,cart,order is present more than once in a (window of a) session.</p>\n<p>Heuristic model uses rules based on these two kind of probabilities.</p>\n<h3>Base Ranker Model (0.573,0.580)</h3>\n<p>Selected first 100 candidates by heuristic model and generated 26 features:  </p>\n<p>10 interaction (between candidate and session) based features of wich the most significative are: </p>\n<ul>\n<li><em>i_count</em>: occurences of candidate in session</li>\n<li><em>i_self_aids_d</em>: sum of probabilities that the candidate is present more than once in the session </li>\n<li><em>i_sims_aids</em>: sum of probabilities that the candidate is a next-aid in the session</li>\n<li>…</li>\n</ul>\n<p>10 session based features : </p>\n<ul>\n<li><em>s_type_mean</em>: mean of clicks , cart and buys (see <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/379631#2108619\" target=\"_blank\">link</a>)</li>\n<li>…</li>\n</ul>\n<p>6 candidated based features : </p>\n<ul>\n<li><em>a_self_buy_total</em>: occurences of orders for  candidate in all sessions  </li>\n<li>…</li>\n</ul>\n<h3>Feature engineering 1 (0.581, 0.588)</h3>\n<p>Added custom features for last 5 aids,carts,orders of each session (for a total of 81 feature):</p>\n<ul>\n<li><em>i_aids_count_last_aid</em>: value of covisitation matrix  (last-aid-session,candidate)  </li>\n<li><em>i_aids_total_last_aid</em>: occurence of last-aid in all sessions </li>\n<li><em>i_aids_count_last_aid/i_aids_total_last_aid</em>: probability that candidate follows last-aid in all sessions</li>\n<li>… </li>\n</ul>\n<h3>Feature engineering 2: 0.583, 0.591</h3>\n<p>Generated covisitation matrices based on last two weeks (validation + test) and added corresponding features (137 features)</p>\n<h3>Stacking : 0.585, 0.592</h3>\n<p>Selected best 50 candidate from previous best model, added cross stacked predictions  from previous lgb models.<br>\nAdded some interaction features based on word2vec model and used xgb as stacked model. </p>",
  "messages": [
    {
      "id": "2130661",
      "postDate": "02/05/2023 15:59:42",
      "content": "<p>Thank you to Otto and Kaggle for hosting this competition with such a challenging dataset.<br>\nA special thank to kagglers who shared their work making the competition even more stimulating  and - as usual - congratulations to the winners and everyone who enjoied the competition !</p>\n<p>This competition introduced me to the Marlin daloader - a super efficient NVIDIA <a href=\"https://github.com/NVIDIA-Merlin/dataloader\" target=\"_blank\">dataloder </a> for recommender systems (thank you <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> !) and <a href=\"https://www.pola.rs/\" target=\"_blank\">Polars</a>:  a blazingly fast DataFrame library for huge datasets that natively supports multithreading and whose syntax is - for me - more intuitive than pandas. During the competition I was able to easily replace all the pandas code with polars code with a significant gain in performance (RAM and  CPU).</p>\n<p>I think Polars it's a library that will become more and more important in the data ecosystem.</p>\n<h2>My solution</h2>\n<p>This is the visualization of the steps of my solution:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fb0b946b4770066127596d3b62d326f34%2Fjourney.png?generation=1675611218130910&amp;alt=media\" alt=\"\"></p>\n<h3>Validation</h3>\n<p>1_000_000 random truncated sessions out of 1_801_251 sessions from 4th week</p>\n<h3>Candidates Selection (Heuristic Model) (val:0.570, test:0.576)</h3>\n<p>Generated covisitation matrices for all aids, carts and orders and computed probability <strong>P(aid,next-aid)</strong> = <strong>next-aid</strong> follow <strong>aid</strong> in a (window of a) session containing <strong>aid</strong>  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F92de3828cb209e69ad8227cce705dabe%2Fprobabilities.png?generation=1675611468571181&amp;alt=media\" alt=\"\"></p>\n<p>Calculated the probability that an aid,cart,order is present more than once in a (window of a) session.</p>\n<p>Heuristic model uses rules based on these two kind of probabilities.</p>\n<h3>Base Ranker Model (0.573,0.580)</h3>\n<p>Selected first 100 candidates by heuristic model and generated 26 features:  </p>\n<p>10 interaction (between candidate and session) based features of wich the most significative are: </p>\n<ul>\n<li><em>i_count</em>: occurences of candidate in session</li>\n<li><em>i_self_aids_d</em>: sum of probabilities that the candidate is present more than once in the session </li>\n<li><em>i_sims_aids</em>: sum of probabilities that the candidate is a next-aid in the session</li>\n<li>…</li>\n</ul>\n<p>10 session based features : </p>\n<ul>\n<li><em>s_type_mean</em>: mean of clicks , cart and buys (see <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/379631#2108619\" target=\"_blank\">link</a>)</li>\n<li>…</li>\n</ul>\n<p>6 candidated based features : </p>\n<ul>\n<li><em>a_self_buy_total</em>: occurences of orders for  candidate in all sessions  </li>\n<li>…</li>\n</ul>\n<h3>Feature engineering 1 (0.581, 0.588)</h3>\n<p>Added custom features for last 5 aids,carts,orders of each session (for a total of 81 feature):</p>\n<ul>\n<li><em>i_aids_count_last_aid</em>: value of covisitation matrix  (last-aid-session,candidate)  </li>\n<li><em>i_aids_total_last_aid</em>: occurence of last-aid in all sessions </li>\n<li><em>i_aids_count_last_aid/i_aids_total_last_aid</em>: probability that candidate follows last-aid in all sessions</li>\n<li>… </li>\n</ul>\n<h3>Feature engineering 2: 0.583, 0.591</h3>\n<p>Generated covisitation matrices based on last two weeks (validation + test) and added corresponding features (137 features)</p>\n<h3>Stacking : 0.585, 0.592</h3>\n<p>Selected best 50 candidate from previous best model, added cross stacked predictions  from previous lgb models.<br>\nAdded some interaction features based on word2vec model and used xgb as stacked model. </p>",
      "rawMarkdown": "Thank you to Otto and Kaggle for hosting this competition with such a challenging dataset.\nA special thank to kagglers who shared their work making the competition even more stimulating  and - as usual - congratulations to the winners and everyone who enjoied the competition !\n\nThis competition introduced me to the Marlin daloader - a super efficient NVIDIA [dataloder ](https://github.com/NVIDIA-Merlin/dataloader) for recommender systems (thank you @radek1 !) and [Polars](https://www.pola.rs/):  a blazingly fast DataFrame library for huge datasets that natively supports multithreading and whose syntax is - for me - more intuitive than pandas. During the competition I was able to easily replace all the pandas code with polars code with a significant gain in performance (RAM and  CPU).\n\nI think Polars it's a library that will become more and more important in the data ecosystem.\n\n## My solution \n\nThis is the visualization of the steps of my solution:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fb0b946b4770066127596d3b62d326f34%2Fjourney.png?generation=1675611218130910&alt=media)\n\n### Validation\n\n1_000_000 random truncated sessions out of 1_801_251 sessions from 4th week\n\n### Candidates Selection (Heuristic Model) (val:0.570, test:0.576)\n\nGenerated covisitation matrices for all aids, carts and orders and computed probability **P(aid,next-aid)** = **next-aid** follow **aid** in a (window of a) session containing **aid**  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F92de3828cb209e69ad8227cce705dabe%2Fprobabilities.png?generation=1675611468571181&alt=media)\n\nCalculated the probability that an aid,cart,order is present more than once in a (window of a) session.\n\nHeuristic model uses rules based on these two kind of probabilities.\n\n### Base Ranker Model (0.573,0.580)\n\nSelected first 100 candidates by heuristic model and generated 26 features:  \n\n10 interaction (between candidate and session) based features of wich the most significative are: \n- *i_count*: occurences of candidate in session\n- *i_self_aids_d*: sum of probabilities that the candidate is present more than once in the session \n- *i_sims_aids*: sum of probabilities that the candidate is a next-aid in the session\n- ...\n\n10 session based features : \n- *s_type_mean*: mean of clicks , cart and buys (see [link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/379631#2108619))\n- ...\n\n6 candidated based features : \n- *a_self_buy_total*: occurences of orders for  candidate in all sessions  \n- ...\n\n### Feature engineering 1 (0.581, 0.588)\n\nAdded custom features for last 5 aids,carts,orders of each session (for a total of 81 feature):\n- *i_aids_count_last_aid*: value of covisitation matrix  (last-aid-session,candidate)  \n- *i_aids_total_last_aid*: occurence of last-aid in all sessions \n- *i_aids_count_last_aid/i_aids_total_last_aid*: probability that candidate follows last-aid in all sessions\n- ... \n\n### Feature engineering 2: 0.583, 0.591\n\nGenerated covisitation matrices based on last two weeks (validation + test) and added corresponding features (137 features)\n\n### Stacking : 0.585, 0.592 \nSelected best 50 candidate from previous best model, added cross stacked predictions  from previous lgb models.\nAdded some interaction features based on word2vec model and used xgb as stacked model.",
      "votes": null
    },
    {
      "id": "2132187",
      "postDate": "02/06/2023 16:45:40",
      "content": "<p>Awesome, thanks for sharing. Simple and effective solution 🎉</p>\n<p>By any chance, are you planning on sharing your code? </p>",
      "rawMarkdown": "Awesome, thanks for sharing. Simple and effective solution 🎉\n\nBy any chance, are you planning on sharing your code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2132187,
      "author_name": "parthpankajtiwary",
      "author_url": "",
      "post_date": "02/06/2023 16:45:40",
      "content": "<p>Awesome, thanks for sharing. Simple and effective solution 🎉</p>\n<p>By any chance, are you planning on sharing your code? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2130661": "Thank you to Otto and Kaggle for hosting this competition with such a challenging dataset.\nA special thank to kagglers who shared their work making the competition even more stimulating  and - as usual - congratulations to the winners and everyone who enjoied the competition !\n\nThis competition introduced me to the Marlin daloader - a super efficient NVIDIA [dataloder ](https://github.com/NVIDIA-Merlin/dataloader) for recommender systems (thank you @radek1 !) and [Polars](https://www.pola.rs/):  a blazingly fast DataFrame library for huge datasets that natively supports multithreading and whose syntax is - for me - more intuitive than pandas. During the competition I was able to easily replace all the pandas code with polars code with a significant gain in performance (RAM and  CPU).\n\nI think Polars it's a library that will become more and more important in the data ecosystem.\n\n## My solution \n\nThis is the visualization of the steps of my solution:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fb0b946b4770066127596d3b62d326f34%2Fjourney.png?generation=1675611218130910&alt=media)\n\n### Validation\n\n1_000_000 random truncated sessions out of 1_801_251 sessions from 4th week\n\n### Candidates Selection (Heuristic Model) (val:0.570, test:0.576)\n\nGenerated covisitation matrices for all aids, carts and orders and computed probability **P(aid,next-aid)** = **next-aid** follow **aid** in a (window of a) session containing **aid**  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F92de3828cb209e69ad8227cce705dabe%2Fprobabilities.png?generation=1675611468571181&alt=media)\n\nCalculated the probability that an aid,cart,order is present more than once in a (window of a) session.\n\nHeuristic model uses rules based on these two kind of probabilities.\n\n### Base Ranker Model (0.573,0.580)\n\nSelected first 100 candidates by heuristic model and generated 26 features:  \n\n10 interaction (between candidate and session) based features of wich the most significative are: \n- *i_count*: occurences of candidate in session\n- *i_self_aids_d*: sum of probabilities that the candidate is present more than once in the session \n- *i_sims_aids*: sum of probabilities that the candidate is a next-aid in the session\n- ...\n\n10 session based features : \n- *s_type_mean*: mean of clicks , cart and buys (see [link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/379631#2108619))\n- ...\n\n6 candidated based features : \n- *a_self_buy_total*: occurences of orders for  candidate in all sessions  \n- ...\n\n### Feature engineering 1 (0.581, 0.588)\n\nAdded custom features for last 5 aids,carts,orders of each session (for a total of 81 feature):\n- *i_aids_count_last_aid*: value of covisitation matrix  (last-aid-session,candidate)  \n- *i_aids_total_last_aid*: occurence of last-aid in all sessions \n- *i_aids_count_last_aid/i_aids_total_last_aid*: probability that candidate follows last-aid in all sessions\n- ... \n\n### Feature engineering 2: 0.583, 0.591\n\nGenerated covisitation matrices based on last two weeks (validation + test) and added corresponding features (137 features)\n\n### Stacking : 0.585, 0.592 \nSelected best 50 candidate from previous best model, added cross stacked predictions  from previous lgb models.\nAdded some interaction features based on word2vec model and used xgb as stacked model.",
    "2132187": "Awesome, thanks for sharing. Simple and effective solution 🎉\n\nBy any chance, are you planning on sharing your code?"
  },
  "source": "meta"
}