{
  "id": 382839,
  "title": "2nd Place Solution(senkin13&30CrMnSiA part)",
  "url": "/competitions/otto-recommender-system/discussion/382839",
  "author_name": "",
  "post_date": "2023-02-01T07:32:18.177913Z",
  "votes": 69,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Thanks to OTTO and Kaggle Team giving us such a wonderful competition very close to real business.Thanks my team mates <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> ,we did many creative work and work hard util the last hours before end.<br>\nI and 30CrMnSiA won the H&amp;M Personalized Fashion Recommendations last year, we think we can start from midterm,but we found this competition is very hard and very different from H&amp;M competition,so we merge with silonodera team for last 18 days.I am satisfied with the progress and result.</p>\n<h1>Overview</h1>\n<p><a href=\"https://postimg.cc/cgfnMFg0\" target=\"_blank\"><img src=\"https://i.postimg.cc/C1Qs0twB/Screenshot-2023-02-01-at-15-43-31.png\" alt=\"Screenshot-2023-02-01-at-15-43-31.png\"></a></p>\n<h1>Retrieval(aka: candidate generation | recall)</h1>\n<p>This competition data has no user/item information, and popularity don't work, hard to improve retrieval, we found only repeat action,next action and itemcf work,there are nothing to do for repeat action,next action,the most important work is to improve itemcf, we tried kinds of matrix and weight for itemcf and worked.<br>\n[matrix]</p>\n<ul>\n<li>cart-order </li>\n<li>click-cart&amp;order </li>\n<li>click-cart-order </li>\n</ul>\n<p>[weight]</p>\n<ul>\n<li>type </li>\n<li>co-occurrence time distance </li>\n<li>co-occurrence order</li>\n<li>session action order </li>\n<li>popularity </li>\n</ul>\n<h1>Ranking</h1>\n<ul>\n<li><strong>Feature Engineering</strong><br>\nBasiclly,features are created base on retrieval strategies,create user-item interaction for repeat action,next action,aggregated collaborative filtering score from itemcf,similarity for embedding retrieval.<br>\nI used 198 features for order/cart model, 106 features for click model</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>aggregation</td>\n<td>count of session,aid,session-aid,each type aggregated by current session</td>\n</tr>\n<tr>\n<td>next action</td>\n<td>count of aid which is the next of session's last aid</td>\n</tr>\n<tr>\n<td>time</td>\n<td>first,last time of session,aid</td>\n</tr>\n<tr>\n<td>collaborative filtering score</td>\n<td>max,sum,weighted sum of session's last aid,last hour aids,all aids &amp; candidate aid score</td>\n</tr>\n<tr>\n<td>embedding similarity</td>\n<td>cosine similarity of aid2aid(word2vec), cosine similarity of session2aid(ProNE)</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><strong>Validation</strong><br>\nI made 3 weeks' train data with labels, use last week as validation,then retrain all train data with fixed round number for test data prediction,the CV matched LB very well, we don't need to watch LB by submission.</p></li>\n<li><p><strong>Model</strong><br>\nAt H&amp;M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.<br>\nMy best single model is a catboost ranker, CV 0.59066 = 0.6710 * 0.6 + 0.4416 * 0.3 + 0.5558 * 0.1, LB 0.602 = 0.409 + 0.137 + 0.056.<br>\nFinally I blend with lightgbm and catboost,the blend LB is higher 0.602, after blend with teammate's 0.602,LB is 0.604</p></li>\n</ul>\n<h1>Optimization</h1>\n<p>At last days I rewrite all my feature engineering code from pandas to polars,the polars is much faster,especially two huge dataframe join, pandas.merge -&gt; polars.join = 40X faster.<br>\nI use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.</p>",
  "messages": [
    {
      "id": "2124741",
      "postDate": "02/01/2023 07:32:18",
      "content": "<p>Thanks to OTTO and Kaggle Team giving us such a wonderful competition very close to real business.Thanks my team mates <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> ,we did many creative work and work hard util the last hours before end.<br>\nI and 30CrMnSiA won the H&amp;M Personalized Fashion Recommendations last year, we think we can start from midterm,but we found this competition is very hard and very different from H&amp;M competition,so we merge with silonodera team for last 18 days.I am satisfied with the progress and result.</p>\n<h1>Overview</h1>\n<p><a href=\"https://postimg.cc/cgfnMFg0\" target=\"_blank\"><img src=\"https://i.postimg.cc/C1Qs0twB/Screenshot-2023-02-01-at-15-43-31.png\" alt=\"Screenshot-2023-02-01-at-15-43-31.png\"></a></p>\n<h1>Retrieval(aka: candidate generation | recall)</h1>\n<p>This competition data has no user/item information, and popularity don't work, hard to improve retrieval, we found only repeat action,next action and itemcf work,there are nothing to do for repeat action,next action,the most important work is to improve itemcf, we tried kinds of matrix and weight for itemcf and worked.<br>\n[matrix]</p>\n<ul>\n<li>cart-order </li>\n<li>click-cart&amp;order </li>\n<li>click-cart-order </li>\n</ul>\n<p>[weight]</p>\n<ul>\n<li>type </li>\n<li>co-occurrence time distance </li>\n<li>co-occurrence order</li>\n<li>session action order </li>\n<li>popularity </li>\n</ul>\n<h1>Ranking</h1>\n<ul>\n<li><strong>Feature Engineering</strong><br>\nBasiclly,features are created base on retrieval strategies,create user-item interaction for repeat action,next action,aggregated collaborative filtering score from itemcf,similarity for embedding retrieval.<br>\nI used 198 features for order/cart model, 106 features for click model</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>aggregation</td>\n<td>count of session,aid,session-aid,each type aggregated by current session</td>\n</tr>\n<tr>\n<td>next action</td>\n<td>count of aid which is the next of session's last aid</td>\n</tr>\n<tr>\n<td>time</td>\n<td>first,last time of session,aid</td>\n</tr>\n<tr>\n<td>collaborative filtering score</td>\n<td>max,sum,weighted sum of session's last aid,last hour aids,all aids &amp; candidate aid score</td>\n</tr>\n<tr>\n<td>embedding similarity</td>\n<td>cosine similarity of aid2aid(word2vec), cosine similarity of session2aid(ProNE)</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><strong>Validation</strong><br>\nI made 3 weeks' train data with labels, use last week as validation,then retrain all train data with fixed round number for test data prediction,the CV matched LB very well, we don't need to watch LB by submission.</p></li>\n<li><p><strong>Model</strong><br>\nAt H&amp;M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.<br>\nMy best single model is a catboost ranker, CV 0.59066 = 0.6710 * 0.6 + 0.4416 * 0.3 + 0.5558 * 0.1, LB 0.602 = 0.409 + 0.137 + 0.056.<br>\nFinally I blend with lightgbm and catboost,the blend LB is higher 0.602, after blend with teammate's 0.602,LB is 0.604</p></li>\n</ul>\n<h1>Optimization</h1>\n<p>At last days I rewrite all my feature engineering code from pandas to polars,the polars is much faster,especially two huge dataframe join, pandas.merge -&gt; polars.join = 40X faster.<br>\nI use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.</p>",
      "rawMarkdown": "Thanks to OTTO and Kaggle Team giving us such a wonderful competition very close to real business.Thanks my team mates @h4211819 @onodera @psilogram ,we did many creative work and work hard util the last hours before end.\nI and 30CrMnSiA won the H&M Personalized Fashion Recommendations last year, we think we can start from midterm,but we found this competition is very hard and very different from H&M competition,so we merge with silonodera team for last 18 days.I am satisfied with the progress and result.\n\n# Overview\n\n[![Screenshot-2023-02-01-at-15-43-31.png](https://i.postimg.cc/C1Qs0twB/Screenshot-2023-02-01-at-15-43-31.png)](https://postimg.cc/cgfnMFg0)\n\n# Retrieval(aka: candidate generation | recall)\nThis competition data has no user/item information, and popularity don't work, hard to improve retrieval, we found only repeat action,next action and itemcf work,there are nothing to do for repeat action,next action,the most important work is to improve itemcf, we tried kinds of matrix and weight for itemcf and worked.\n[matrix]\n- cart-order \n- click-cart&order \n- click-cart-order \n\n[weight]\n- type \n- co-occurrence time distance \n- co-occurrence order\n- session action order \n- popularity \n\n# Ranking\n- **Feature Engineering**\nBasiclly,features are created base on retrieval strategies,create user-item interaction for repeat action,next action,aggregated collaborative filtering score from itemcf,similarity for embedding retrieval.\nI used 198 features for order/cart model, 106 features for click model\n\n| Type | Description|\n| --- | --- |\n| aggregation |  count of session,aid,session-aid,each type aggregated by current session|last week sessions|previous sessions |\n| next action |  count of aid which is the next of session's last aid|\n| time|  first,last time of session,aid|\n| collaborative filtering score| max,sum,weighted sum of session's last aid,last hour aids,all aids & candidate aid score|\n| embedding similarity| cosine similarity of aid2aid(word2vec), cosine similarity of session2aid(ProNE) |\n\n- **Validation**\nI made 3 weeks' train data with labels, use last week as validation,then retrain all train data with fixed round number for test data prediction,the CV matched LB very well, we don't need to watch LB by submission.\n\n- **Model**\nAt H&M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.\nMy best single model is a catboost ranker, CV 0.59066 = 0.6710 * 0.6 + 0.4416 * 0.3 + 0.5558 * 0.1, LB 0.602 = 0.409 + 0.137 + 0.056.\nFinally I blend with lightgbm and catboost,the blend LB is higher 0.602, after blend with teammate's 0.602,LB is 0.604\n\n# Optimization\nAt last days I rewrite all my feature engineering code from pandas to polars,the polars is much faster,especially two huge dataframe join, pandas.merge -> polars.join = 40X faster.\nI use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.",
      "votes": null
    },
    {
      "id": "2125673",
      "postDate": "02/01/2023 19:42:55",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> we learnt lot from you ideas from the H&amp;M comp . Could you give some details on <code>cosine similarity of session2aid(ProNE)</code> approach .. </p>",
      "rawMarkdown": "Thanks @senkin13 we learnt lot from you ideas from the H&M comp . Could you give some details on `cosine similarity of session2aid(ProNE)` approach ..",
      "votes": null
    },
    {
      "id": "2125911",
      "postDate": "02/02/2023 01:48:23",
      "content": "<p>Congrats on yet another impressive result!!! </p>\n<blockquote>\n  <p>we found only repeat action,next action</p>\n</blockquote>\n<p>I wonder how to calculate this? maybe you can get repeat action based on counts, but how about next action?</p>",
      "rawMarkdown": "Congrats on yet another impressive result!!! \n> we found only repeat action,next action\n\n\nI wonder how to calculate this? maybe you can get repeat action based on counts, but how about next action?",
      "votes": null
    },
    {
      "id": "2125914",
      "postDate": "02/02/2023 01:50:09",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and team. Great job winning top spots in both H&amp;M comp and Otto comp!</p>",
      "rawMarkdown": "Congratulations @senkin13 and team. Great job winning top spots in both H&M comp and Otto comp!",
      "votes": null
    },
    {
      "id": "2126059",
      "postDate": "02/02/2023 04:29:36",
      "content": "<p>let me show you the polars code<br>\ndf: candidates<br>\ntrain: previoud session + current session<br>\ntest: current session</p>\n<pre><code>   train = train.with_column(\n            (pl.col().shift(-).over()).alias()\n            )    \n    train = train.with_column(\n            (pl.col().shift(-).over()).alias()\n            )     \n\n    df = df.join(test\n                 .select([,,])\n                 .unique(subset=, keep=),\n            on=[], how=)  \n    df.columns = [,,,]     \n\n    df = df.join(train                 \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                 \n            ]\n        ), \n        on=[,,], how=)    \n\n    df = df.join(train    \n        .(\n        pl.col()==     \n        )                       \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                  \n            ]\n        ), \n        on=[,,], how=)    \n\n    df = df.join(train    \n        .(\n        pl.col()==    \n        )                       \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                  \n            ]\n        ), \n        on=[,,], how=)        \n\n    df = df.join(train    \n        .(\n        pl.col()==     \n        )                       \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                  \n            ]\n        ), \n        on=[,,], how=)\n</code></pre>",
      "rawMarkdown": "let me show you the polars code\ndf: candidates\ntrain: previoud session + current session\ntest: current session\n```python\n   train = train.with_column(\n            (pl.col('aid').shift(-1).over('session')).alias('aid_next')\n            )    \n    train = train.with_column(\n            (pl.col('type').shift(-1).over('session')).alias('type_next')\n            )     \n        \n    df = df.join(test\n                 .select(['session','aid','type'])\n                 .unique(subset='session', keep='last'),\n            on=['session'], how='left')  \n    df.columns = ['session','aid','aid_next','type']     \n    \n    df = df.join(train                 \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_count'),                 \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')    \n    \n    df = df.join(train    \n        .filter(\n        pl.col('type_next')==0     \n        )                       \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_click_count'),                  \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')    \n    \n    df = df.join(train    \n        .filter(\n        pl.col('type_next')==1    \n        )                       \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_cart_count'),                  \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')        \n    \n    df = df.join(train    \n        .filter(\n        pl.col('type_next')==2     \n        )                       \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_order_count'),                  \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')\n```",
      "votes": null
    },
    {
      "id": "2126065",
      "postDate": "02/02/2023 04:32:36",
      "content": "<p>treat session and aid both as a node of one graph, then train ProNE model to get graph embedding, finally calculate cosine similary of session embedding and aid embedding</p>",
      "rawMarkdown": "treat session and aid both as a node of one graph, then train ProNE model to get graph embedding, finally calculate cosine similary of session embedding and aid embedding",
      "votes": null
    },
    {
      "id": "2126123",
      "postDate": "02/02/2023 05:30:40",
      "content": "<p>Congratulations!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Congratulations!\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2127225",
      "postDate": "02/02/2023 18:55:54",
      "content": "<p>Congratulations !!! Great efforts</p>",
      "rawMarkdown": "Congratulations !!! Great efforts",
      "votes": null
    },
    {
      "id": "2127344",
      "postDate": "02/02/2023 20:23:12",
      "content": "<blockquote>\n  <p>At H&amp;M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.</p>\n</blockquote>\n<p>Wow that's a big boost. I just retrained my XGB model now using CatBoostRanker and <code>PairLogitPairwise</code> loss. The orders boost <code>+0.00063</code> and the carts boost <code>+0.00074</code>. Wow. During the competition i tried CatBoost with YetiRank loss but that did poorly. I didn't know that CatBoost had a pairwise loss until after the comp ended and read winners writeups.</p>",
      "rawMarkdown": ">At H&M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.\n\nWow that's a big boost. I just retrained my XGB model now using CatBoostRanker and `PairLogitPairwise` loss. The orders boost `+0.00063` and the carts boost `+0.00074`. Wow. During the competition i tried CatBoost with YetiRank loss but that did poorly. I didn't know that CatBoost had a pairwise loss until after the comp ended and read winners writeups.",
      "votes": null
    },
    {
      "id": "2127387",
      "postDate": "02/02/2023 21:11:05",
      "content": "<p>For Catboost we also tried <a href=\"https://catboost.ai/en/docs/references/querycrossentropy\" target=\"_blank\">https://catboost.ai/en/docs/references/querycrossentropy</a> loss and it worked better for us .. more than pariwise logit or yeti . Not sure top teams also tried this ?</p>",
      "rawMarkdown": "For Catboost we also tried https://catboost.ai/en/docs/references/querycrossentropy loss and it worked better for us .. more than pariwise logit or yeti . Not sure top teams also tried this ?",
      "votes": null
    },
    {
      "id": "2127388",
      "postDate": "02/02/2023 21:12:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> will keep this in mind for the future, we did generate GNN embeddings but added them directly which didn't work that well .great to learn . </p>",
      "rawMarkdown": "Thanks @senkin13 will keep this in mind for the future, we did generate GNN embeddings but added them directly which didn't work that well .great to learn .",
      "votes": null
    },
    {
      "id": "2127389",
      "postDate": "02/02/2023 21:12:44",
      "content": "<p>Interesting. I'll give it a try and report results here</p>",
      "rawMarkdown": "Interesting. I'll give it a try and report results here",
      "votes": null
    },
    {
      "id": "2127390",
      "postDate": "02/02/2023 21:13:38",
      "content": "<p>I just submitted my CatBoost (with <code>PairLogitPairwise</code>) to LB and it boost LB <code>+0.0006</code> too (compared with XGB) for my single model. Big boost.</p>",
      "rawMarkdown": "I just submitted my CatBoost (with `PairLogitPairwise`) to LB and it boost LB `+0.0006` too (compared with XGB) for my single model. Big boost.",
      "votes": null
    },
    {
      "id": "2127419",
      "postDate": "02/02/2023 21:46:41",
      "content": "<p>sure will be good to know . also it works only on GPU and I think its like a group Log loss .We also had an idea to implement this for some NN rankers but we didnt get the time .</p>",
      "rawMarkdown": "sure will be good to know . also it works only on GPU and I think its like a group Log loss .We also had an idea to implement this for some NN rankers but we didnt get the time .",
      "votes": null
    },
    {
      "id": "2127583",
      "postDate": "02/03/2023 01:27:57",
      "content": "<p>Congratulations !! <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> </p>",
      "rawMarkdown": "Congratulations !! @senkin13",
      "votes": null
    },
    {
      "id": "2127633",
      "postDate": "02/03/2023 02:23:57",
      "content": "<p>I think PairLogitPairwise of catboost's ranking loss fuction is strong, and my lightgbm classifier used variable candidators for each session, there will be more bias for each session,  the ranking model has less bias per session.</p>",
      "rawMarkdown": "I think PairLogitPairwise of catboost's ranking loss fuction is strong, and my lightgbm classifier used variable candidators for each session, there will be more bias for each session,  the ranking model has less bias per session.",
      "votes": null
    },
    {
      "id": "2129490",
      "postDate": "02/04/2023 17:01:45",
      "content": "<p>Congratulations !!!</p>",
      "rawMarkdown": "Congratulations !!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125673,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "02/01/2023 19:42:55",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> we learnt lot from you ideas from the H&amp;M comp . Could you give some details on <code>cosine similarity of session2aid(ProNE)</code> approach .. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2126065,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "02/02/2023 04:32:36",
          "content": "<p>treat session and aid both as a node of one graph, then train ProNE model to get graph embedding, finally calculate cosine similary of session embedding and aid embedding</p>",
          "votes": null,
          "replies": [
            {
              "id": 2127388,
              "author_name": "gauravbrills",
              "author_url": "",
              "post_date": "02/02/2023 21:12:22",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> will keep this in mind for the future, we did generate GNN embeddings but added them directly which didn't work that well .great to learn . </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2125911,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "02/02/2023 01:48:23",
      "content": "<p>Congrats on yet another impressive result!!! </p>\n<blockquote>\n  <p>we found only repeat action,next action</p>\n</blockquote>\n<p>I wonder how to calculate this? maybe you can get repeat action based on counts, but how about next action?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126059,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "02/02/2023 04:29:36",
          "content": "<p>let me show you the polars code<br>\ndf: candidates<br>\ntrain: previoud session + current session<br>\ntest: current session</p>\n<pre><code>   train = train.with_column(\n            (pl.col().shift(-).over()).alias()\n            )    \n    train = train.with_column(\n            (pl.col().shift(-).over()).alias()\n            )     \n\n    df = df.join(test\n                 .select([,,])\n                 .unique(subset=, keep=),\n            on=[], how=)  \n    df.columns = [,,,]     \n\n    df = df.join(train                 \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                 \n            ]\n        ), \n        on=[,,], how=)    \n\n    df = df.join(train    \n        .(\n        pl.col()==     \n        )                       \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                  \n            ]\n        ), \n        on=[,,], how=)    \n\n    df = df.join(train    \n        .(\n        pl.col()==    \n        )                       \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                  \n            ]\n        ), \n        on=[,,], how=)        \n\n    df = df.join(train    \n        .(\n        pl.col()==     \n        )                       \n        .groupby(by=[,,])\n        .agg(\n            [\n                pl.count().alias(),                  \n            ]\n        ), \n        on=[,,], how=)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125914,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/02/2023 01:50:09",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and team. Great job winning top spots in both H&amp;M comp and Otto comp!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2126123,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "02/02/2023 05:30:40",
      "content": "<p>Congratulations!</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127225,
      "author_name": "sohilsharma1996",
      "author_url": "",
      "post_date": "02/02/2023 18:55:54",
      "content": "<p>Congratulations !!! Great efforts</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127344,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/02/2023 20:23:12",
      "content": "<blockquote>\n  <p>At H&amp;M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.</p>\n</blockquote>\n<p>Wow that's a big boost. I just retrained my XGB model now using CatBoostRanker and <code>PairLogitPairwise</code> loss. The orders boost <code>+0.00063</code> and the carts boost <code>+0.00074</code>. Wow. During the competition i tried CatBoost with YetiRank loss but that did poorly. I didn't know that CatBoost had a pairwise loss until after the comp ended and read winners writeups.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2127387,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "02/02/2023 21:11:05",
          "content": "<p>For Catboost we also tried <a href=\"https://catboost.ai/en/docs/references/querycrossentropy\" target=\"_blank\">https://catboost.ai/en/docs/references/querycrossentropy</a> loss and it worked better for us .. more than pariwise logit or yeti . Not sure top teams also tried this ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2127389,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/02/2023 21:12:44",
              "content": "<p>Interesting. I'll give it a try and report results here</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2127419,
                  "author_name": "gauravbrills",
                  "author_url": "",
                  "post_date": "02/02/2023 21:46:41",
                  "content": "<p>sure will be good to know . also it works only on GPU and I think its like a group Log loss .We also had an idea to implement this for some NN rankers but we didnt get the time .</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2127390,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/02/2023 21:13:38",
          "content": "<p>I just submitted my CatBoost (with <code>PairLogitPairwise</code>) to LB and it boost LB <code>+0.0006</code> too (compared with XGB) for my single model. Big boost.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2127633,
              "author_name": "senkin13",
              "author_url": "",
              "post_date": "02/03/2023 02:23:57",
              "content": "<p>I think PairLogitPairwise of catboost's ranking loss fuction is strong, and my lightgbm classifier used variable candidators for each session, there will be more bias for each session,  the ranking model has less bias per session.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2127583,
      "author_name": "lynnswt",
      "author_url": "",
      "post_date": "02/03/2023 01:27:57",
      "content": "<p>Congratulations !! <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2129490,
      "author_name": "valentasarbaciauskas",
      "author_url": "",
      "post_date": "02/04/2023 17:01:45",
      "content": "<p>Congratulations !!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2124741": "Thanks to OTTO and Kaggle Team giving us such a wonderful competition very close to real business.Thanks my team mates @h4211819 @onodera @psilogram ,we did many creative work and work hard util the last hours before end.\nI and 30CrMnSiA won the H&M Personalized Fashion Recommendations last year, we think we can start from midterm,but we found this competition is very hard and very different from H&M competition,so we merge with silonodera team for last 18 days.I am satisfied with the progress and result.\n\n# Overview\n\n[![Screenshot-2023-02-01-at-15-43-31.png](https://i.postimg.cc/C1Qs0twB/Screenshot-2023-02-01-at-15-43-31.png)](https://postimg.cc/cgfnMFg0)\n\n# Retrieval(aka: candidate generation | recall)\nThis competition data has no user/item information, and popularity don't work, hard to improve retrieval, we found only repeat action,next action and itemcf work,there are nothing to do for repeat action,next action,the most important work is to improve itemcf, we tried kinds of matrix and weight for itemcf and worked.\n[matrix]\n- cart-order \n- click-cart&order \n- click-cart-order \n\n[weight]\n- type \n- co-occurrence time distance \n- co-occurrence order\n- session action order \n- popularity \n\n# Ranking\n- **Feature Engineering**\nBasiclly,features are created base on retrieval strategies,create user-item interaction for repeat action,next action,aggregated collaborative filtering score from itemcf,similarity for embedding retrieval.\nI used 198 features for order/cart model, 106 features for click model\n\n| Type | Description|\n| --- | --- |\n| aggregation |  count of session,aid,session-aid,each type aggregated by current session|last week sessions|previous sessions |\n| next action |  count of aid which is the next of session's last aid|\n| time|  first,last time of session,aid|\n| collaborative filtering score| max,sum,weighted sum of session's last aid,last hour aids,all aids & candidate aid score|\n| embedding similarity| cosine similarity of aid2aid(word2vec), cosine similarity of session2aid(ProNE) |\n\n- **Validation**\nI made 3 weeks' train data with labels, use last week as validation,then retrain all train data with fixed round number for test data prediction,the CV matched LB very well, we don't need to watch LB by submission.\n\n- **Model**\nAt H&M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.\nMy best single model is a catboost ranker, CV 0.59066 = 0.6710 * 0.6 + 0.4416 * 0.3 + 0.5558 * 0.1, LB 0.602 = 0.409 + 0.137 + 0.056.\nFinally I blend with lightgbm and catboost,the blend LB is higher 0.602, after blend with teammate's 0.602,LB is 0.604\n\n# Optimization\nAt last days I rewrite all my feature engineering code from pandas to polars,the polars is much faster,especially two huge dataframe join, pandas.merge -> polars.join = 40X faster.\nI use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.",
    "2125673": "Thanks @senkin13 we learnt lot from you ideas from the H&M comp . Could you give some details on `cosine similarity of session2aid(ProNE)` approach ..",
    "2125911": "Congrats on yet another impressive result!!! \n> we found only repeat action,next action\n\n\nI wonder how to calculate this? maybe you can get repeat action based on counts, but how about next action?",
    "2125914": "Congratulations @senkin13 and team. Great job winning top spots in both H&M comp and Otto comp!",
    "2126059": "let me show you the polars code\ndf: candidates\ntrain: previoud session + current session\ntest: current session\n```python\n   train = train.with_column(\n            (pl.col('aid').shift(-1).over('session')).alias('aid_next')\n            )    \n    train = train.with_column(\n            (pl.col('type').shift(-1).over('session')).alias('type_next')\n            )     \n        \n    df = df.join(test\n                 .select(['session','aid','type'])\n                 .unique(subset='session', keep='last'),\n            on=['session'], how='left')  \n    df.columns = ['session','aid','aid_next','type']     \n    \n    df = df.join(train                 \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_count'),                 \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')    \n    \n    df = df.join(train    \n        .filter(\n        pl.col('type_next')==0     \n        )                       \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_click_count'),                  \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')    \n    \n    df = df.join(train    \n        .filter(\n        pl.col('type_next')==1    \n        )                       \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_cart_count'),                  \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')        \n    \n    df = df.join(train    \n        .filter(\n        pl.col('type_next')==2     \n        )                       \n        .groupby(by=['aid','type','aid_next'])\n        .agg(\n            [\n                pl.count().alias('next_order_count'),                  \n            ]\n        ), \n        on=['aid','type','aid_next'], how='left')\n```",
    "2126065": "treat session and aid both as a node of one graph, then train ProNE model to get graph embedding, finally calculate cosine similary of session embedding and aid embedding",
    "2126123": "Congratulations!\n\nThe Devastator.",
    "2127225": "Congratulations !!! Great efforts",
    "2127344": ">At H&M competition, my best model is a lightgbm binary classifier,I also start with lightgbm at this competition,at last days I tried catboost ranker, the improvement surprised me, order model: 0.0007 up, cart model: 0.002 up, clicks model: 0.0012 up.\n\nWow that's a big boost. I just retrained my XGB model now using CatBoostRanker and `PairLogitPairwise` loss. The orders boost `+0.00063` and the carts boost `+0.00074`. Wow. During the competition i tried CatBoost with YetiRank loss but that did poorly. I didn't know that CatBoost had a pairwise loss until after the comp ended and read winners writeups.",
    "2127387": "For Catboost we also tried https://catboost.ai/en/docs/references/querycrossentropy loss and it worked better for us .. more than pariwise logit or yeti . Not sure top teams also tried this ?",
    "2127388": "Thanks @senkin13 will keep this in mind for the future, we did generate GNN embeddings but added them directly which didn't work that well .great to learn .",
    "2127389": "Interesting. I'll give it a try and report results here",
    "2127390": "I just submitted my CatBoost (with `PairLogitPairwise`) to LB and it boost LB `+0.0006` too (compared with XGB) for my single model. Big boost.",
    "2127419": "sure will be good to know . also it works only on GPU and I think its like a group Log loss .We also had an idea to implement this for some NN rankers but we didnt get the time .",
    "2127583": "Congratulations !! @senkin13",
    "2127633": "I think PairLogitPairwise of catboost's ranking loss fuction is strong, and my lightgbm classifier used variable candidators for each session, there will be more bias for each session,  the ranking model has less bias per session.",
    "2129490": "Congratulations !!!"
  },
  "source": "meta"
}