{
  "id": 382905,
  "title": "15th Place Solution",
  "url": "/competitions/otto-recommender-system/writeups/hjam-15th-place-solution",
  "author_name": "",
  "post_date": "2023-02-10T01:19:19.613Z",
  "votes": 19,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, I would thank the organizers and those who shared their knowledge and ideas. Especially, I would like to thank <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> for providing great numba notebook. Without him, I must give up this competition and enjoy  my New year holiday lying in the bed everyday (LOL.)</p>\n<h2>Approch Summary</h2>\n<ul>\n<li>Retrieval &amp; Re-rank 2stage model</li>\n<li>local cv: train on week3/ valid on week 4</li>\n<li>Retrieval recall20: 0.585(pulic LB)/0.585(private LB)</li>\n<li>Re-rank models: lightgbm ranker with 182 features. (single model get 0.601(public LB)/0.600(private LB))</li>\n<li>ensemble: 5 lightgbm ranker with same features with different seed  for negative sampling and model training. ( 0.601(public LB)/0.601(private LB))</li>\n</ul>\n<h2>Retrieval/candidate stage</h2>\n<ul>\n<li>100 candidates for each session. (candidates number 50 -&gt; 100, which boosts 0.0007. have no time and memory to test 200.)<ul>\n<li>historical aids.</li>\n<li>i2i2i based on <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> 's notebook </li>\n<li>most popular aids to make up for 100 candidates.</li></ul></li>\n<li>local cv on week4 [LB 0.585]</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>num of candidates</th>\n<th>click</th>\n<th>cart</th>\n<th>order</th>\n<th>sum</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>20</td>\n<td>0.5397</td>\n<td>0.4220</td>\n<td>0.6580</td>\n<td>0.5754</td>\n</tr>\n<tr>\n<td>50</td>\n<td>0.6223</td>\n<td>0.4838</td>\n<td>0.69498</td>\n<td>0.6243</td>\n</tr>\n<tr>\n<td>100</td>\n<td>0.6745</td>\n<td>0.5265</td>\n<td>0.7188</td>\n<td>0.6567</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>main modification of i2i<ul>\n<li>order weight: adjacent aids in interactions have larger weight. (session 1 has a,b,c aids, the weight of b,c is 1 and the weight of a,c is 1/2.) Use those weight in both i2i similarity and u2i recommandation. I didn't use time weight here, because I notice most of interactions are within 1hour.</li>\n<li>Normalize the i2i similarity dict. The popular aids have larger i2i score with other aids. It's not fair in the u2i stage. Just use the mean of i2i scores of popular aids to normalize.</li></ul></li>\n</ul>\n<h2>Rerank-models</h2>\n<p>lightgbm ranker with 182 features. 0.2 negative ratio. (single model get 0.601(public LB)/0.600(private LB))</p>\n<ul>\n<li>session features(adding more session features didn't work for me)<ul>\n<li>event count/ type count</li>\n<li>aid_nunique/ type aid_nunique</li>\n<li>first time, last time, time range</li></ul></li>\n<li>aids features <ul>\n<li>kinds of count/nunique from (1week,2week)</li>\n<li>hour features to detect some \"promotion\" aids which only exists for 1-2 hour in a week.</li>\n<li>ratio of buy2cart, buy2order, cart2order</li>\n<li>ratio of re-cart, re-order, re-click</li>\n<li>first last time of aids </li></ul></li>\n<li>aid X session features<ul>\n<li>kinds of count, order decayed counts</li>\n<li>time diff</li></ul></li>\n<li>retrieval features<ul>\n<li>revisit rank/scores</li>\n<li>i2i rank/scores</li>\n<li>(avg/median/max/last one / last two/ last 3 avg) i2i scores</li></ul></li>\n<li>embedding similarity<ul>\n<li>(avg/median/max/last one / last two/ last 3 avg) w2v similarity</li>\n<li>(avg/median/max/last one / last two/ last 3 avg) bpr similarity</li></ul></li>\n</ul>\n<h2>ensemble</h2>\n<ul>\n<li><p>5 lightgbm ranker with same features with different seed  for negative sampling and model training. </p>\n<table>\n<thead>\n<tr>\n<th>LB</th>\n<th>click</th>\n<th>cart</th>\n<th>order</th>\n<th>sum</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>public</td>\n<td>0.056</td>\n<td>0.136</td>\n<td>0.408</td>\n<td>0.601</td>\n</tr>\n<tr>\n<td>private</td>\n<td>0.056</td>\n<td>0.136</td>\n<td>0.408</td>\n<td>0.601</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>catboost&amp;&amp;MLP ensemble can boost 0.0003 in local cv but have no time to run it.</p></li>\n</ul>\n<h2>what didn't work for me.</h2>\n<ul>\n<li>prone similarity</li>\n<li>click/cart/order 3 targets are correlated. So I use pred score of other 2 targets as features.(i.e. when predicting order, I use pred score of click and cart). In local cv, it boosted almost 0.001 but didn't work in LB. don't know why.</li>\n<li>combining week3 data and week4 data as final trainset didn't work. I don't create week2 data to test it in local cv. I just used 1.25 * num of rounds of lgb using only week4 to train.</li>\n<li>adding more session feas and interaction feas.</li>\n<li>adding k nearest neighbor of w2v embedding in the retrieval stage didn't improve hit rate.</li>\n</ul>",
  "messages": [
    {
      "id": "2125160",
      "postDate": "02/01/2023 13:27:30",
      "content": "<p>First of all, I would thank the organizers and those who shared their knowledge and ideas. Especially, I would like to thank <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> for providing great numba notebook. Without him, I must give up this competition and enjoy  my New year holiday lying in the bed everyday (LOL.)</p>\n<h2>Approch Summary</h2>\n<ul>\n<li>Retrieval &amp; Re-rank 2stage model</li>\n<li>local cv: train on week3/ valid on week 4</li>\n<li>Retrieval recall20: 0.585(pulic LB)/0.585(private LB)</li>\n<li>Re-rank models: lightgbm ranker with 182 features. (single model get 0.601(public LB)/0.600(private LB))</li>\n<li>ensemble: 5 lightgbm ranker with same features with different seed  for negative sampling and model training. ( 0.601(public LB)/0.601(private LB))</li>\n</ul>\n<h2>Retrieval/candidate stage</h2>\n<ul>\n<li>100 candidates for each session. (candidates number 50 -&gt; 100, which boosts 0.0007. have no time and memory to test 200.)<ul>\n<li>historical aids.</li>\n<li>i2i2i based on <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> 's notebook </li>\n<li>most popular aids to make up for 100 candidates.</li></ul></li>\n<li>local cv on week4 [LB 0.585]</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>num of candidates</th>\n<th>click</th>\n<th>cart</th>\n<th>order</th>\n<th>sum</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>20</td>\n<td>0.5397</td>\n<td>0.4220</td>\n<td>0.6580</td>\n<td>0.5754</td>\n</tr>\n<tr>\n<td>50</td>\n<td>0.6223</td>\n<td>0.4838</td>\n<td>0.69498</td>\n<td>0.6243</td>\n</tr>\n<tr>\n<td>100</td>\n<td>0.6745</td>\n<td>0.5265</td>\n<td>0.7188</td>\n<td>0.6567</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>main modification of i2i<ul>\n<li>order weight: adjacent aids in interactions have larger weight. (session 1 has a,b,c aids, the weight of b,c is 1 and the weight of a,c is 1/2.) Use those weight in both i2i similarity and u2i recommandation. I didn't use time weight here, because I notice most of interactions are within 1hour.</li>\n<li>Normalize the i2i similarity dict. The popular aids have larger i2i score with other aids. It's not fair in the u2i stage. Just use the mean of i2i scores of popular aids to normalize.</li></ul></li>\n</ul>\n<h2>Rerank-models</h2>\n<p>lightgbm ranker with 182 features. 0.2 negative ratio. (single model get 0.601(public LB)/0.600(private LB))</p>\n<ul>\n<li>session features(adding more session features didn't work for me)<ul>\n<li>event count/ type count</li>\n<li>aid_nunique/ type aid_nunique</li>\n<li>first time, last time, time range</li></ul></li>\n<li>aids features <ul>\n<li>kinds of count/nunique from (1week,2week)</li>\n<li>hour features to detect some \"promotion\" aids which only exists for 1-2 hour in a week.</li>\n<li>ratio of buy2cart, buy2order, cart2order</li>\n<li>ratio of re-cart, re-order, re-click</li>\n<li>first last time of aids </li></ul></li>\n<li>aid X session features<ul>\n<li>kinds of count, order decayed counts</li>\n<li>time diff</li></ul></li>\n<li>retrieval features<ul>\n<li>revisit rank/scores</li>\n<li>i2i rank/scores</li>\n<li>(avg/median/max/last one / last two/ last 3 avg) i2i scores</li></ul></li>\n<li>embedding similarity<ul>\n<li>(avg/median/max/last one / last two/ last 3 avg) w2v similarity</li>\n<li>(avg/median/max/last one / last two/ last 3 avg) bpr similarity</li></ul></li>\n</ul>\n<h2>ensemble</h2>\n<ul>\n<li><p>5 lightgbm ranker with same features with different seed  for negative sampling and model training. </p>\n<table>\n<thead>\n<tr>\n<th>LB</th>\n<th>click</th>\n<th>cart</th>\n<th>order</th>\n<th>sum</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>public</td>\n<td>0.056</td>\n<td>0.136</td>\n<td>0.408</td>\n<td>0.601</td>\n</tr>\n<tr>\n<td>private</td>\n<td>0.056</td>\n<td>0.136</td>\n<td>0.408</td>\n<td>0.601</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>catboost&amp;&amp;MLP ensemble can boost 0.0003 in local cv but have no time to run it.</p></li>\n</ul>\n<h2>what didn't work for me.</h2>\n<ul>\n<li>prone similarity</li>\n<li>click/cart/order 3 targets are correlated. So I use pred score of other 2 targets as features.(i.e. when predicting order, I use pred score of click and cart). In local cv, it boosted almost 0.001 but didn't work in LB. don't know why.</li>\n<li>combining week3 data and week4 data as final trainset didn't work. I don't create week2 data to test it in local cv. I just used 1.25 * num of rounds of lgb using only week4 to train.</li>\n<li>adding more session feas and interaction feas.</li>\n<li>adding k nearest neighbor of w2v embedding in the retrieval stage didn't improve hit rate.</li>\n</ul>",
      "rawMarkdown": "First of all, I would thank the organizers and those who shared their knowledge and ideas. Especially, I would like to thank @carnozhao for providing great numba notebook. Without him, I must give up this competition and enjoy  my New year holiday lying in the bed everyday (LOL.)\n\n## Approch Summary\n- Retrieval & Re-rank 2stage model\n- local cv: train on week3/ valid on week 4\n- Retrieval recall20: 0.585(pulic LB)/0.585(private LB)\n- Re-rank models: lightgbm ranker with 182 features. (single model get 0.601(public LB)/0.600(private LB))\n- ensemble: 5 lightgbm ranker with same features with different seed  for negative sampling and model training. ( 0.601(public LB)/0.601(private LB))\n## Retrieval/candidate stage\n- 100 candidates for each session. (candidates number 50 -> 100, which boosts 0.0007. have no time and memory to test 200.)\n    - historical aids.\n    - i2i2i based on @carnozhao 's notebook \n    - most popular aids to make up for 100 candidates.\n- local cv on week4 [LB 0.585]\n\n| num of candidates | click |cart|order|sum|\n| --- | --- |--- |--- |--- |\n| 20 | 0.5397 |0.4220|0.6580|0.5754|\n|50|0.6223|0.4838|0.69498|0.6243|\n|100|0.6745|0.5265|0.7188|0.6567|\n\n- main modification of i2i\n    - order weight: adjacent aids in interactions have larger weight. (session 1 has a,b,c aids, the weight of b,c is 1 and the weight of a,c is 1/2.) Use those weight in both i2i similarity and u2i recommandation. I didn't use time weight here, because I notice most of interactions are within 1hour.\n    -  Normalize the i2i similarity dict. The popular aids have larger i2i score with other aids. It's not fair in the u2i stage. Just use the mean of i2i scores of popular aids to normalize.\n\n## Rerank-models\nlightgbm ranker with 182 features. 0.2 negative ratio. (single model get 0.601(public LB)/0.600(private LB))\n\n- session features(adding more session features didn't work for me)\n    - event count/ type count\n    - aid_nunique/ type aid_nunique\n    - first time, last time, time range\n- aids features \n    - kinds of count/nunique from (1week,2week)\n    - hour features to detect some \"promotion\" aids which only exists for 1-2 hour in a week.\n    - ratio of buy2cart, buy2order, cart2order\n    - ratio of re-cart, re-order, re-click\n    - first last time of aids \n- aid X session features\n    - kinds of count, order decayed counts\n    - time diff\n- retrieval features\n    - revisit rank/scores\n    - i2i rank/scores\n    - (avg/median/max/last one / last two/ last 3 avg) i2i scores\n- embedding similarity\n    - (avg/median/max/last one / last two/ last 3 avg) w2v similarity\n    - (avg/median/max/last one / last two/ last 3 avg) bpr similarity\n## ensemble\n- 5 lightgbm ranker with same features with different seed  for negative sampling and model training. \n| LB|  click|cart|order|sum|\n| --- | --- |--- |--- |--- |\n| public |  0.056|0.136|0.408|0.601|\n| private|  0.056|0.136|0.408|0.601|\n\n- catboost&&MLP ensemble can boost 0.0003 in local cv but have no time to run it.\n## what didn't work for me.\n- prone similarity\n- click/cart/order 3 targets are correlated. So I use pred score of other 2 targets as features.(i.e. when predicting order, I use pred score of click and cart). In local cv, it boosted almost 0.001 but didn't work in LB. don't know why.\n- combining week3 data and week4 data as final trainset didn't work. I don't create week2 data to test it in local cv. I just used 1.25 * num of rounds of lgb using only week4 to train.\n- adding more session feas and interaction feas.\n- adding k nearest neighbor of w2v embedding in the retrieval stage didn't improve hit rate.",
      "votes": null
    },
    {
      "id": "2125316",
      "postDate": "02/01/2023 15:31:00",
      "content": "<p>Congras and thanks for sharing. May I ask did you train your model on CV data and make inference on test data directly ?</p>",
      "rawMarkdown": "Congras and thanks for sharing. May I ask did you train your model on CV data and make inference on test data directly ?",
      "votes": null
    },
    {
      "id": "2125329",
      "postDate": "02/01/2023 15:36:39",
      "content": "<p>Sure. For LB submit, I would re-train the model on week4 data and predict the test data. (for cv, I train on week3 and predict week4.)</p>",
      "rawMarkdown": "Sure. For LB submit, I would re-train the model on week4 data and predict the test data. (for cv, I train on week3 and predict week4.)",
      "votes": null
    },
    {
      "id": "2126146",
      "postDate": "02/02/2023 05:43:08",
      "content": "<p>Congratulations on a strong finish!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Congratulations on a strong finish!\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2126239",
      "postDate": "02/02/2023 06:21:52",
      "content": "<p>Thanks, Devastator!</p>",
      "rawMarkdown": "Thanks, Devastator!",
      "votes": null
    },
    {
      "id": "2126424",
      "postDate": "02/02/2023 08:17:24",
      "content": "<p>Let's fight for gold next game, my bro! 👍🎉</p>",
      "rawMarkdown": "Let's fight for gold next game, my bro! 👍🎉",
      "votes": null
    },
    {
      "id": "2126437",
      "postDate": "02/02/2023 08:24:51",
      "content": "<p>hahaha. good luck to us.😜</p>",
      "rawMarkdown": "hahaha. good luck to us.😜",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125316,
      "author_name": "weiqiu",
      "author_url": "",
      "post_date": "02/01/2023 15:31:00",
      "content": "<p>Congras and thanks for sharing. May I ask did you train your model on CV data and make inference on test data directly ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125329,
          "author_name": "hookman",
          "author_url": "",
          "post_date": "02/01/2023 15:36:39",
          "content": "<p>Sure. For LB submit, I would re-train the model on week4 data and predict the test data. (for cv, I train on week3 and predict week4.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2126146,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "02/02/2023 05:43:08",
      "content": "<p>Congratulations on a strong finish!</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126239,
          "author_name": "hookman",
          "author_url": "",
          "post_date": "02/02/2023 06:21:52",
          "content": "<p>Thanks, Devastator!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2126424,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "02/02/2023 08:17:24",
      "content": "<p>Let's fight for gold next game, my bro! 👍🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126437,
          "author_name": "hookman",
          "author_url": "",
          "post_date": "02/02/2023 08:24:51",
          "content": "<p>hahaha. good luck to us.😜</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2125160": "First of all, I would thank the organizers and those who shared their knowledge and ideas. Especially, I would like to thank @carnozhao for providing great numba notebook. Without him, I must give up this competition and enjoy  my New year holiday lying in the bed everyday (LOL.)\n\n## Approch Summary\n- Retrieval & Re-rank 2stage model\n- local cv: train on week3/ valid on week 4\n- Retrieval recall20: 0.585(pulic LB)/0.585(private LB)\n- Re-rank models: lightgbm ranker with 182 features. (single model get 0.601(public LB)/0.600(private LB))\n- ensemble: 5 lightgbm ranker with same features with different seed  for negative sampling and model training. ( 0.601(public LB)/0.601(private LB))\n## Retrieval/candidate stage\n- 100 candidates for each session. (candidates number 50 -> 100, which boosts 0.0007. have no time and memory to test 200.)\n    - historical aids.\n    - i2i2i based on @carnozhao 's notebook \n    - most popular aids to make up for 100 candidates.\n- local cv on week4 [LB 0.585]\n\n| num of candidates | click |cart|order|sum|\n| --- | --- |--- |--- |--- |\n| 20 | 0.5397 |0.4220|0.6580|0.5754|\n|50|0.6223|0.4838|0.69498|0.6243|\n|100|0.6745|0.5265|0.7188|0.6567|\n\n- main modification of i2i\n    - order weight: adjacent aids in interactions have larger weight. (session 1 has a,b,c aids, the weight of b,c is 1 and the weight of a,c is 1/2.) Use those weight in both i2i similarity and u2i recommandation. I didn't use time weight here, because I notice most of interactions are within 1hour.\n    -  Normalize the i2i similarity dict. The popular aids have larger i2i score with other aids. It's not fair in the u2i stage. Just use the mean of i2i scores of popular aids to normalize.\n\n## Rerank-models\nlightgbm ranker with 182 features. 0.2 negative ratio. (single model get 0.601(public LB)/0.600(private LB))\n\n- session features(adding more session features didn't work for me)\n    - event count/ type count\n    - aid_nunique/ type aid_nunique\n    - first time, last time, time range\n- aids features \n    - kinds of count/nunique from (1week,2week)\n    - hour features to detect some \"promotion\" aids which only exists for 1-2 hour in a week.\n    - ratio of buy2cart, buy2order, cart2order\n    - ratio of re-cart, re-order, re-click\n    - first last time of aids \n- aid X session features\n    - kinds of count, order decayed counts\n    - time diff\n- retrieval features\n    - revisit rank/scores\n    - i2i rank/scores\n    - (avg/median/max/last one / last two/ last 3 avg) i2i scores\n- embedding similarity\n    - (avg/median/max/last one / last two/ last 3 avg) w2v similarity\n    - (avg/median/max/last one / last two/ last 3 avg) bpr similarity\n## ensemble\n- 5 lightgbm ranker with same features with different seed  for negative sampling and model training. \n| LB|  click|cart|order|sum|\n| --- | --- |--- |--- |--- |\n| public |  0.056|0.136|0.408|0.601|\n| private|  0.056|0.136|0.408|0.601|\n\n- catboost&&MLP ensemble can boost 0.0003 in local cv but have no time to run it.\n## what didn't work for me.\n- prone similarity\n- click/cart/order 3 targets are correlated. So I use pred score of other 2 targets as features.(i.e. when predicting order, I use pred score of click and cart). In local cv, it boosted almost 0.001 but didn't work in LB. don't know why.\n- combining week3 data and week4 data as final trainset didn't work. I don't create week2 data to test it in local cv. I just used 1.25 * num of rounds of lgb using only week4 to train.\n- adding more session feas and interaction feas.\n- adding k nearest neighbor of w2v embedding in the retrieval stage didn't improve hit rate.",
    "2125316": "Congras and thanks for sharing. May I ask did you train your model on CV data and make inference on test data directly ?",
    "2125329": "Sure. For LB submit, I would re-train the model on week4 data and predict the test data. (for cv, I train on week3 and predict week4.)",
    "2126146": "Congratulations on a strong finish!\n\nThe Devastator.",
    "2126239": "Thanks, Devastator!",
    "2126424": "Let's fight for gold next game, my bro! 👍🎉",
    "2126437": "hahaha. good luck to us.😜"
  },
  "source": "meta"
}