{
  "id": 379980,
  "title": "My GBDT Ranker can't beat heuristic approach !?",
  "url": "/competitions/otto-recommender-system/discussion/379980",
  "author_name": "",
  "post_date": "2023-01-21T22:04:12.155807300Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have been working so hard in this challenge for the past two weeks ,but unfortunately my ranker model can’t beat the heuristic approach shared in public notebooks.<br>\nMy pipeline is for each target :<br>\n     1. Generate up to 50/200 aids with co-visitation matrices .<br>\n     2. I used suggests function  to generate candidates . =&gt;Top 20<br>\n     3. I train on half of the session of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> data ,for train (<code>%2==0</code>) and validate (<code>%2==1</code>)<br>\n     4. I create user features and  item features(using train + test) (around 90 features)<br>\n     5. before merging with the true labels , under sample negatives by 15%<br>\n     6. after merging my candidate data with ground truth i got :</p>\n<blockquote>\n  <ul>\n  <li>2633153 negatives &amp; 68725 positives for clicks</li>\n  <li>2684244 negatives  &amp; 17634 positives for carts</li>\n  <li>2686587 negatives &amp; 15291 positives for orders</li>\n  </ul>\n  <ol>\n  <li>I train LGBMRanker for <code>20 rounds</code><br>\n  My CV is :<br>\n  clicks recall = 0.5231477262371275<br>\n  carts recall = 0.40543254663219636<br>\n  orders recall = 0.6483631646548363<br>\n  <strong>Overall Recall = 0.5629624354062734</strong></li>\n  </ol>\n</blockquote>\n<p>I tried to generate more candidates with more co-vis matrices  but CV always decrease , top 20 give best CV .<br>\nAlso I tried to pass GC_ranking as feature but CV decrease.<br>\nAnything is wrong with my pipeline ?</p>",
  "messages": [
    {
      "id": "2110039",
      "postDate": "01/21/2023 22:04:12",
      "content": "<p>I have been working so hard in this challenge for the past two weeks ,but unfortunately my ranker model can’t beat the heuristic approach shared in public notebooks.<br>\nMy pipeline is for each target :<br>\n     1. Generate up to 50/200 aids with co-visitation matrices .<br>\n     2. I used suggests function  to generate candidates . =&gt;Top 20<br>\n     3. I train on half of the session of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> data ,for train (<code>%2==0</code>) and validate (<code>%2==1</code>)<br>\n     4. I create user features and  item features(using train + test) (around 90 features)<br>\n     5. before merging with the true labels , under sample negatives by 15%<br>\n     6. after merging my candidate data with ground truth i got :</p>\n<blockquote>\n  <ul>\n  <li>2633153 negatives &amp; 68725 positives for clicks</li>\n  <li>2684244 negatives  &amp; 17634 positives for carts</li>\n  <li>2686587 negatives &amp; 15291 positives for orders</li>\n  </ul>\n  <ol>\n  <li>I train LGBMRanker for <code>20 rounds</code><br>\n  My CV is :<br>\n  clicks recall = 0.5231477262371275<br>\n  carts recall = 0.40543254663219636<br>\n  orders recall = 0.6483631646548363<br>\n  <strong>Overall Recall = 0.5629624354062734</strong></li>\n  </ol>\n</blockquote>\n<p>I tried to generate more candidates with more co-vis matrices  but CV always decrease , top 20 give best CV .<br>\nAlso I tried to pass GC_ranking as feature but CV decrease.<br>\nAnything is wrong with my pipeline ?</p>",
      "rawMarkdown": "I have been working so hard in this challenge for the past two weeks ,but unfortunately my ranker model can’t beat the heuristic approach shared in public notebooks.\nMy pipeline is for each target :\n     1. Generate up to 50/200 aids with co-visitation matrices .\n     2. I used suggests function  to generate candidates . =>Top 20\n     3. I train on half of the session of @cdeotte data ,for train (`%2==0`) and validate (`%2==1`)\n     4. I create user features and  item features(using train + test) (around 90 features)\n     5. before merging with the true labels , under sample negatives by 15%\n     6. after merging my candidate data with ground truth i got :\n >  * 2633153 negatives & 68725 positives for clicks\n >  *  2684244 negatives  & 17634 positives for carts\n >  * 2686587 negatives & 15291 positives for orders\n4. I train LGBMRanker for `20 rounds`\nMy CV is :\nclicks recall = 0.5231477262371275\ncarts recall = 0.40543254663219636\norders recall = 0.6483631646548363\n**Overall Recall = 0.5629624354062734**\n\nI tried to generate more candidates with more co-vis matrices  but CV always decrease , top 20 give best CV .\nAlso I tried to pass GC_ranking as feature but CV decrease.\nAnything is wrong with my pipeline ?",
      "votes": null
    },
    {
      "id": "2110745",
      "postDate": "01/22/2023 11:52:19",
      "content": "<p>HI, I also experimented this. From my point of view, here's what you can try ( that works for me ):</p>\n<ul>\n<li>Create similarity scores variables between users|candidates, and candidates | last item seen by the user.</li>\n<li>Increase the learning rate ( it gave me the little boost I needed to reach 0.6521).</li>\n</ul>\n<p>Also, your recall can be improved by mixing the predictions of your ranker and the heuristic approach, simply because co-visitation matrices are much powerful in some cases.</p>",
      "rawMarkdown": "HI, I also experimented this. From my point of view, here's what you can try ( that works for me ):\n\n- Create similarity scores variables between users|candidates, and candidates | last item seen by the user.\n- Increase the learning rate ( it gave me the little boost I needed to reach 0.6521).\n\n\n\nAlso, your recall can be improved by mixing the predictions of your ranker and the heuristic approach, simply because co-visitation matrices are much powerful in some cases.",
      "votes": null
    },
    {
      "id": "2110789",
      "postDate": "01/22/2023 12:33:30",
      "content": "<p>thank you 😊 </p>",
      "rawMarkdown": "thank you 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2110745,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/22/2023 11:52:19",
      "content": "<p>HI, I also experimented this. From my point of view, here's what you can try ( that works for me ):</p>\n<ul>\n<li>Create similarity scores variables between users|candidates, and candidates | last item seen by the user.</li>\n<li>Increase the learning rate ( it gave me the little boost I needed to reach 0.6521).</li>\n</ul>\n<p>Also, your recall can be improved by mixing the predictions of your ranker and the heuristic approach, simply because co-visitation matrices are much powerful in some cases.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2110789,
          "author_name": "ihebch",
          "author_url": "",
          "post_date": "01/22/2023 12:33:30",
          "content": "<p>thank you 😊 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2110039": "I have been working so hard in this challenge for the past two weeks ,but unfortunately my ranker model can’t beat the heuristic approach shared in public notebooks.\nMy pipeline is for each target :\n     1. Generate up to 50/200 aids with co-visitation matrices .\n     2. I used suggests function  to generate candidates . =>Top 20\n     3. I train on half of the session of @cdeotte data ,for train (`%2==0`) and validate (`%2==1`)\n     4. I create user features and  item features(using train + test) (around 90 features)\n     5. before merging with the true labels , under sample negatives by 15%\n     6. after merging my candidate data with ground truth i got :\n >  * 2633153 negatives & 68725 positives for clicks\n >  *  2684244 negatives  & 17634 positives for carts\n >  * 2686587 negatives & 15291 positives for orders\n4. I train LGBMRanker for `20 rounds`\nMy CV is :\nclicks recall = 0.5231477262371275\ncarts recall = 0.40543254663219636\norders recall = 0.6483631646548363\n**Overall Recall = 0.5629624354062734**\n\nI tried to generate more candidates with more co-vis matrices  but CV always decrease , top 20 give best CV .\nAlso I tried to pass GC_ranking as feature but CV decrease.\nAnything is wrong with my pipeline ?",
    "2110745": "HI, I also experimented this. From my point of view, here's what you can try ( that works for me ):\n\n- Create similarity scores variables between users|candidates, and candidates | last item seen by the user.\n- Increase the learning rate ( it gave me the little boost I needed to reach 0.6521).\n\n\n\nAlso, your recall can be improved by mixing the predictions of your ranker and the heuristic approach, simply because co-visitation matrices are much powerful in some cases.",
    "2110789": "thank you 😊"
  },
  "source": "meta"
}