{
  "id": 381717,
  "title": "Questions for the TOP 100 : Are you using only ranker  ?",
  "url": "/competitions/otto-recommender-system/discussion/381717",
  "author_name": "",
  "post_date": "2023-01-27T23:48:49.855393700Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello community, I know I posted a lot of topics, but  I found this competition very interesting and at the same time rather intriguing, given that it's the first time on kaggle that I've been so involved in a competition.</p>\n<p>I'm writing this topic because even when using a GBDT Ranker with the appropriate features ( I suppose haha ), <strong>I couldn't improve my recall over 0.565 using the Ranker</strong>, which means that the heuristic approach is better than my model.</p>\n<p>I used:</p>\n<ul>\n<li>Items and Users Features.</li>\n<li>Interactions Features.</li>\n<li>Similarities between entities.</li>\n<li>GroupKFOLD.</li>\n<li>Feature selection</li>\n</ul>\n<p>I took care of :</p>\n<ul>\n<li>Not include the future in the training --&gt; NO LEAK !!</li>\n</ul>\n<p>I tried:</p>\n<ul>\n<li>DART / GOSS Booster.</li>\n<li>Regularization using L1 L2 / bagging / Feature sampling.</li>\n</ul>\n<p>Another point, is that my last model scored <strong>0.667 on orders</strong> in my local validation, but It still scored  0.565 on LB, using the Ranker, no matter if I  use 100  / 300 or  / 1000 estimators.</p>\n<p>However, I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices, which means that the Ranker scores better than the heuristic model in some cases, which also means that the model is doing the job somehow !</p>\n<p>A big thanks to the kagglers of this competitions for all the tips and sharing !</p>",
  "messages": [
    {
      "id": "2118324",
      "postDate": "01/27/2023 23:48:49",
      "content": "<p>Hello community, I know I posted a lot of topics, but  I found this competition very interesting and at the same time rather intriguing, given that it's the first time on kaggle that I've been so involved in a competition.</p>\n<p>I'm writing this topic because even when using a GBDT Ranker with the appropriate features ( I suppose haha ), <strong>I couldn't improve my recall over 0.565 using the Ranker</strong>, which means that the heuristic approach is better than my model.</p>\n<p>I used:</p>\n<ul>\n<li>Items and Users Features.</li>\n<li>Interactions Features.</li>\n<li>Similarities between entities.</li>\n<li>GroupKFOLD.</li>\n<li>Feature selection</li>\n</ul>\n<p>I took care of :</p>\n<ul>\n<li>Not include the future in the training --&gt; NO LEAK !!</li>\n</ul>\n<p>I tried:</p>\n<ul>\n<li>DART / GOSS Booster.</li>\n<li>Regularization using L1 L2 / bagging / Feature sampling.</li>\n</ul>\n<p>Another point, is that my last model scored <strong>0.667 on orders</strong> in my local validation, but It still scored  0.565 on LB, using the Ranker, no matter if I  use 100  / 300 or  / 1000 estimators.</p>\n<p>However, I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices, which means that the Ranker scores better than the heuristic model in some cases, which also means that the model is doing the job somehow !</p>\n<p>A big thanks to the kagglers of this competitions for all the tips and sharing !</p>",
      "rawMarkdown": "Hello community, I know I posted a lot of topics, but  I found this competition very interesting and at the same time rather intriguing, given that it's the first time on kaggle that I've been so involved in a competition.\n\nI'm writing this topic because even when using a GBDT Ranker with the appropriate features ( I suppose haha ), **I couldn't improve my recall over 0.565 using the Ranker**, which means that the heuristic approach is better than my model.\n\nI used:\n- Items and Users Features.\n- Interactions Features.\n- Similarities between entities.\n- GroupKFOLD.\n- Feature selection\n\nI took care of :\n- Not include the future in the training --> NO LEAK !!\n\nI tried:\n- DART / GOSS Booster.\n- Regularization using L1 L2 / bagging / Feature sampling.\n\n\n\nAnother point, is that my last model scored **0.667 on orders** in my local validation, but It still scored  0.565 on LB, using the Ranker, no matter if I  use 100  / 300 or  / 1000 estimators.\n\nHowever, I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices, which means that the Ranker scores better than the heuristic model in some cases, which also means that the model is doing the job somehow !\n\nA big thanks to the kagglers of this competitions for all the tips and sharing !",
      "votes": null
    },
    {
      "id": "2118330",
      "postDate": "01/28/2023 00:09:52",
      "content": "<p><code>I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices</code><br>\nInstead of mixing them manually, co-visitation matrices can be a feature of LGBM Ranker</p>",
      "rawMarkdown": "`I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices`\nInstead of mixing them manually, co-visitation matrices can be a feature of LGBM Ranker",
      "votes": null
    },
    {
      "id": "2118336",
      "postDate": "01/28/2023 00:15:32",
      "content": "<p>Sorry I didn't mention that , I used some features from the co-visitation matrices, but maybe I should extract more valuable features ! thanks <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a>  ! :)</p>",
      "rawMarkdown": "Sorry I didn't mention that , I used some features from the co-visitation matrices, but maybe I should extract more valuable features ! thanks @sirius81  ! :)",
      "votes": null
    },
    {
      "id": "2118342",
      "postDate": "01/28/2023 00:21:51",
      "content": "<p><code>which means that the heuristic approach is better than my model</code><br>\nAll the heuristic approaches can be features of model. Usually, a model with those features performs better than the heuristic approaches.<br>\nBesides extracting more features, how to generate the candidates and make training samples may be also important for the performance of ranker.</p>",
      "rawMarkdown": "`which means that the heuristic approach is better than my model`\nAll the heuristic approaches can be features of model. Usually, a model with those features performs better than the heuristic approaches.\nBesides extracting more features, how to generate the candidates and make training samples may be also important for the performance of ranker.",
      "votes": null
    },
    {
      "id": "2118600",
      "postDate": "01/28/2023 06:19:19",
      "content": "<p>I got 0.565 too when I train orders model mark orders or carts to the ground truth. I got 0.572 when I remove ground truth came from carts. maybe helpful for u</p>",
      "rawMarkdown": "I got 0.565 too when I train orders model mark orders or carts to the ground truth. I got 0.572 when I remove ground truth came from carts. maybe helpful for u",
      "votes": null
    },
    {
      "id": "2118888",
      "postDate": "01/28/2023 10:58:43",
      "content": "<p>I got 0.580 with a model for cart prediction. I added a feature that is the order of aid in rerank model and so gbt model works better.</p>",
      "rawMarkdown": "I got 0.580 with a model for cart prediction. I added a feature that is the order of aid in rerank model and so gbt model works better.",
      "votes": null
    },
    {
      "id": "2118900",
      "postDate": "01/28/2023 11:23:14",
      "content": "<p>In fact, I scored 0.565 using only ground truth from orders. But thanks for the hint </p>",
      "rawMarkdown": "In fact, I scored 0.565 using only ground truth from orders. But thanks for the hint",
      "votes": null
    },
    {
      "id": "2119595",
      "postDate": "01/28/2023 23:37:45",
      "content": "<p>There are some important skills from recall to rank. I will share my codes after the competition. Perhaps you will find the points in our previous codes  <a href=\"https://github.com/ChuanyuXue/CIKM-2019-AnalytiCup\" target=\"_blank\">https://github.com/ChuanyuXue/CIKM-2019-AnalytiCup</a></p>\n<p>For example, page 12 in \"答辩ppt.pptx\"</p>",
      "rawMarkdown": "There are some important skills from recall to rank. I will share my codes after the competition. Perhaps you will find the points in our previous codes  https://github.com/ChuanyuXue/CIKM-2019-AnalytiCup\n\nFor example, page 12 in \"答辩ppt.pptx\"",
      "votes": null
    },
    {
      "id": "2119795",
      "postDate": "01/29/2023 05:38:42",
      "content": "<p>I had a similar issue, and solved it with this question. Did I lose any information from manual approach to the ranker approach? If I'm a human, and I'm trying to rank the candidates, can I perform as well as the heuristic? If not, what simple feature can I add to get at least the same performance?</p>",
      "rawMarkdown": "I had a similar issue, and solved it with this question. Did I lose any information from manual approach to the ranker approach? If I'm a human, and I'm trying to rank the candidates, can I perform as well as the heuristic? If not, what simple feature can I add to get at least the same performance?",
      "votes": null
    },
    {
      "id": "2119865",
      "postDate": "01/29/2023 06:55:24",
      "content": "<p>I think you can check whether you did incorrect fillna. I fix it in my code and I just train the orders model, then my score up to 0.581.  Did you train three separate models for clicks/carts/orders now? I plan to do it right now.</p>",
      "rawMarkdown": "I think you can check whether you did incorrect fillna. I fix it in my code and I just train the orders model, then my score up to 0.581.  Did you train three separate models for clicks/carts/orders now? I plan to do it right now.",
      "votes": null
    },
    {
      "id": "2120621",
      "postDate": "01/29/2023 18:05:44",
      "content": "<p>Did you consider it as a float or integer ?</p>",
      "rawMarkdown": "Did you consider it as a float or integer ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2118330,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "01/28/2023 00:09:52",
      "content": "<p><code>I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices</code><br>\nInstead of mixing them manually, co-visitation matrices can be a feature of LGBM Ranker</p>",
      "votes": null,
      "replies": [
        {
          "id": 2118336,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/28/2023 00:15:32",
          "content": "<p>Sorry I didn't mention that , I used some features from the co-visitation matrices, but maybe I should extract more valuable features ! thanks <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a>  ! :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2118342,
              "author_name": "sirius81",
              "author_url": "",
              "post_date": "01/28/2023 00:21:51",
              "content": "<p><code>which means that the heuristic approach is better than my model</code><br>\nAll the heuristic approaches can be features of model. Usually, a model with those features performs better than the heuristic approaches.<br>\nBesides extracting more features, how to generate the candidates and make training samples may be also important for the performance of ranker.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2118600,
      "author_name": "earlee0412",
      "author_url": "",
      "post_date": "01/28/2023 06:19:19",
      "content": "<p>I got 0.565 too when I train orders model mark orders or carts to the ground truth. I got 0.572 when I remove ground truth came from carts. maybe helpful for u</p>",
      "votes": null,
      "replies": [
        {
          "id": 2118900,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/28/2023 11:23:14",
          "content": "<p>In fact, I scored 0.565 using only ground truth from orders. But thanks for the hint </p>",
          "votes": null,
          "replies": [
            {
              "id": 2119865,
              "author_name": "earlee0412",
              "author_url": "",
              "post_date": "01/29/2023 06:55:24",
              "content": "<p>I think you can check whether you did incorrect fillna. I fix it in my code and I just train the orders model, then my score up to 0.581.  Did you train three separate models for clicks/carts/orders now? I plan to do it right now.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2118888,
      "author_name": "bibanh",
      "author_url": "",
      "post_date": "01/28/2023 10:58:43",
      "content": "<p>I got 0.580 with a model for cart prediction. I added a feature that is the order of aid in rerank model and so gbt model works better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2120621,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/29/2023 18:05:44",
          "content": "<p>Did you consider it as a float or integer ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2119595,
      "author_name": "bruceqdu",
      "author_url": "",
      "post_date": "01/28/2023 23:37:45",
      "content": "<p>There are some important skills from recall to rank. I will share my codes after the competition. Perhaps you will find the points in our previous codes  <a href=\"https://github.com/ChuanyuXue/CIKM-2019-AnalytiCup\" target=\"_blank\">https://github.com/ChuanyuXue/CIKM-2019-AnalytiCup</a></p>\n<p>For example, page 12 in \"答辩ppt.pptx\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2119795,
      "author_name": "datadote",
      "author_url": "",
      "post_date": "01/29/2023 05:38:42",
      "content": "<p>I had a similar issue, and solved it with this question. Did I lose any information from manual approach to the ranker approach? If I'm a human, and I'm trying to rank the candidates, can I perform as well as the heuristic? If not, what simple feature can I add to get at least the same performance?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2118324": "Hello community, I know I posted a lot of topics, but  I found this competition very interesting and at the same time rather intriguing, given that it's the first time on kaggle that I've been so involved in a competition.\n\nI'm writing this topic because even when using a GBDT Ranker with the appropriate features ( I suppose haha ), **I couldn't improve my recall over 0.565 using the Ranker**, which means that the heuristic approach is better than my model.\n\nI used:\n- Items and Users Features.\n- Interactions Features.\n- Similarities between entities.\n- GroupKFOLD.\n- Feature selection\n\nI took care of :\n- Not include the future in the training --> NO LEAK !!\n\nI tried:\n- DART / GOSS Booster.\n- Regularization using L1 L2 / bagging / Feature sampling.\n\n\n\nAnother point, is that my last model scored **0.667 on orders** in my local validation, but It still scored  0.565 on LB, using the Ranker, no matter if I  use 100  / 300 or  / 1000 estimators.\n\nHowever, I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices, which means that the Ranker scores better than the heuristic model in some cases, which also means that the model is doing the job somehow !\n\nA big thanks to the kagglers of this competitions for all the tips and sharing !",
    "2118330": "`I could score 0.579 using a mixture of LGBM Ranker and co-visitation matrices`\nInstead of mixing them manually, co-visitation matrices can be a feature of LGBM Ranker",
    "2118336": "Sorry I didn't mention that , I used some features from the co-visitation matrices, but maybe I should extract more valuable features ! thanks @sirius81  ! :)",
    "2118342": "`which means that the heuristic approach is better than my model`\nAll the heuristic approaches can be features of model. Usually, a model with those features performs better than the heuristic approaches.\nBesides extracting more features, how to generate the candidates and make training samples may be also important for the performance of ranker.",
    "2118600": "I got 0.565 too when I train orders model mark orders or carts to the ground truth. I got 0.572 when I remove ground truth came from carts. maybe helpful for u",
    "2118888": "I got 0.580 with a model for cart prediction. I added a feature that is the order of aid in rerank model and so gbt model works better.",
    "2118900": "In fact, I scored 0.565 using only ground truth from orders. But thanks for the hint",
    "2119595": "There are some important skills from recall to rank. I will share my codes after the competition. Perhaps you will find the points in our previous codes  https://github.com/ChuanyuXue/CIKM-2019-AnalytiCup\n\nFor example, page 12 in \"答辩ppt.pptx\"",
    "2119795": "I had a similar issue, and solved it with this question. Did I lose any information from manual approach to the ranker approach? If I'm a human, and I'm trying to rank the candidates, can I perform as well as the heuristic? If not, what simple feature can I add to get at least the same performance?",
    "2119865": "I think you can check whether you did incorrect fillna. I fix it in my code and I just train the orders model, then my score up to 0.581.  Did you train three separate models for clicks/carts/orders now? I plan to do it right now.",
    "2120621": "Did you consider it as a float or integer ?"
  },
  "source": "meta"
}