{
  "id": 376693,
  "title": "Do you expect better recall after using a Ranker ?",
  "url": "/competitions/otto-recommender-system/discussion/376693",
  "author_name": "",
  "post_date": "2023-01-07T20:45:30.133722200Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, I have a question about the last step of the pipeline which is the re-ranking of the top K predictions, it really got me confused.</p>\n<p>I'm using an  LGBMRanker using more than 60 features ( user/item/interactions features ) I got a validation score for the top20 predictions selected by the ranker  = 0.519 ( for clicks ) which is worse than the actual baseline that is around 0.525 ( for clicks ), whereas my metrics are :<br>\nNDCG@20 : 0.89<br>\nbinary loss : 0.3<br>\nproportion of negative sample : 0.2</p>\n<p>Also my clicks recall from the top 50 candidates is : 0.61</p>\n<p>Next, I noticed that the use of interactions features between users and items ( without leaks ), and its join ( with respect to \"session\" and \"aid\" ) with the train dataset ( dataset with exploded candidates in rows ) results into more than 90% of NaN in my interactions features.</p>\n<p>Do anyone have an idea about the problem that might be affect my validation score ? </p>",
  "messages": [
    {
      "id": "2090947",
      "postDate": "01/07/2023 20:45:30",
      "content": "<p>Hello, I have a question about the last step of the pipeline which is the re-ranking of the top K predictions, it really got me confused.</p>\n<p>I'm using an  LGBMRanker using more than 60 features ( user/item/interactions features ) I got a validation score for the top20 predictions selected by the ranker  = 0.519 ( for clicks ) which is worse than the actual baseline that is around 0.525 ( for clicks ), whereas my metrics are :<br>\nNDCG@20 : 0.89<br>\nbinary loss : 0.3<br>\nproportion of negative sample : 0.2</p>\n<p>Also my clicks recall from the top 50 candidates is : 0.61</p>\n<p>Next, I noticed that the use of interactions features between users and items ( without leaks ), and its join ( with respect to \"session\" and \"aid\" ) with the train dataset ( dataset with exploded candidates in rows ) results into more than 90% of NaN in my interactions features.</p>\n<p>Do anyone have an idea about the problem that might be affect my validation score ? </p>",
      "rawMarkdown": "Hello, I have a question about the last step of the pipeline which is the re-ranking of the top K predictions, it really got me confused.\n\nI'm using an  LGBMRanker using more than 60 features ( user/item/interactions features ) I got a validation score for the top20 predictions selected by the ranker  = 0.519 ( for clicks ) which is worse than the actual baseline that is around 0.525 ( for clicks ), whereas my metrics are :\nNDCG@20 : 0.89\nbinary loss : 0.3\nproportion of negative sample : 0.2\n\nAlso my clicks recall from the top 50 candidates is : 0.61\n\nNext, I noticed that the use of interactions features between users and items ( without leaks ), and its join ( with respect to \"session\" and \"aid\" ) with the train dataset ( dataset with exploded candidates in rows ) results into more than 90% of NaN in my interactions features.\n\nDo anyone have an idea about the problem that might be affect my validation score ?",
      "votes": null
    },
    {
      "id": "2091474",
      "postDate": "01/08/2023 12:21:30",
      "content": "<p>I guess something is wrong with your interaction features. You need something better than features with more than 90% NaN.<br>\nNote that most sessions for clicks are very short and have no buys in history. About 90% of sessions have less than 10 unique aids.<br>\nSo, if your feature relies on buys in history or needs many unique aids it will be NaN for most sessions.<br>\nMy first catboost model also had worse results than the baseline 0.519, but adding a few more interaction features improved model to 53,6%, and, I guess it can be improved further.</p>",
      "rawMarkdown": "I guess something is wrong with your interaction features. You need something better than features with more than 90% NaN.\nNote that most sessions for clicks are very short and have no buys in history. About 90% of sessions have less than 10 unique aids.\nSo, if your feature relies on buys in history or needs many unique aids it will be NaN for most sessions.\nMy first catboost model also had worse results than the baseline 0.519, but adding a few more interaction features improved model to 53,6%, and, I guess it can be improved further.",
      "votes": null
    },
    {
      "id": "2120881",
      "postDate": "01/29/2023 21:53:48",
      "content": "<p>Thanks ! however, when doing the left join between your candidates tables and interactions features, you ends up with a lot of NaNs anyways ? isn't ?</p>",
      "rawMarkdown": "Thanks ! however, when doing the left join between your candidates tables and interactions features, you ends up with a lot of NaNs anyways ? isn't ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2091474,
      "author_name": "artemfedorov",
      "author_url": "",
      "post_date": "01/08/2023 12:21:30",
      "content": "<p>I guess something is wrong with your interaction features. You need something better than features with more than 90% NaN.<br>\nNote that most sessions for clicks are very short and have no buys in history. About 90% of sessions have less than 10 unique aids.<br>\nSo, if your feature relies on buys in history or needs many unique aids it will be NaN for most sessions.<br>\nMy first catboost model also had worse results than the baseline 0.519, but adding a few more interaction features improved model to 53,6%, and, I guess it can be improved further.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2120881,
          "author_name": "zemirlinelynda",
          "author_url": "",
          "post_date": "01/29/2023 21:53:48",
          "content": "<p>Thanks ! however, when doing the left join between your candidates tables and interactions features, you ends up with a lot of NaNs anyways ? isn't ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2090947": "Hello, I have a question about the last step of the pipeline which is the re-ranking of the top K predictions, it really got me confused.\n\nI'm using an  LGBMRanker using more than 60 features ( user/item/interactions features ) I got a validation score for the top20 predictions selected by the ranker  = 0.519 ( for clicks ) which is worse than the actual baseline that is around 0.525 ( for clicks ), whereas my metrics are :\nNDCG@20 : 0.89\nbinary loss : 0.3\nproportion of negative sample : 0.2\n\nAlso my clicks recall from the top 50 candidates is : 0.61\n\nNext, I noticed that the use of interactions features between users and items ( without leaks ), and its join ( with respect to \"session\" and \"aid\" ) with the train dataset ( dataset with exploded candidates in rows ) results into more than 90% of NaN in my interactions features.\n\nDo anyone have an idea about the problem that might be affect my validation score ?",
    "2091474": "I guess something is wrong with your interaction features. You need something better than features with more than 90% NaN.\nNote that most sessions for clicks are very short and have no buys in history. About 90% of sessions have less than 10 unique aids.\nSo, if your feature relies on buys in history or needs many unique aids it will be NaN for most sessions.\nMy first catboost model also had worse results than the baseline 0.519, but adding a few more interaction features improved model to 53,6%, and, I guess it can be improved further.",
    "2120881": "Thanks ! however, when doing the left join between your candidates tables and interactions features, you ends up with a lot of NaNs anyways ? isn't ?"
  },
  "source": "meta"
}