{
  "id": 322493,
  "title": "LGBMRanker feature importance",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/322493",
  "author_name": "",
  "post_date": "2022-05-02T14:22:33.432633100Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello everyone, I've been stuck on LGBMRanker in recent days. The map12 is decreased after ranker, and I find the article popularity is the most important feature in my ranker model, however, I think the most important feature should be the similarities between articles and customers, but I don't know what's wrong with my rank feature. I'm really confused about this problem😭, did you have met this problem?</p>\n<p>BTW, the similarity features I used between customers and articles are cosine similarities of customer embeedings and article embeedings.</p>",
  "messages": [
    {
      "id": "1774867",
      "postDate": "05/02/2022 14:22:33",
      "content": "<p>Hello everyone, I've been stuck on LGBMRanker in recent days. The map12 is decreased after ranker, and I find the article popularity is the most important feature in my ranker model, however, I think the most important feature should be the similarities between articles and customers, but I don't know what's wrong with my rank feature. I'm really confused about this problem😭, did you have met this problem?</p>\n<p>BTW, the similarity features I used between customers and articles are cosine similarities of customer embeedings and article embeedings.</p>",
      "rawMarkdown": "Hello everyone, I've been stuck on LGBMRanker in recent days. The map12 is decreased after ranker, and I find the article popularity is the most important feature in my ranker model, however, I think the most important feature should be the similarities between articles and customers, but I don't know what's wrong with my rank feature. I'm really confused about this problem😭, did you have met this problem?\n\nBTW, the similarity features I used between customers and articles are cosine similarities of customer embeedings and article embeedings.",
      "votes": null
    },
    {
      "id": "1775432",
      "postDate": "05/03/2022 02:48:31",
      "content": "<p>I think it could make sense that article popularity is still the most important feature.</p>\n<p>I also found that LGBMRanker was descreasing the score, even on the training set. But somehow it's still giving me a better score on the LB.<br>\nStill haven't figured it out - <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374\" target=\"_blank\">see my post here</a>.</p>",
      "rawMarkdown": "I think it could make sense that article popularity is still the most important feature.\n\nI also found that LGBMRanker was descreasing the score, even on the training set. But somehow it's still giving me a better score on the LB.\nStill haven't figured it out - [see my post here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374).",
      "votes": null
    },
    {
      "id": "1775437",
      "postDate": "05/03/2022 03:12:17",
      "content": "<p>Thanks for your reply! I think it may because I haven't find the right way to mesuare the similarity between customer and article, and I'm going to try deep learning method like youtubednn. By the way, may I ask what your LB score is before ranker? Thank you!</p>",
      "rawMarkdown": "Thanks for your reply! I think it may because I haven't find the right way to mesuare the similarity between customer and article, and I'm going to try deep learning method like youtubednn. By the way, may I ask what your LB score is before ranker? Thank you!",
      "votes": null
    },
    {
      "id": "1775546",
      "postDate": "05/03/2022 06:25:08",
      "content": "<p>Hope you are accounting for recency of popularity as well. What I mean is what is more popular in September 2020 is far more important than what was popular in September 2019. <br>\nIt is also possible the most popular item overall has not made a single sale in the last month of the dataset. So it may not account for much.</p>",
      "rawMarkdown": "Hope you are accounting for recency of popularity as well. What I mean is what is more popular in September 2020 is far more important than what was popular in September 2019. \nIt is also possible the most popular item overall has not made a single sale in the last month of the dataset. So it may not account for much.",
      "votes": null
    },
    {
      "id": "1776207",
      "postDate": "05/03/2022 19:16:15",
      "content": "<p>Highest I got without LGBMRanker was .0243</p>",
      "rawMarkdown": "Highest I got without LGBMRanker was .0243",
      "votes": null
    },
    {
      "id": "1777967",
      "postDate": "05/05/2022 00:55:50",
      "content": "<p><a href=\"https://www.kaggle.com/rock139\" target=\"_blank\">@rock139</a> - In my case, I figured it out - see <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">this post</a>.</p>",
      "rawMarkdown": "rock139 - In my case, I figured it out - see [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081).",
      "votes": null
    },
    {
      "id": "1778998",
      "postDate": "05/06/2022 01:50:09",
      "content": "<p>Congratulations!! I also met this problem when using cudf, which bothered me for a long time, and I'm glad you figured it out!<br>\nAnd my problem is also solved. Because there was something wrong with the way I verify the results. I thought the score has decreased, but in fact, the score has increased. This has troubled me for a long time, and maybe I should have more confidence in my code😂.</p>",
      "rawMarkdown": "Congratulations!! I also met this problem when using cudf, which bothered me for a long time, and I'm glad you figured it out!\nAnd my problem is also solved. Because there was something wrong with the way I verify the results. I thought the score has decreased, but in fact, the score has increased. This has troubled me for a long time, and maybe I should have more confidence in my code😂.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1775432,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/03/2022 02:48:31",
      "content": "<p>I think it could make sense that article popularity is still the most important feature.</p>\n<p>I also found that LGBMRanker was descreasing the score, even on the training set. But somehow it's still giving me a better score on the LB.<br>\nStill haven't figured it out - <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374\" target=\"_blank\">see my post here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1775437,
          "author_name": "rock139",
          "author_url": "",
          "post_date": "05/03/2022 03:12:17",
          "content": "<p>Thanks for your reply! I think it may because I haven't find the right way to mesuare the similarity between customer and article, and I'm going to try deep learning method like youtubednn. By the way, may I ask what your LB score is before ranker? Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1776207,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/03/2022 19:16:15",
          "content": "<p>Highest I got without LGBMRanker was .0243</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1777967,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/05/2022 00:55:50",
          "content": "<p><a href=\"https://www.kaggle.com/rock139\" target=\"_blank\">@rock139</a> - In my case, I figured it out - see <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">this post</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1778998,
          "author_name": "rock139",
          "author_url": "",
          "post_date": "05/06/2022 01:50:09",
          "content": "<p>Congratulations!! I also met this problem when using cudf, which bothered me for a long time, and I'm glad you figured it out!<br>\nAnd my problem is also solved. Because there was something wrong with the way I verify the results. I thought the score has decreased, but in fact, the score has increased. This has troubled me for a long time, and maybe I should have more confidence in my code😂.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1775546,
      "author_name": "srinivassateesh",
      "author_url": "",
      "post_date": "05/03/2022 06:25:08",
      "content": "<p>Hope you are accounting for recency of popularity as well. What I mean is what is more popular in September 2020 is far more important than what was popular in September 2019. <br>\nIt is also possible the most popular item overall has not made a single sale in the last month of the dataset. So it may not account for much.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1774867": "Hello everyone, I've been stuck on LGBMRanker in recent days. The map12 is decreased after ranker, and I find the article popularity is the most important feature in my ranker model, however, I think the most important feature should be the similarities between articles and customers, but I don't know what's wrong with my rank feature. I'm really confused about this problem😭, did you have met this problem?\n\nBTW, the similarity features I used between customers and articles are cosine similarities of customer embeedings and article embeedings.",
    "1775432": "I think it could make sense that article popularity is still the most important feature.\n\nI also found that LGBMRanker was descreasing the score, even on the training set. But somehow it's still giving me a better score on the LB.\nStill haven't figured it out - [see my post here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374).",
    "1775437": "Thanks for your reply! I think it may because I haven't find the right way to mesuare the similarity between customer and article, and I'm going to try deep learning method like youtubednn. By the way, may I ask what your LB score is before ranker? Thank you!",
    "1775546": "Hope you are accounting for recency of popularity as well. What I mean is what is more popular in September 2020 is far more important than what was popular in September 2019. \nIt is also possible the most popular item overall has not made a single sale in the last month of the dataset. So it may not account for much.",
    "1776207": "Highest I got without LGBMRanker was .0243",
    "1777967": "rock139 - In my case, I figured it out - see [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081).",
    "1778998": "Congratulations!! I also met this problem when using cudf, which bothered me for a long time, and I'm glad you figured it out!\nAnd my problem is also solved. Because there was something wrong with the way I verify the results. I thought the score has decreased, but in fact, the score has increased. This has troubled me for a long time, and maybe I should have more confidence in my code😂."
  },
  "source": "meta"
}