{
  "id": 321649,
  "title": "Why the dnn recall system cause bad result for LightgbmRanker?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/321649",
  "author_name": "",
  "post_date": "2022-04-28T01:50:02.415383200Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The baseline is** Radek's LGBMRanker starter-pack** with only put lastbuy items and random items as negative candidates.And the map12 is 0.024. I used dnn recall system which proposed by Youtube work in 2016.In CV, the recall is 0.065 in top 50 for test week. But when I put the **recall candidates **to LightgbmRanker as test candidates. .The map12 drops from 0.024 to 0.009. In fact, the  situation appears when I put bestsalers as test candidates, and the map12 drops to 0.0189.</p>",
  "messages": [
    {
      "id": "1770108",
      "postDate": "04/28/2022 01:50:02",
      "content": "<p>The baseline is** Radek's LGBMRanker starter-pack** with only put lastbuy items and random items as negative candidates.And the map12 is 0.024. I used dnn recall system which proposed by Youtube work in 2016.In CV, the recall is 0.065 in top 50 for test week. But when I put the **recall candidates **to LightgbmRanker as test candidates. .The map12 drops from 0.024 to 0.009. In fact, the  situation appears when I put bestsalers as test candidates, and the map12 drops to 0.0189.</p>",
      "rawMarkdown": "The baseline is** Radek's LGBMRanker starter-pack** with only put lastbuy items and random items as negative candidates.And the map12 is 0.024. I used dnn recall system which proposed by Youtube work in 2016.In CV, the recall is 0.065 in top 50 for test week. But when I put the **recall candidates **to LightgbmRanker as test candidates. .The map12 drops from 0.024 to 0.009. In fact, the  situation appears when I put bestsalers as test candidates, and the map12 drops to 0.0189.",
      "votes": null
    },
    {
      "id": "1771089",
      "postDate": "04/28/2022 21:55:06",
      "content": "<p>Try using user and item embeddings (or user-item similarity calculated from the embeddings) from the YouTubeDNN as your LGB features😀 My thought is that DL models for retrieve/recall tasks may only perform well when k is big enough, and 12 is too small.</p>",
      "rawMarkdown": "Try using user and item embeddings (or user-item similarity calculated from the embeddings) from the YouTubeDNN as your LGB features😀 My thought is that DL models for retrieve/recall tasks may only perform well when k is big enough, and 12 is too small.",
      "votes": null
    },
    {
      "id": "1771219",
      "postDate": "04/29/2022 02:25:55",
      "content": "<p>Thanks for you help! I will try your strategy, and I want to konw whether it's useful to use DNN to generate negative samples for LGB?</p>",
      "rawMarkdown": "Thanks for you help! I will try your strategy, and I want to konw whether it's useful to use DNN to generate negative samples for LGB?",
      "votes": null
    },
    {
      "id": "1771222",
      "postDate": "04/29/2022 02:29:37",
      "content": "<p>Sorry I haven't tried it yet, but I think it is worth a try!</p>",
      "rawMarkdown": "Sorry I haven't tried it yet, but I think it is worth a try!",
      "votes": null
    },
    {
      "id": "1771234",
      "postDate": "04/29/2022 02:52:06",
      "content": "<p>Hello, thanks for your sharing, I also use embeddings to measure the similarity between customers and articles, and use similarity as one of features for LGBM, but I find the importance of this similarity feature is smaller than articles popularity and other features, I don't konw if it is normal, have you had similar problems before?</p>",
      "rawMarkdown": "Hello, thanks for your sharing, I also use embeddings to measure the similarity between customers and articles, and use similarity as one of features for LGBM, but I find the importance of this similarity feature is smaller than articles popularity and other features, I don't konw if it is normal, have you had similar problems before?",
      "votes": null
    },
    {
      "id": "1771236",
      "postDate": "04/29/2022 03:02:29",
      "content": "<p>My case is that the feature importance of similarity is high but the CV score won't change much if it's removed.  Maybe it depends on how you trained the DL model and its performance, or how strong is your article popularity features, but I haven't dug into it yet.</p>",
      "rawMarkdown": "My case is that the feature importance of similarity is high but the CV score won't change much if it's removed.  Maybe it depends on how you trained the DL model and its performance, or how strong is your article popularity features, but I haven't dug into it yet.",
      "votes": null
    },
    {
      "id": "1771262",
      "postDate": "04/29/2022 03:56:56",
      "content": "<p>Thanks for your reply!</p>",
      "rawMarkdown": "Thanks for your reply!",
      "votes": null
    },
    {
      "id": "1773177",
      "postDate": "04/30/2022 23:49:27",
      "content": "<p>It looks like the dnn recall system is doing a very poor job of detecting negative examples. This is causing the LightgbmRanker to perform much worse than it should.</p>",
      "rawMarkdown": "It looks like the dnn recall system is doing a very poor job of detecting negative examples. This is causing the LightgbmRanker to perform much worse than it should.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1771089,
      "author_name": "weipengzhang",
      "author_url": "",
      "post_date": "04/28/2022 21:55:06",
      "content": "<p>Try using user and item embeddings (or user-item similarity calculated from the embeddings) from the YouTubeDNN as your LGB features😀 My thought is that DL models for retrieve/recall tasks may only perform well when k is big enough, and 12 is too small.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1771219,
          "author_name": "dengniewei",
          "author_url": "",
          "post_date": "04/29/2022 02:25:55",
          "content": "<p>Thanks for you help! I will try your strategy, and I want to konw whether it's useful to use DNN to generate negative samples for LGB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771222,
          "author_name": "weipengzhang",
          "author_url": "",
          "post_date": "04/29/2022 02:29:37",
          "content": "<p>Sorry I haven't tried it yet, but I think it is worth a try!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771234,
          "author_name": "rock139",
          "author_url": "",
          "post_date": "04/29/2022 02:52:06",
          "content": "<p>Hello, thanks for your sharing, I also use embeddings to measure the similarity between customers and articles, and use similarity as one of features for LGBM, but I find the importance of this similarity feature is smaller than articles popularity and other features, I don't konw if it is normal, have you had similar problems before?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771236,
          "author_name": "weipengzhang",
          "author_url": "",
          "post_date": "04/29/2022 03:02:29",
          "content": "<p>My case is that the feature importance of similarity is high but the CV score won't change much if it's removed.  Maybe it depends on how you trained the DL model and its performance, or how strong is your article popularity features, but I haven't dug into it yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771262,
          "author_name": "rock139",
          "author_url": "",
          "post_date": "04/29/2022 03:56:56",
          "content": "<p>Thanks for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1773177,
      "author_name": "",
      "author_url": "",
      "post_date": "04/30/2022 23:49:27",
      "content": "<p>It looks like the dnn recall system is doing a very poor job of detecting negative examples. This is causing the LightgbmRanker to perform much worse than it should.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1770108": "The baseline is** Radek's LGBMRanker starter-pack** with only put lastbuy items and random items as negative candidates.And the map12 is 0.024. I used dnn recall system which proposed by Youtube work in 2016.In CV, the recall is 0.065 in top 50 for test week. But when I put the **recall candidates **to LightgbmRanker as test candidates. .The map12 drops from 0.024 to 0.009. In fact, the  situation appears when I put bestsalers as test candidates, and the map12 drops to 0.0189.",
    "1771089": "Try using user and item embeddings (or user-item similarity calculated from the embeddings) from the YouTubeDNN as your LGB features😀 My thought is that DL models for retrieve/recall tasks may only perform well when k is big enough, and 12 is too small.",
    "1771219": "Thanks for you help! I will try your strategy, and I want to konw whether it's useful to use DNN to generate negative samples for LGB?",
    "1771222": "Sorry I haven't tried it yet, but I think it is worth a try!",
    "1771234": "Hello, thanks for your sharing, I also use embeddings to measure the similarity between customers and articles, and use similarity as one of features for LGBM, but I find the importance of this similarity feature is smaller than articles popularity and other features, I don't konw if it is normal, have you had similar problems before?",
    "1771236": "My case is that the feature importance of similarity is high but the CV score won't change much if it's removed.  Maybe it depends on how you trained the DL model and its performance, or how strong is your article popularity features, but I haven't dug into it yet.",
    "1771262": "Thanks for your reply!",
    "1773177": "It looks like the dnn recall system is doing a very poor job of detecting negative examples. This is causing the LightgbmRanker to perform much worse than it should."
  },
  "source": "meta"
}