{
  "id": 322322,
  "title": "How to choose negative samples",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/322322",
  "author_name": "",
  "post_date": "2022-05-01T12:55:59.877164700Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm new to Kaggle. I use itemcf as recall mothod(recommend customer with high similarity with <br>\nhistory buying articles,the similarity measured by how many customer  buy both  articles).Use LGBMRanker to rank candidates.For every customer who bought something in the last two week ,  I use what they bought in the last two week as positive sample labeling 1 ,and  some  last articles recalled and the most unpopular articles in the last two week as negative sample labeling 0,<br>\nthese three part proportion are about 1:1:2.I got 0.0047 score,which is lower than choose  top popular  12 articles  for all customers(0.0054 score).So I am so confused about How to choose negative sample.<br>\nAny comments from you would be very helpful.Thanks.</p>",
  "messages": [
    {
      "id": "1773793",
      "postDate": "05/01/2022 12:55:59",
      "content": "<p>I'm new to Kaggle. I use itemcf as recall mothod(recommend customer with high similarity with <br>\nhistory buying articles,the similarity measured by how many customer  buy both  articles).Use LGBMRanker to rank candidates.For every customer who bought something in the last two week ,  I use what they bought in the last two week as positive sample labeling 1 ,and  some  last articles recalled and the most unpopular articles in the last two week as negative sample labeling 0,<br>\nthese three part proportion are about 1:1:2.I got 0.0047 score,which is lower than choose  top popular  12 articles  for all customers(0.0054 score).So I am so confused about How to choose negative sample.<br>\nAny comments from you would be very helpful.Thanks.</p>",
      "rawMarkdown": "I'm new to Kaggle. I use itemcf as recall mothod(recommend customer with high similarity with \nhistory buying articles,the similarity measured by how many customer  buy both  articles).Use LGBMRanker to rank candidates.For every customer who bought something in the last two week ,  I use what they bought in the last two week as positive sample labeling 1 ,and  some  last articles recalled and the most unpopular articles in the last two week as negative sample labeling 0,\nthese three part proportion are about 1:1:2.I got 0.0047 score,which is lower than choose  top popular  12 articles  for all customers(0.0054 score).So I am so confused about How to choose negative sample.\nAny comments from you would be very helpful.Thanks.",
      "votes": null
    },
    {
      "id": "1773996",
      "postDate": "05/01/2022 17:13:40",
      "content": "<p>In general, you want your training data to match the data you'll be predicting on for the LB submission.</p>\n<p>You don't have labels for the LB week, so you'll generate candidates blindly- some will be positive, most will be negative.<br>\nSimilarly, when training/evaluating, you just want to generate candidates (for example, for your strategy, getting 50 most similar items using itemcf), and just label them as positive/negative based on whether they're in the evaluation week.</p>\n<p>(You want to use one week for the labels, not two weeks, so that it's similar to the LB sub data)</p>",
      "rawMarkdown": "In general, you want your training data to match the data you'll be predicting on for the LB submission.\n\nYou don't have labels for the LB week, so you'll generate candidates blindly- some will be positive, most will be negative.\nSimilarly, when training/evaluating, you just want to generate candidates (for example, for your strategy, getting 50 most similar items using itemcf), and just label them as positive/negative based on whether they're in the evaluation week.\n\n(You want to use one week for the labels, not two weeks, so that it's similar to the LB sub data)",
      "votes": null
    },
    {
      "id": "1774262",
      "postDate": "05/02/2022 03:01:48",
      "content": "<p>Thanks for your patience to reply.<br>\nI will change to last one week and try again.<br>\nAnd I also want to ask whether choosing negative samples from the most unpopular articles  is meaningful,or just choosing popular articles which customer didn't buy as negative samples?</p>",
      "rawMarkdown": "Thanks for your patience to reply.\nI will change to last one week and try again.\nAnd I also want to ask whether choosing negative samples from the most unpopular articles  is meaningful,or just choosing popular articles which customer didn't buy as negative samples?",
      "votes": null
    },
    {
      "id": "1774832",
      "postDate": "05/02/2022 13:32:33",
      "content": "<p>I don't think you should \"choose\" negative examples - you just choose promising candidates, and the ones that are negative are negative.</p>",
      "rawMarkdown": "I don't think you should \"choose\" negative examples - you just choose promising candidates, and the ones that are negative are negative.",
      "votes": null
    },
    {
      "id": "1775517",
      "postDate": "05/03/2022 05:55:50",
      "content": "<p>What \"the ones that are negative are negative \" mean? I can't understand.LGBMRanker need data labeled 0 and 1.How should I get label 0 when transactions given are all 1? </p>",
      "rawMarkdown": "What \"the ones that are negative are negative \" mean? I can't understand.LGBMRanker need data labeled 0 and 1.How should I get label 0 when transactions given are all 1?",
      "votes": null
    },
    {
      "id": "1775652",
      "postDate": "05/03/2022 08:33:00",
      "content": "<p>As a clarification, negative observations could be interpreted as the most likely to be bought items, which however were not bought in that basket (for training). For application, they can be considered as most likely candidates. This is how I thought of the problem and it really helped me build more strategies. In general, I was suggested not to create fake-purchased items (fake y = 1). I hope it helps!</p>",
      "rawMarkdown": "As a clarification, negative observations could be interpreted as the most likely to be bought items, which however were not bought in that basket (for training). For application, they can be considered as most likely candidates. This is how I thought of the problem and it really helped me build more strategies. In general, I was suggested not to create fake-purchased items (fake y = 1). I hope it helps!",
      "votes": null
    },
    {
      "id": "1775683",
      "postDate": "05/03/2022 09:14:10",
      "content": "<p>Thanks a lot.I understand that items the most likely to buy but actually not to buy can be  as  negative  y=0\"</p>",
      "rawMarkdown": "Thanks a lot.I understand that items the most likely to buy but actually not to buy can be  as  negative  y=0\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1773996,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/01/2022 17:13:40",
      "content": "<p>In general, you want your training data to match the data you'll be predicting on for the LB submission.</p>\n<p>You don't have labels for the LB week, so you'll generate candidates blindly- some will be positive, most will be negative.<br>\nSimilarly, when training/evaluating, you just want to generate candidates (for example, for your strategy, getting 50 most similar items using itemcf), and just label them as positive/negative based on whether they're in the evaluation week.</p>\n<p>(You want to use one week for the labels, not two weeks, so that it's similar to the LB sub data)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1774262,
          "author_name": "jiahongxie",
          "author_url": "",
          "post_date": "05/02/2022 03:01:48",
          "content": "<p>Thanks for your patience to reply.<br>\nI will change to last one week and try again.<br>\nAnd I also want to ask whether choosing negative samples from the most unpopular articles  is meaningful,or just choosing popular articles which customer didn't buy as negative samples?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1774832,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/02/2022 13:32:33",
          "content": "<p>I don't think you should \"choose\" negative examples - you just choose promising candidates, and the ones that are negative are negative.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1775517,
          "author_name": "jiahongxie",
          "author_url": "",
          "post_date": "05/03/2022 05:55:50",
          "content": "<p>What \"the ones that are negative are negative \" mean? I can't understand.LGBMRanker need data labeled 0 and 1.How should I get label 0 when transactions given are all 1? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1775652,
      "author_name": "lorenzopagliaro01",
      "author_url": "",
      "post_date": "05/03/2022 08:33:00",
      "content": "<p>As a clarification, negative observations could be interpreted as the most likely to be bought items, which however were not bought in that basket (for training). For application, they can be considered as most likely candidates. This is how I thought of the problem and it really helped me build more strategies. In general, I was suggested not to create fake-purchased items (fake y = 1). I hope it helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1775683,
          "author_name": "jiahongxie",
          "author_url": "",
          "post_date": "05/03/2022 09:14:10",
          "content": "<p>Thanks a lot.I understand that items the most likely to buy but actually not to buy can be  as  negative  y=0\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1773793": "I'm new to Kaggle. I use itemcf as recall mothod(recommend customer with high similarity with \nhistory buying articles,the similarity measured by how many customer  buy both  articles).Use LGBMRanker to rank candidates.For every customer who bought something in the last two week ,  I use what they bought in the last two week as positive sample labeling 1 ,and  some  last articles recalled and the most unpopular articles in the last two week as negative sample labeling 0,\nthese three part proportion are about 1:1:2.I got 0.0047 score,which is lower than choose  top popular  12 articles  for all customers(0.0054 score).So I am so confused about How to choose negative sample.\nAny comments from you would be very helpful.Thanks.",
    "1773996": "In general, you want your training data to match the data you'll be predicting on for the LB submission.\n\nYou don't have labels for the LB week, so you'll generate candidates blindly- some will be positive, most will be negative.\nSimilarly, when training/evaluating, you just want to generate candidates (for example, for your strategy, getting 50 most similar items using itemcf), and just label them as positive/negative based on whether they're in the evaluation week.\n\n(You want to use one week for the labels, not two weeks, so that it's similar to the LB sub data)",
    "1774262": "Thanks for your patience to reply.\nI will change to last one week and try again.\nAnd I also want to ask whether choosing negative samples from the most unpopular articles  is meaningful,or just choosing popular articles which customer didn't buy as negative samples?",
    "1774832": "I don't think you should \"choose\" negative examples - you just choose promising candidates, and the ones that are negative are negative.",
    "1775517": "What \"the ones that are negative are negative \" mean? I can't understand.LGBMRanker need data labeled 0 and 1.How should I get label 0 when transactions given are all 1?",
    "1775652": "As a clarification, negative observations could be interpreted as the most likely to be bought items, which however were not bought in that basket (for training). For application, they can be considered as most likely candidates. This is how I thought of the problem and it really helped me build more strategies. In general, I was suggested not to create fake-purchased items (fake y = 1). I hope it helps!",
    "1775683": "Thanks a lot.I understand that items the most likely to buy but actually not to buy can be  as  negative  y=0\""
  },
  "source": "meta"
}