{
  "id": 320810,
  "title": "Candidates and negative samples",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/320810",
  "author_name": "",
  "post_date": "2022-04-23T16:58:16.785310800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Many codes and discussions talk about adding candidates to dataset, negative samples for lightgbmranker model.<br>\nBut I don’t understand difference of them(They just concatenated to dataset with their labels set 0). <br>\nIn my opinion, feeding candidates like weekly bestsellers can show model that bestsellers are preferred to customers, but their label is 0 , same as negative sample’s label.<br>\nHow this method work?(refer to code below)<br>\n<a href=\"https://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack\" target=\"_blank\">https://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack</a></p>",
  "messages": [
    {
      "id": "1765575",
      "postDate": "04/23/2022 16:58:16",
      "content": "<p>Many codes and discussions talk about adding candidates to dataset, negative samples for lightgbmranker model.<br>\nBut I don’t understand difference of them(They just concatenated to dataset with their labels set 0). <br>\nIn my opinion, feeding candidates like weekly bestsellers can show model that bestsellers are preferred to customers, but their label is 0 , same as negative sample’s label.<br>\nHow this method work?(refer to code below)<br>\n<a href=\"https://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack\" target=\"_blank\">https://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack</a></p>",
      "rawMarkdown": "Many codes and discussions talk about adding candidates to dataset, negative samples for lightgbmranker model.\nBut I don’t understand difference of them(They just concatenated to dataset with their labels set 0). \nIn my opinion, feeding candidates like weekly bestsellers can show model that bestsellers are preferred to customers, but their label is 0 , same as negative sample’s label.\nHow this method work?(refer to code below)\nhttps://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack",
      "votes": null
    },
    {
      "id": "1766163",
      "postDate": "04/24/2022 09:09:58",
      "content": "<p>If runtime performance was not an issue, the model should predict for each combination of article and customer whether the customer bought that article or not in the week to be predicted.</p>\n<p>Using a brute force approach this would result in 105,542 * 1,371,980 = 144,801,513,160 records just for the final prediction. If you can predict 300,000 records per second this would take 15 years for your submission.</p>\n<p>Training a model on all weeks for these combinations would result in 105 * 144,801,513,160 = 15,204,158,881,800 training data records. If you can train 10,000 records per second the training would take 48,212 years.</p>\n<p>Hence, we need to reduce both the data sets for training and the final submission. Instead of considering all customers and all articles we only take a look at good candidates.</p>\n<p>The positive samples are combinations of articles and customer where the customer actually bought this article.</p>\n<p>Only using positive samples would result in a model that always predicts \"YES\". Hence, we add negative samples that are combinations of articles and customer where the customer actually did not buy this article.</p>\n<p>Apart from the feature engineering an important part of the competition is to create a good heuristic that generates candidate articles for each customer. You can create as many candidates as your machine is able to handle for training and prediction. The best selling articles are a good starting point. Other ideas:</p>\n<ul>\n<li>articles the customer purchased before</li>\n<li>articles of the same product the customer purchased before</li>\n<li>articles bought by similar customers</li>\n<li>articles that complement the articles the customer has previously purchased</li>\n</ul>",
      "rawMarkdown": "If runtime performance was not an issue, the model should predict for each combination of article and customer whether the customer bought that article or not in the week to be predicted.\n\nUsing a brute force approach this would result in 105,542 * 1,371,980 = 144,801,513,160 records just for the final prediction. If you can predict 300,000 records per second this would take 15 years for your submission.\n\nTraining a model on all weeks for these combinations would result in 105 * 144,801,513,160 = 15,204,158,881,800 training data records. If you can train 10,000 records per second the training would take 48,212 years.\n\nHence, we need to reduce both the data sets for training and the final submission. Instead of considering all customers and all articles we only take a look at good candidates.\n\nThe positive samples are combinations of articles and customer where the customer actually bought this article.\n\nOnly using positive samples would result in a model that always predicts \"YES\". Hence, we add negative samples that are combinations of articles and customer where the customer actually did not buy this article.\n\nApart from the feature engineering an important part of the competition is to create a good heuristic that generates candidate articles for each customer. You can create as many candidates as your machine is able to handle for training and prediction. The best selling articles are a good starting point. Other ideas:\n\n- articles the customer purchased before\n- articles of the same product the customer purchased before\n- articles bought by similar customers\n- articles that complement the articles the customer has previously purchased",
      "votes": null
    },
    {
      "id": "1766360",
      "postDate": "04/24/2022 13:26:47",
      "content": "<p>Thank you for your reply!!!  Since I'm new to recsys, it sound really weird, though. It's quite strange for me giving negative label to data that we assumed to be related with ground truth.<br>\nIf my understanding is correct, generating candidates as many as possible can help model, whatever the label is?</p>",
      "rawMarkdown": "Thank you for your reply!!!  Since I'm new to recsys, it sound really weird, though. It's quite strange for me giving negative label to data that we assumed to be related with ground truth.\nIf my understanding is correct, generating candidates as many as possible can help model, whatever the label is?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1766163,
      "author_name": "tbierhance",
      "author_url": "",
      "post_date": "04/24/2022 09:09:58",
      "content": "<p>If runtime performance was not an issue, the model should predict for each combination of article and customer whether the customer bought that article or not in the week to be predicted.</p>\n<p>Using a brute force approach this would result in 105,542 * 1,371,980 = 144,801,513,160 records just for the final prediction. If you can predict 300,000 records per second this would take 15 years for your submission.</p>\n<p>Training a model on all weeks for these combinations would result in 105 * 144,801,513,160 = 15,204,158,881,800 training data records. If you can train 10,000 records per second the training would take 48,212 years.</p>\n<p>Hence, we need to reduce both the data sets for training and the final submission. Instead of considering all customers and all articles we only take a look at good candidates.</p>\n<p>The positive samples are combinations of articles and customer where the customer actually bought this article.</p>\n<p>Only using positive samples would result in a model that always predicts \"YES\". Hence, we add negative samples that are combinations of articles and customer where the customer actually did not buy this article.</p>\n<p>Apart from the feature engineering an important part of the competition is to create a good heuristic that generates candidate articles for each customer. You can create as many candidates as your machine is able to handle for training and prediction. The best selling articles are a good starting point. Other ideas:</p>\n<ul>\n<li>articles the customer purchased before</li>\n<li>articles of the same product the customer purchased before</li>\n<li>articles bought by similar customers</li>\n<li>articles that complement the articles the customer has previously purchased</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1766360,
          "author_name": "lionsheep24",
          "author_url": "",
          "post_date": "04/24/2022 13:26:47",
          "content": "<p>Thank you for your reply!!!  Since I'm new to recsys, it sound really weird, though. It's quite strange for me giving negative label to data that we assumed to be related with ground truth.<br>\nIf my understanding is correct, generating candidates as many as possible can help model, whatever the label is?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1765575": "Many codes and discussions talk about adding candidates to dataset, negative samples for lightgbmranker model.\nBut I don’t understand difference of them(They just concatenated to dataset with their labels set 0). \nIn my opinion, feeding candidates like weekly bestsellers can show model that bestsellers are preferred to customers, but their label is 0 , same as negative sample’s label.\nHow this method work?(refer to code below)\nhttps://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack",
    "1766163": "If runtime performance was not an issue, the model should predict for each combination of article and customer whether the customer bought that article or not in the week to be predicted.\n\nUsing a brute force approach this would result in 105,542 * 1,371,980 = 144,801,513,160 records just for the final prediction. If you can predict 300,000 records per second this would take 15 years for your submission.\n\nTraining a model on all weeks for these combinations would result in 105 * 144,801,513,160 = 15,204,158,881,800 training data records. If you can train 10,000 records per second the training would take 48,212 years.\n\nHence, we need to reduce both the data sets for training and the final submission. Instead of considering all customers and all articles we only take a look at good candidates.\n\nThe positive samples are combinations of articles and customer where the customer actually bought this article.\n\nOnly using positive samples would result in a model that always predicts \"YES\". Hence, we add negative samples that are combinations of articles and customer where the customer actually did not buy this article.\n\nApart from the feature engineering an important part of the competition is to create a good heuristic that generates candidate articles for each customer. You can create as many candidates as your machine is able to handle for training and prediction. The best selling articles are a good starting point. Other ideas:\n\n- articles the customer purchased before\n- articles of the same product the customer purchased before\n- articles bought by similar customers\n- articles that complement the articles the customer has previously purchased",
    "1766360": "Thank you for your reply!!!  Since I'm new to recsys, it sound really weird, though. It's quite strange for me giving negative label to data that we assumed to be related with ground truth.\nIf my understanding is correct, generating candidates as many as possible can help model, whatever the label is?"
  },
  "source": "meta"
}