{
  "id": 317981,
  "title": "positive, negative sample?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/317981",
  "author_name": "",
  "post_date": "2022-04-10T01:44:45.307633500Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all,<br>\nSince Recys is totally new for me, I don't understand positive, negative sampling in many threads.<br>\nRefer to many code examples, they just sample data and give label 0 or 1 to them, randomly.<br>\nWhat does it mean to train a model this way?</p>",
  "messages": [
    {
      "id": "1750714",
      "postDate": "04/10/2022 01:44:45",
      "content": "<p>Hi all,<br>\nSince Recys is totally new for me, I don't understand positive, negative sampling in many threads.<br>\nRefer to many code examples, they just sample data and give label 0 or 1 to them, randomly.<br>\nWhat does it mean to train a model this way?</p>",
      "rawMarkdown": "Hi all,\nSince Recys is totally new for me, I don't understand positive, negative sampling in many threads.\nRefer to many code examples, they just sample data and give label 0 or 1 to them, randomly.\nWhat does it mean to train a model this way?",
      "votes": null
    },
    {
      "id": "1752629",
      "postDate": "04/12/2022 02:11:07",
      "content": "<p>Can you give an example of the code you're talking about?<br>\nUsually, the 0/1 label comes from whether the purchase happened in the week being used for validation.</p>",
      "rawMarkdown": "Can you give an example of the code you're talking about?\nUsually, the 0/1 label comes from whether the purchase happened in the week being used for validation.",
      "votes": null
    },
    {
      "id": "1753067",
      "postDate": "04/12/2022 12:49:11",
      "content": "<p>Let me share some code snippet as below, from <a href=\"https://www.kaggle.com/code/masaponto/h-m-lightgbm-train-and-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/masaponto/h-m-lightgbm-train-and-feature-importance</a><br>\ndf_false = df_truth.copy() <br>\ndf_false.loc[:, \"article_id\"] = df_false[\"article_id\"].sample(frac=1).tolist()<br>\ndf_truth.loc[:, Config.label] = 1<br>\ndf_false.loc[:, Config.label] = 0</p>",
      "rawMarkdown": "Let me share some code snippet as below, from https://www.kaggle.com/code/masaponto/h-m-lightgbm-train-and-feature-importance\ndf_false = df_truth.copy() \ndf_false.loc[:, \"article_id\"] = df_false[\"article_id\"].sample(frac=1).tolist()\ndf_truth.loc[:, Config.label] = 1\ndf_false.loc[:, Config.label] = 0",
      "votes": null
    },
    {
      "id": "1753077",
      "postDate": "04/12/2022 13:08:19",
      "content": "<p>In that notebook, <code>df_truth</code> are real transactions that happened.<br>\n<code>df_truth = df_trans[[\"customer_id\", \"article_id\"]]</code></p>\n<p><code>df_false</code> starts off as a copy of <code>df_true</code>, but then the article_id column gets randomly shuffled. <br>\n<code>df[\"article_id\"].sample(frac=1)</code> randomly shuffles the article ids.</p>\n<p>Once shuffled, it can be assumed that most of the customer_id/article_id pairs are not real transactions that happened.</p>",
      "rawMarkdown": "In that notebook, `df_truth` are real transactions that happened.\n`df_truth = df_trans[[\"customer_id\", \"article_id\"]]`\n\n`df_false` starts off as a copy of `df_true`, but then the article_id column gets randomly shuffled. \n`df[\"article_id\"].sample(frac=1)` randomly shuffles the article ids.\n\nOnce shuffled, it can be assumed that most of the customer_id/article_id pairs are not real transactions that happened.",
      "votes": null
    },
    {
      "id": "1753142",
      "postDate": "04/12/2022 14:25:23",
      "content": "<p>I see. thank you for your help!<br>\nThis method might have augmentation effect, I guess.</p>",
      "rawMarkdown": "I see. thank you for your help!\nThis method might have augmentation effect, I guess.",
      "votes": null
    },
    {
      "id": "1757431",
      "postDate": "04/16/2022 15:54:10",
      "content": "<p>hei,guy! Dose the strategy of making negative samples make a good result?</p>",
      "rawMarkdown": "hei,guy! Dose the strategy of making negative samples make a good result?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1752629,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "04/12/2022 02:11:07",
      "content": "<p>Can you give an example of the code you're talking about?<br>\nUsually, the 0/1 label comes from whether the purchase happened in the week being used for validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1753067,
          "author_name": "lionsheep24",
          "author_url": "",
          "post_date": "04/12/2022 12:49:11",
          "content": "<p>Let me share some code snippet as below, from <a href=\"https://www.kaggle.com/code/masaponto/h-m-lightgbm-train-and-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/masaponto/h-m-lightgbm-train-and-feature-importance</a><br>\ndf_false = df_truth.copy() <br>\ndf_false.loc[:, \"article_id\"] = df_false[\"article_id\"].sample(frac=1).tolist()<br>\ndf_truth.loc[:, Config.label] = 1<br>\ndf_false.loc[:, Config.label] = 0</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1753077,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/12/2022 13:08:19",
          "content": "<p>In that notebook, <code>df_truth</code> are real transactions that happened.<br>\n<code>df_truth = df_trans[[\"customer_id\", \"article_id\"]]</code></p>\n<p><code>df_false</code> starts off as a copy of <code>df_true</code>, but then the article_id column gets randomly shuffled. <br>\n<code>df[\"article_id\"].sample(frac=1)</code> randomly shuffles the article ids.</p>\n<p>Once shuffled, it can be assumed that most of the customer_id/article_id pairs are not real transactions that happened.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1753142,
          "author_name": "lionsheep24",
          "author_url": "",
          "post_date": "04/12/2022 14:25:23",
          "content": "<p>I see. thank you for your help!<br>\nThis method might have augmentation effect, I guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1757431,
          "author_name": "dengniewei",
          "author_url": "",
          "post_date": "04/16/2022 15:54:10",
          "content": "<p>hei,guy! Dose the strategy of making negative samples make a good result?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1750714": "Hi all,\nSince Recys is totally new for me, I don't understand positive, negative sampling in many threads.\nRefer to many code examples, they just sample data and give label 0 or 1 to them, randomly.\nWhat does it mean to train a model this way?",
    "1752629": "Can you give an example of the code you're talking about?\nUsually, the 0/1 label comes from whether the purchase happened in the week being used for validation.",
    "1753067": "Let me share some code snippet as below, from https://www.kaggle.com/code/masaponto/h-m-lightgbm-train-and-feature-importance\ndf_false = df_truth.copy() \ndf_false.loc[:, \"article_id\"] = df_false[\"article_id\"].sample(frac=1).tolist()\ndf_truth.loc[:, Config.label] = 1\ndf_false.loc[:, Config.label] = 0",
    "1753077": "In that notebook, `df_truth` are real transactions that happened.\n`df_truth = df_trans[[\"customer_id\", \"article_id\"]]`\n\n`df_false` starts off as a copy of `df_true`, but then the article_id column gets randomly shuffled. \n`df[\"article_id\"].sample(frac=1)` randomly shuffles the article ids.\n\nOnce shuffled, it can be assumed that most of the customer_id/article_id pairs are not real transactions that happened.",
    "1753142": "I see. thank you for your help!\nThis method might have augmentation effect, I guess.",
    "1757431": "hei,guy! Dose the strategy of making negative samples make a good result?"
  },
  "source": "meta"
}