{
  "id": 382048,
  "title": "What is your best negative sampling rate ？ ",
  "url": "/competitions/otto-recommender-system/discussion/382048",
  "author_name": "",
  "post_date": "2023-01-29T11:20:16.951201700Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I used 1:20 right now and I just used 50 candidates for each type.  Has anyone used a higher sample rate?</p>",
  "messages": [
    {
      "id": "2120104",
      "postDate": "01/29/2023 11:20:16",
      "content": "<p>I used 1:20 right now and I just used 50 candidates for each type.  Has anyone used a higher sample rate?</p>",
      "rawMarkdown": "I used 1:20 right now and I just used 50 candidates for each type.  Has anyone used a higher sample rate?",
      "votes": null
    },
    {
      "id": "2120209",
      "postDate": "01/29/2023 12:32:58",
      "content": "<p>My suggestion would be to try out different ratios and see how it affects your local validation -  it depends on an array of hyper-parameters. </p>\n<p>20:1 comes with some issues as described in this thread <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442#2099887\" target=\"_blank\">here</a>.</p>\n<p>You might want to take session length into consideration after down-sampling. </p>",
      "rawMarkdown": "My suggestion would be to try out different ratios and see how it affects your local validation -  it depends on an array of hyper-parameters. \n\n20:1 comes with some issues as described in this thread [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442#2099887).\n\nYou might want to take session length into consideration after down-sampling.",
      "votes": null
    },
    {
      "id": "2120346",
      "postDate": "01/29/2023 14:40:13",
      "content": "<p>Yup, I just random drop a part of negative sample for train data. And I still never test another ratio. I plan to test it only tomorrow and ensemble model the day after tomorrow. </p>",
      "rawMarkdown": "Yup, I just random drop a part of negative sample for train data. And I still never test another ratio. I plan to test it only tomorrow and ensemble model the day after tomorrow.",
      "votes": null
    },
    {
      "id": "2120653",
      "postDate": "01/29/2023 18:27:33",
      "content": "<p>From my personal experience, I tried with 5 / 10 / 25 / 50 / 100 %, but the result didn't change at all,  I didn't notice any improvements in my recall.</p>\n<p>Also, I'm still facing the issue of 0.99% NDCG@20 despite using all those ratios, and still stuck in 0.565 recall</p>",
      "rawMarkdown": "From my personal experience, I tried with 5 / 10 / 25 / 50 / 100 %, but the result didn't change at all,  I didn't notice any improvements in my recall.\n\nAlso, I'm still facing the issue of 0.99% NDCG@20 despite using all those ratios, and still stuck in 0.565 recall",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2120209,
      "author_name": "parthpankajtiwary",
      "author_url": "",
      "post_date": "01/29/2023 12:32:58",
      "content": "<p>My suggestion would be to try out different ratios and see how it affects your local validation -  it depends on an array of hyper-parameters. </p>\n<p>20:1 comes with some issues as described in this thread <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442#2099887\" target=\"_blank\">here</a>.</p>\n<p>You might want to take session length into consideration after down-sampling. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2120346,
          "author_name": "earlee0412",
          "author_url": "",
          "post_date": "01/29/2023 14:40:13",
          "content": "<p>Yup, I just random drop a part of negative sample for train data. And I still never test another ratio. I plan to test it only tomorrow and ensemble model the day after tomorrow. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2120653,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/29/2023 18:27:33",
      "content": "<p>From my personal experience, I tried with 5 / 10 / 25 / 50 / 100 %, but the result didn't change at all,  I didn't notice any improvements in my recall.</p>\n<p>Also, I'm still facing the issue of 0.99% NDCG@20 despite using all those ratios, and still stuck in 0.565 recall</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2120104": "I used 1:20 right now and I just used 50 candidates for each type.  Has anyone used a higher sample rate?",
    "2120209": "My suggestion would be to try out different ratios and see how it affects your local validation -  it depends on an array of hyper-parameters. \n\n20:1 comes with some issues as described in this thread [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442#2099887).\n\nYou might want to take session length into consideration after down-sampling.",
    "2120346": "Yup, I just random drop a part of negative sample for train data. And I still never test another ratio. I plan to test it only tomorrow and ensemble model the day after tomorrow.",
    "2120653": "From my personal experience, I tried with 5 / 10 / 25 / 50 / 100 %, but the result didn't change at all,  I didn't notice any improvements in my recall.\n\nAlso, I'm still facing the issue of 0.99% NDCG@20 despite using all those ratios, and still stuck in 0.565 recall"
  },
  "source": "meta"
}