{
  "id": 55166,
  "title": "New Random Seed Increased LB Score by 0.0004. Can I Use it?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55166",
  "author_name": "",
  "post_date": "2018-04-23T04:01:40.594759300Z",
  "votes": 3,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Until today, I was using a random seed <strong>X</strong> and casually changed it to <strong>Y</strong> and observed some changes in the submission file along with an improved validation score. So, I went ahead and submitted the output to realize that the LB score improved by 0.0004.</p>\n\n<p>So, do you think is it OK to use the new random seed? BTW is it normal to see such a huge change (at least I think 0.0004 is a huge improvement in this competition.)</p>",
  "messages": [
    {
      "id": "318027",
      "postDate": "04/23/2018 04:01:40",
      "content": "<p>Until today, I was using a random seed <strong>X</strong> and casually changed it to <strong>Y</strong> and observed some changes in the submission file along with an improved validation score. So, I went ahead and submitted the output to realize that the LB score improved by 0.0004.</p>\n\n<p>So, do you think is it OK to use the new random seed? BTW is it normal to see such a huge change (at least I think 0.0004 is a huge improvement in this competition.)</p>",
      "rawMarkdown": "Until today, I was using a random seed **X** and casually changed it to **Y** and observed some changes in the submission file along with an improved validation score. So, I went ahead and submitted the output to realize that the LB score improved by 0.0004.\n\nSo, do you think is it OK to use the new random seed? BTW is it normal to see such a huge change (at least I think 0.0004 is a huge improvement in this competition.)",
      "votes": null
    },
    {
      "id": "318033",
      "postDate": "04/23/2018 04:29:49",
      "content": "<p>Keep calm and play with seed:) In the past competition (with small data size) people will ran different random seed and average the result from different random seed to get the final output. Which will make the output more reliable, but It seems like \"impossible\" in this competition (with such big bars size). I think new random seed are welcome, if it improve the score both in local CV and leaderboard. Good luck bro:)</p>",
      "rawMarkdown": "Keep calm and play with seed:) In the past competition (with small data size) people will ran different random seed and average the result from different random seed to get the final output. Which will make the output more reliable, but It seems like \"impossible\" in this competition (with such big bars size). I think new random seed are welcome, if it improve the score both in local CV and leaderboard. Good luck bro:)",
      "votes": null
    },
    {
      "id": "318061",
      "postDate": "04/23/2018 05:16:34",
      "content": "<p>Thanks a lot CHEN... Thats helpful...</p>",
      "rawMarkdown": "Thanks a lot CHEN... Thats helpful...",
      "votes": null
    },
    {
      "id": "318082",
      "postDate": "04/23/2018 06:57:16",
      "content": "<p>It shows that a 0.0004 change in LB is random noise...</p>\n\n<p>This is only  a half joke.  If your local CV score also increases then this may not be noise.</p>",
      "rawMarkdown": "It shows that a 0.0004 change in LB is random noise...\n\nThis is only  a half joke.  If your local CV score also increases then this may not be noise.",
      "votes": null
    },
    {
      "id": "318111",
      "postDate": "04/23/2018 07:47:32",
      "content": "<p>@CPMP Yeah I saw that the local validation score is also improved...</p>",
      "rawMarkdown": "CPMP Yeah I saw that the local validation score is also improved...",
      "votes": null
    },
    {
      "id": "318114",
      "postDate": "04/23/2018 07:49:33",
      "content": "<p>Then go for it. It might make the difference between gold and silver in few days :-).</p>",
      "rawMarkdown": "Then go for it. It might make the difference between gold and silver in few days :-).",
      "votes": null
    },
    {
      "id": "318129",
      "postDate": "04/23/2018 08:07:47",
      "content": "<p>Being my first competition, I'll be more than happy to receive a bronze and forget about Gold... that's way too much for me :)</p>",
      "rawMarkdown": "Being my first competition, I'll be more than happy to receive a bronze and forget about Gold... that's way too much for me :)",
      "votes": null
    },
    {
      "id": "318132",
      "postDate": "04/23/2018 08:10:36",
      "content": "<p>Actually I do think you will have chance to get a gold:-)</p>",
      "rawMarkdown": "Actually I do think you will have chance to get a gold:-)",
      "votes": null
    },
    {
      "id": "318138",
      "postDate": "04/23/2018 08:39:29",
      "content": "<p>You could always average over several random seeds to reduce the noise</p>",
      "rawMarkdown": "You could always average over several random seeds to reduce the noise",
      "votes": null
    },
    {
      "id": "318142",
      "postDate": "04/23/2018 08:51:33",
      "content": "<p>Sure.. will definitely try that.</p>",
      "rawMarkdown": "Sure.. will definitely try that.",
      "votes": null
    },
    {
      "id": "318175",
      "postDate": "04/23/2018 10:24:46",
      "content": "<p>I think it's noise.\nHow much data do you use to train?</p>\n\n<p>Depends on sample size, noise order would be:</p>\n\n<pre><code>10,000,000 sample -&gt;  noise 0.0003\n50,000,000 sample -&gt;  noise 0.0001 \n</code></pre>\n\n<p>This is not from statistics,  just from my experience about this competition data.</p>\n\n<p>In addition to averaging several random seeds, just more sample would help to reduce noise.</p>\n\n<p>good luck.</p>",
      "rawMarkdown": "I think it's noise.\nHow much data do you use to train?\n\nDepends on sample size, noise order would be:\n\n    10,000,000 sample -&gt;  noise 0.0003\n    50,000,000 sample -&gt;  noise 0.0001 \n\nThis is not from statistics,  just from my experience about this competition data.\n\nIn addition to averaging several random seeds, just more sample would help to reduce noise.\n\ngood luck.",
      "votes": null
    },
    {
      "id": "318180",
      "postDate": "04/23/2018 10:31:02",
      "content": "<p>Is that mean with the bigger training sample, the random seeds will have less impact on the final result?</p>",
      "rawMarkdown": "Is that mean with the bigger training sample, the random seeds will have less impact on the final result?",
      "votes": null
    },
    {
      "id": "318183",
      "postDate": "04/23/2018 10:51:06",
      "content": "<p>@CuteChibiko I'm using all the data except the last few....</p>",
      "rawMarkdown": "CuteChibiko I'm using all the data except the last few....",
      "votes": null
    },
    {
      "id": "318215",
      "postDate": "04/23/2018 12:17:05",
      "content": "<p>@CHEN, @Samrat,</p>\n\n<p>Sorry, the noise which I observed was from both train sample size and validation sample size.\nSo, this is not the case for Samrat LB score improved, and I could not understand the reason of this big improvement of 0.0004. :(\nAny way congrats for big boost! :)</p>",
      "rawMarkdown": "CHEN, @Samrat,\n\nSorry, the noise which I observed was from both train sample size and validation sample size.\nSo, this is not the case for Samrat LB score improved, and I could not understand the reason of this big improvement of 0.0004. :(\nAny way congrats for big boost! :)",
      "votes": null
    },
    {
      "id": "318232",
      "postDate": "04/23/2018 12:59:17",
      "content": "<p>@CuteChibiko Nothing to be sorry... Even I have no clue on this improvement. I was struggling for days to get an improvement of 0.0002 and 0.0004 improvement with a small change is a much welcomed luck!!</p>",
      "rawMarkdown": "CuteChibiko Nothing to be sorry... Even I have no clue on this improvement. I was struggling for days to get an improvement of 0.0002 and 0.0004 improvement with a small change is a much welcomed luck!!",
      "votes": null
    },
    {
      "id": "318238",
      "postDate": "04/23/2018 13:07:34",
      "content": "<p>I bet everybody is now frantically trying different random seeds :)</p>",
      "rawMarkdown": "I bet everybody is now frantically trying different random seeds :)",
      "votes": null
    },
    {
      "id": "318244",
      "postDate": "04/23/2018 13:15:30",
      "content": "<p>@CPMP Even I'm trying with my/spouse/kid date of birth/ month of birth/ year of birth ;)</p>",
      "rawMarkdown": "CPMP Even I'm trying with my/spouse/kid date of birth/ month of birth/ year of birth ;)",
      "votes": null
    },
    {
      "id": "318247",
      "postDate": "04/23/2018 13:20:10",
      "content": "<p>@CuteChibiko Don't feel sorry about it. Your guys are pretty awesome in this competition. 😉</p>",
      "rawMarkdown": "CuteChibiko Don't feel sorry about it. Your guys are pretty awesome in this competition. 😉",
      "votes": null
    },
    {
      "id": "318266",
      "postDate": "04/23/2018 14:06:34",
      "content": "<p>what is your birthday date then ? :) haha</p>",
      "rawMarkdown": "what is your birthday date then ? :) haha",
      "votes": null
    },
    {
      "id": "320531",
      "postDate": "04/29/2018 00:39:43",
      "content": "<p>I learn something new from a fellow Kaggler everyday.  I will try that on my own kernel.</p>",
      "rawMarkdown": "I learn something new from a fellow Kaggler everyday.  I will try that on my own kernel.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 318033,
      "author_name": "a1339398734",
      "author_url": "",
      "post_date": "04/23/2018 04:29:49",
      "content": "<p>Keep calm and play with seed:) In the past competition (with small data size) people will ran different random seed and average the result from different random seed to get the final output. Which will make the output more reliable, but It seems like \"impossible\" in this competition (with such big bars size). I think new random seed are welcome, if it improve the score both in local CV and leaderboard. Good luck bro:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 318061,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 05:16:34",
          "content": "<p>Thanks a lot CHEN... Thats helpful...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318082,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/23/2018 06:57:16",
      "content": "<p>It shows that a 0.0004 change in LB is random noise...</p>\n\n<p>This is only  a half joke.  If your local CV score also increases then this may not be noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 318111,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 07:47:32",
          "content": "<p>@CPMP Yeah I saw that the local validation score is also improved...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318114,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "04/23/2018 07:49:33",
          "content": "<p>Then go for it. It might make the difference between gold and silver in few days :-).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318129,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 08:07:47",
          "content": "<p>Being my first competition, I'll be more than happy to receive a bronze and forget about Gold... that's way too much for me :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318132,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 08:10:36",
          "content": "<p>Actually I do think you will have chance to get a gold:-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318138,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "04/23/2018 08:39:29",
      "content": "<p>You could always average over several random seeds to reduce the noise</p>",
      "votes": null,
      "replies": [
        {
          "id": 318142,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 08:51:33",
          "content": "<p>Sure.. will definitely try that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320531,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "04/29/2018 00:39:43",
          "content": "<p>I learn something new from a fellow Kaggler everyday.  I will try that on my own kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318175,
      "author_name": "its7171",
      "author_url": "",
      "post_date": "04/23/2018 10:24:46",
      "content": "<p>I think it's noise.\nHow much data do you use to train?</p>\n\n<p>Depends on sample size, noise order would be:</p>\n\n<pre><code>10,000,000 sample -&gt;  noise 0.0003\n50,000,000 sample -&gt;  noise 0.0001 \n</code></pre>\n\n<p>This is not from statistics,  just from my experience about this competition data.</p>\n\n<p>In addition to averaging several random seeds, just more sample would help to reduce noise.</p>\n\n<p>good luck.</p>",
      "votes": null,
      "replies": [
        {
          "id": 318180,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 10:31:02",
          "content": "<p>Is that mean with the bigger training sample, the random seeds will have less impact on the final result?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318183,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 10:51:06",
          "content": "<p>@CuteChibiko I'm using all the data except the last few....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318215,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/23/2018 12:17:05",
          "content": "<p>@CHEN, @Samrat,</p>\n\n<p>Sorry, the noise which I observed was from both train sample size and validation sample size.\nSo, this is not the case for Samrat LB score improved, and I could not understand the reason of this big improvement of 0.0004. :(\nAny way congrats for big boost! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318232,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 12:59:17",
          "content": "<p>@CuteChibiko Nothing to be sorry... Even I have no clue on this improvement. I was struggling for days to get an improvement of 0.0002 and 0.0004 improvement with a small change is a much welcomed luck!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318238,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/23/2018 13:07:34",
          "content": "<p>I bet everybody is now frantically trying different random seeds :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318244,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/23/2018 13:15:30",
          "content": "<p>@CPMP Even I'm trying with my/spouse/kid date of birth/ month of birth/ year of birth ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318247,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 13:20:10",
          "content": "<p>@CuteChibiko Don't feel sorry about it. Your guys are pretty awesome in this competition. 😉</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318266,
          "author_name": "kesjien",
          "author_url": "",
          "post_date": "04/23/2018 14:06:34",
          "content": "<p>what is your birthday date then ? :) haha</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "318027": "Until today, I was using a random seed **X** and casually changed it to **Y** and observed some changes in the submission file along with an improved validation score. So, I went ahead and submitted the output to realize that the LB score improved by 0.0004.\n\nSo, do you think is it OK to use the new random seed? BTW is it normal to see such a huge change (at least I think 0.0004 is a huge improvement in this competition.)",
    "318033": "Keep calm and play with seed:) In the past competition (with small data size) people will ran different random seed and average the result from different random seed to get the final output. Which will make the output more reliable, but It seems like \"impossible\" in this competition (with such big bars size). I think new random seed are welcome, if it improve the score both in local CV and leaderboard. Good luck bro:)",
    "318061": "Thanks a lot CHEN... Thats helpful...",
    "318082": "It shows that a 0.0004 change in LB is random noise...\n\nThis is only  a half joke.  If your local CV score also increases then this may not be noise.",
    "318111": "CPMP Yeah I saw that the local validation score is also improved...",
    "318114": "Then go for it. It might make the difference between gold and silver in few days :-).",
    "318129": "Being my first competition, I'll be more than happy to receive a bronze and forget about Gold... that's way too much for me :)",
    "318132": "Actually I do think you will have chance to get a gold:-)",
    "318138": "You could always average over several random seeds to reduce the noise",
    "318142": "Sure.. will definitely try that.",
    "318175": "I think it's noise.\nHow much data do you use to train?\n\nDepends on sample size, noise order would be:\n\n    10,000,000 sample -&gt;  noise 0.0003\n    50,000,000 sample -&gt;  noise 0.0001 \n\nThis is not from statistics,  just from my experience about this competition data.\n\nIn addition to averaging several random seeds, just more sample would help to reduce noise.\n\ngood luck.",
    "318180": "Is that mean with the bigger training sample, the random seeds will have less impact on the final result?",
    "318183": "CuteChibiko I'm using all the data except the last few....",
    "318215": "CHEN, @Samrat,\n\nSorry, the noise which I observed was from both train sample size and validation sample size.\nSo, this is not the case for Samrat LB score improved, and I could not understand the reason of this big improvement of 0.0004. :(\nAny way congrats for big boost! :)",
    "318232": "CuteChibiko Nothing to be sorry... Even I have no clue on this improvement. I was struggling for days to get an improvement of 0.0002 and 0.0004 improvement with a small change is a much welcomed luck!!",
    "318238": "I bet everybody is now frantically trying different random seeds :)",
    "318244": "CPMP Even I'm trying with my/spouse/kid date of birth/ month of birth/ year of birth ;)",
    "318247": "CuteChibiko Don't feel sorry about it. Your guys are pretty awesome in this competition. 😉",
    "318266": "what is your birthday date then ? :) haha",
    "320531": "I learn something new from a fellow Kaggler everyday.  I will try that on my own kernel."
  },
  "source": "meta"
}