{
  "id": 329787,
  "title": "Can you find the best seed?",
  "url": "/competitions/amex-default-prediction/discussion/329787",
  "author_name": "AmbrosM",
  "post_date": "2022-06-08T18:17:23.784000",
  "votes": 127,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Have you ever thought about training your model several times with different seeds and then selecting the seed with the best score for submission? Wouldn't it be nice to increase the lb score so easily? I wanted to know whether this idea works.</p>\n<p>For the experiment, I split the training data into three parts: A training set (80 %) and two validation sets (10 % each, i.e. 45891 customers). (This is the same split as in the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782\" target=\"_blank\">discussion post about the noisy metric</a>.) Then I trained 600 instances of my gradient-boosting model with 600 different seeds on this training set. For all 600 instances, I measured the score on both validation sets. If it were possible to select the best seed based on the cv score (or based on the public lb score) and so get a higher private lb score, this process could be simulated with the two validation sets. For the method to work, a good score on validation set A should predict a good score on validation set B. In other words, the scores on validation set A and B would need to be correlated.</p>\n<p>A 2d scatterplot of the validation scores immediately destroys every hope: The two scores are fully independent. Whether you select a model with a low score on dataset A (left part of the diagram) or one with a high score on dataset A (right part of the diagram), the expected outcome on dataset B is the same.</p>\n<p><img src=\"https://i.imgur.com/gNW5Fxt.png\" alt=\"correlation-scatter\"></p>\n<p>For a formal statistical test, we can use the function <code>scipy.stats.pearsonr()</code> to compute the Pearson correlation coefficient \\(r = -0.02\\) and the two-sided p-value, which is \\(p = 0.65\\). A p-value less than \\(0.05\\) might be considered a significant correlation; \\(0.65\\) is far above that.</p>\n<p>Summary: Don't try to find the best seed - it's impossible!</p>",
  "messages": [
    {
      "id": 1815153,
      "postDate": "2022-06-08T18:17:23.783Z",
      "content": "<p>Have you ever thought about training your model several times with different seeds and then selecting the seed with the best score for submission? Wouldn't it be nice to increase the lb score so easily? I wanted to know whether this idea works.</p>\n<p>For the experiment, I split the training data into three parts: A training set (80 %) and two validation sets (10 % each, i.e. 45891 customers). (This is the same split as in the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782\" target=\"_blank\">discussion post about the noisy metric</a>.) Then I trained 600 instances of my gradient-boosting model with 600 different seeds on this training set. For all 600 instances, I measured the score on both validation sets. If it were possible to select the best seed based on the cv score (or based on the public lb score) and so get a higher private lb score, this process could be simulated with the two validation sets. For the method to work, a good score on validation set A should predict a good score on validation set B. In other words, the scores on validation set A and B would need to be correlated.</p>\n<p>A 2d scatterplot of the validation scores immediately destroys every hope: The two scores are fully independent. Whether you select a model with a low score on dataset A (left part of the diagram) or one with a high score on dataset A (right part of the diagram), the expected outcome on dataset B is the same.</p>\n<p><img src=\"https://i.imgur.com/gNW5Fxt.png\" alt=\"correlation-scatter\"></p>\n<p>For a formal statistical test, we can use the function <code>scipy.stats.pearsonr()</code> to compute the Pearson correlation coefficient \\(r = -0.02\\) and the two-sided p-value, which is \\(p = 0.65\\). A p-value less than \\(0.05\\) might be considered a significant correlation; \\(0.65\\) is far above that.</p>\n<p>Summary: Don't try to find the best seed - it's impossible!</p>",
      "rawMarkdown": "Have you ever thought about training your model several times with different seeds and then selecting the seed with the best score for submission? Wouldn't it be nice to increase the lb score so easily? I wanted to know whether this idea works.\n\nFor the experiment, I split the training data into three parts: A training set (80 %) and two validation sets (10 % each, i.e. 45891 customers). (This is the same split as in the [discussion post about the noisy metric](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782).) Then I trained 600 instances of my gradient-boosting model with 600 different seeds on this training set. For all 600 instances, I measured the score on both validation sets. If it were possible to select the best seed based on the cv score (or based on the public lb score) and so get a higher private lb score, this process could be simulated with the two validation sets. For the method to work, a good score on validation set A should predict a good score on validation set B. In other words, the scores on validation set A and B would need to be correlated.\n\nA 2d scatterplot of the validation scores immediately destroys every hope: The two scores are fully independent. Whether you select a model with a low score on dataset A (left part of the diagram) or one with a high score on dataset A (right part of the diagram), the expected outcome on dataset B is the same.\n\n![correlation-scatter](https://i.imgur.com/gNW5Fxt.png)\n\nFor a formal statistical test, we can use the function `scipy.stats.pearsonr()` to compute the Pearson correlation coefficient \\\\(r = -0.02\\\\) and the two-sided p-value, which is \\\\(p = 0.65\\\\). A p-value less than \\\\(0.05\\\\) might be considered a significant correlation; \\\\(0.65\\\\) is far above that.\n\nSummary: Don't try to find the best seed - it's impossible!",
      "votes": 126
    },
    {
      "id": 1815215,
      "postDate": "2022-06-08T19:47:03.720Z",
      "content": "<p>We all know the best Seed -&gt; 42</p>",
      "rawMarkdown": "We all know the best Seed -> 42",
      "votes": 37,
      "replies": [
        {
          "id": 1854720,
          "postDate": "2022-07-14T00:27:56.537Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1815316,
      "postDate": "2022-06-08T23:06:04.127Z",
      "content": "<p>Seeing this for gradient boosting makes sense to me, but I wonder if it might be a different story for neural network approaches. Randomness plays a pretty small role in the typical booster -- maybe some column and row subsampling -- but can play a bigger role in neural nets through parameter initialization and the downstream impact on model convergence in a wonky, nonconvex cost space. That's not to say that seed hacking is a good plan for DL, but there's a good reason why e.g. averaging models over multiple random weight initializations is a common ensembling technique.</p>",
      "rawMarkdown": "Seeing this for gradient boosting makes sense to me, but I wonder if it might be a different story for neural network approaches. Randomness plays a pretty small role in the typical booster -- maybe some column and row subsampling -- but can play a bigger role in neural nets through parameter initialization and the downstream impact on model convergence in a wonky, nonconvex cost space. That's not to say that seed hacking is a good plan for DL, but there's a good reason why e.g. averaging models over multiple random weight initializations is a common ensembling technique.",
      "votes": 14,
      "replies": [
        {
          "id": 1815700,
          "postDate": "2022-06-09T10:27:09.627Z",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> Right, we have to be careful not to overgeneralize the result of the experiment.</p>",
          "rawMarkdown": "@aquatic Right, we have to be careful not to overgeneralize the result of the experiment.",
          "votes": 3
        },
        {
          "id": 1815850,
          "postDate": "2022-06-09T14:43:03.223Z",
          "content": "<p>In the past using different seeds and stacking the results makes a slight improvement, which can give you a few places even in the silver metal range.</p>",
          "rawMarkdown": "In the past using different seeds and stacking the results makes a slight improvement, which can give you a few places even in the silver metal range.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1833492,
      "postDate": "2022-06-26T04:16:05.937Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <br>\nA similar discussion happened in an earlier competition- Petfinder where <br>\n<a href=\"https://www.kaggle.com/haoge233\" target=\"_blank\">@haoge233</a> discussed how different seeds affected his score <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/285580\" target=\"_blank\">here</a><br>\n<a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> had raised a similar point, <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/296706\" target=\"_blank\">here</a> and referenced the  paper - <a href=\"https://arxiv.org/pdf/2109.08203.pdf\" target=\"_blank\">https://arxiv.org/pdf/2109.08203.pdf</a>   which suggets </p>\n<blockquote>\n  <p>there are indeed seeds that produce scores sufficiently good to be considered as a significant improvement by the computer vision community. This is a worrying result as the community is currently very much score driven, and yet these can just be artifacts of randomness.</p>\n</blockquote>\n<p>However, this was not useful in the end. As many pointed out and saw post the private LB release and final results.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> noted <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/280302#1551958\" target=\"_blank\">here</a>:</p>\n<blockquote>\n  <p>If using a different seed changes your CV score, that means that the train data is small and/or your model has large variance in accuracy after training.<br>\n  When this happens, you can either try changing your model (lower its variance), or you may have to run many different seeds (or same seed multiple times) and average all the CVs together to get a reliable CV score. </p>\n</blockquote>\n<p>As a result, best is to stick your preferred seed and try to lower variance of the model.<br>\nHope this helps!</p>",
      "rawMarkdown": "Dear @ambrosm \nA similar discussion happened in an earlier competition- Petfinder where \n@haoge233 discussed how different seeds affected his score [here](https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/285580)\n@allohvk had raised a similar point, [here](https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/296706) and referenced the  paper - https://arxiv.org/pdf/2109.08203.pdf   which suggets \n>there are indeed seeds that produce scores sufficiently good to be considered as a significant improvement by the computer vision community. This is a worrying result as the community is currently very much score driven, and yet these can just be artifacts of randomness.\n\nHowever, this was not useful in the end. As many pointed out and saw post the private LB release and final results.\n\n@cdeotte noted [here](https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/280302#1551958):\n>If using a different seed changes your CV score, that means that the train data is small and/or your model has large variance in accuracy after training.\n> When this happens, you can either try changing your model (lower its variance), or you may have to run many different seeds (or same seed multiple times) and average all the CVs together to get a reliable CV score. \n\nAs a result, best is to stick your preferred seed and try to lower variance of the model.\nHope this helps!\n",
      "votes": 11
    },
    {
      "id": 1815183,
      "postDate": "2022-06-08T19:04:18.720Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p>You seem to have built up a sort of  \"ergodic hypothesis\" for this competition 😄!<br>\nIn the \"niosy\" post you keep the seed constant and observe the fluctuations over time (<em>i.e.</em> number of iterations), and in this experiment you keep the time constant and create an ensemble of seeds. Both approaches indicating similar expectation values. Wonderful work!</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @ambrosm \n\nYou seem to have built up a sort of  \"ergodic hypothesis\" for this competition 😄!\nIn the \"niosy\" post you keep the seed constant and observe the fluctuations over time (*i.e.* number of iterations), and in this experiment you keep the time constant and create an ensemble of seeds. Both approaches indicating similar expectation values. Wonderful work!\n\nAll the best,\ncarl",
      "votes": 9
    },
    {
      "id": 1815670,
      "postDate": "2022-06-09T09:37:20.820Z",
      "content": "<p>For fun =)))<br>\n<a href=\"https://arxiv.org/pdf/2109.08203.pdf\" target=\"_blank\">https://arxiv.org/pdf/2109.08203.pdf</a></p>",
      "rawMarkdown": "For fun =)))\nhttps://arxiv.org/pdf/2109.08203.pdf",
      "votes": 8
    },
    {
      "id": 1815571,
      "postDate": "2022-06-09T07:07:39.367Z",
      "content": "<blockquote>\n  <p><strong>random_state</strong> should only be used for reproducibility and should not give a better model.</p>\n</blockquote>\n<p>I had this notion in my mind, but today I saw it's experimental proof. thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "> **random_state** should only be used for reproducibility and should not give a better model.\n\nI had this notion in my mind, but today I saw it's experimental proof. thanks @ambrosm ",
      "votes": 6,
      "replies": [
        {
          "id": 1941236,
          "postDate": "2022-09-15T21:26:49.680Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1816682,
      "postDate": "2022-06-10T12:45:29.270Z",
      "content": "<p>İ don't think that Seed is one of the important parameters that should be optimized in ML model. </p>",
      "rawMarkdown": "İ don't think that Seed is one of the important parameters that should be optimized in ML model. ",
      "votes": 4,
      "replies": [
        {
          "id": 1819228,
          "postDate": "2022-06-13T14:53:38.170Z",
          "content": "<p>The seed actually wouldn't be an important parameters for a ML model. But you will get gold in leaderborard if you improve your LB score 0.001. Sometimes, seed worked.</p>",
          "rawMarkdown": "The seed actually wouldn't be an important parameters for a ML model. But you will get gold in leaderborard if you improve your LB score 0.001. Sometimes, seed worked.",
          "votes": 1
        },
        {
          "id": 1823493,
          "postDate": "2022-06-17T12:32:38.190Z",
          "content": "<p>That is true in theory, but in practice changing the seed sometimes can yield slightly better results. That might be related to the nondeterministic nature of some machine learning algorithms!<br>\nBut we all know the best seed is 42 😄. </p>",
          "rawMarkdown": "That is true in theory, but in practice changing the seed sometimes can yield slightly better results. That might be related to the nondeterministic nature of some machine learning algorithms!\nBut we all know the best seed is 42 😄. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1838921,
      "postDate": "2022-07-01T02:15:29.723Z",
      "content": "<p>I always wanted to try this. Thanks for sharing your work!</p>",
      "rawMarkdown": "I always wanted to try this. Thanks for sharing your work!",
      "votes": 1
    },
    {
      "id": 1825781,
      "postDate": "2022-06-19T17:46:53.863Z",
      "content": "<p>That's Interesting.</p>",
      "rawMarkdown": "That's Interesting.",
      "votes": 1
    },
    {
      "id": 1822492,
      "postDate": "2022-06-16T12:45:19.867Z",
      "content": "<p>42 indeed is the best seed :P</p>",
      "rawMarkdown": "42 indeed is the best seed :P",
      "votes": 1
    },
    {
      "id": 1819443,
      "postDate": "2022-06-13T19:11:07.497Z",
      "content": "<p>I just choose a random seed. But good to know the seed can sometimes improve score a little bit.</p>",
      "rawMarkdown": "I just choose a random seed. But good to know the seed can sometimes improve score a little bit.",
      "votes": 1
    },
    {
      "id": 1818620,
      "postDate": "2022-06-13T01:26:20.497Z",
      "content": "<p>why 42 is best seed?</p>",
      "rawMarkdown": "why 42 is best seed?",
      "votes": 1,
      "replies": [
        {
          "id": 1818773,
          "postDate": "2022-06-13T06:30:18.390Z",
          "content": "<p>The number 42 is especially significant to fans of science fiction novelist Douglas Adams’ “The Hitchhiker’s Guide to the Galaxy,” because that number is the answer given by a supercomputer to “the Ultimate Question of Life, the Universe, and Everything.” <br>\n<a href=\"https://news.mit.edu/2019/answer-life-universe-and-everything-sum-three-cubes-mathematics-0910\" target=\"_blank\">https://news.mit.edu/2019/answer-life-universe-and-everything-sum-three-cubes-mathematics-0910</a></p>",
          "rawMarkdown": "The number 42 is especially significant to fans of science fiction novelist Douglas Adams’ “The Hitchhiker’s Guide to the Galaxy,” because that number is the answer given by a supercomputer to “the Ultimate Question of Life, the Universe, and Everything.” \n[https://news.mit.edu/2019/answer-life-universe-and-everything-sum-three-cubes-mathematics-0910](https://news.mit.edu/2019/answer-life-universe-and-everything-sum-three-cubes-mathematics-0910)",
          "votes": 12
        },
        {
          "id": 1819222,
          "postDate": "2022-06-13T14:46:20.287Z",
          "content": "<p>Wow! That is an amazing answer. Maybe we are training the things not are models but life.😂</p>",
          "rawMarkdown": "Wow! That is an amazing answer. Maybe we are training the things not are models but life.😂"
        }
      ]
    },
    {
      "id": 1903417,
      "postDate": "2022-08-17T12:22:12.333Z",
      "content": "<p>Thanks for sharing this; nice to confirm that randomness doesn't have a huge impact on the overall results</p>",
      "rawMarkdown": "Thanks for sharing this; nice to confirm that randomness doesn't have a huge impact on the overall results",
      "votes": 2
    },
    {
      "id": 1868705,
      "postDate": "2022-07-24T06:59:43.553Z",
      "content": "<p>This is a great effort, thanks for sharing the results! </p>",
      "rawMarkdown": "This is a great effort, thanks for sharing the results! "
    },
    {
      "id": 1823804,
      "postDate": "2022-06-17T16:54:08.867Z",
      "content": "<p>Thank you for putting in all of this effort</p>",
      "rawMarkdown": "Thank you for putting in all of this effort\n",
      "votes": 2
    },
    {
      "id": 1821369,
      "postDate": "2022-06-15T13:28:20.400Z",
      "content": "<p>How about using 600 seeds for validation set split and looking for correlation :S</p>\n<p>On more serious notice do you think dropping features with adversarial validation analysis can help with this? I still believe that may be the reason</p>",
      "rawMarkdown": "How about using 600 seeds for validation set split and looking for correlation :S\n\nOn more serious notice do you think dropping features with adversarial validation analysis can help with this? I still believe that may be the reason",
      "votes": 2,
      "replies": [
        {
          "id": 1825415,
          "postDate": "2022-06-19T10:16:01.640Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/huseyincot\" target=\"_blank\">@huseyincot</a> Adversarial validation analysis is certainly a good thing to do - I haven't yet started it.</p>",
          "rawMarkdown": "Hi @huseyincot Adversarial validation analysis is certainly a good thing to do - I haven't yet started it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1818714,
      "postDate": "2022-06-13T05:01:36.913Z",
      "content": "<p>42 is all u need</p>",
      "rawMarkdown": "42 is all u need",
      "votes": 2
    },
    {
      "id": 1817729,
      "postDate": "2022-06-11T17:43:55.017Z",
      "content": "<p>Someone had to do this test! Thanks for the heads up…</p>",
      "rawMarkdown": "Someone had to do this test! Thanks for the heads up...",
      "votes": 2
    },
    {
      "id": 1817703,
      "postDate": "2022-06-11T17:03:46.180Z",
      "content": "<p>What it seems from the Plot is that, for a specific split, the val_score across seeds kinda resemble a Gaussian distribution (close enough). In this case, for A the mean_val_score seems to be somewhere between ~0.792-0.793 and likewise for B. So why not just averaging results across seeds instead of searching the best seed? Just a pretty standard way to make the submissions robust to noise, imo. (As far as ‘seed’ is concerned.)</p>",
      "rawMarkdown": "What it seems from the Plot is that, for a specific split, the val_score across seeds kinda resemble a Gaussian distribution (close enough). In this case, for A the mean_val_score seems to be somewhere between ~0.792-0.793 and likewise for B. So why not just averaging results across seeds instead of searching the best seed? Just a pretty standard way to make the submissions robust to noise, imo. (As far as ‘seed’ is concerned.)",
      "votes": 2,
      "replies": [
        {
          "id": 1818539,
          "postDate": "2022-06-12T20:07:32.700Z",
          "content": "<p><a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> Sure, ensembling across seeds is always a good method. With the experiment I wanted to investigate whether it makes sense to select good seeds before ensembling.</p>",
          "rawMarkdown": "@mrutyunjaybiswal Sure, ensembling across seeds is always a good method. With the experiment I wanted to investigate whether it makes sense to select good seeds before ensembling.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1817444,
      "postDate": "2022-06-11T10:57:34.530Z",
      "content": "<p>I think this is the best seed -&gt; 42</p>",
      "rawMarkdown": "I think this is the best seed -> 42",
      "votes": 2
    },
    {
      "id": 1816973,
      "postDate": "2022-06-10T18:05:52.903Z",
      "content": "<p>Another way to think about this is you will need to win by more than 0.005 on the public leaderboard to be confident it's not just noise.</p>",
      "rawMarkdown": "Another way to think about this is you will need to win by more than 0.005 on the public leaderboard to be confident it's not just noise.",
      "votes": 2
    },
    {
      "id": 1816718,
      "postDate": "2022-06-10T13:48:35.830Z",
      "content": "<p>Thanks for sharing an insightful experiment</p>",
      "rawMarkdown": "Thanks for sharing an insightful experiment",
      "votes": 2
    },
    {
      "id": 1815642,
      "postDate": "2022-06-09T09:00:21.173Z",
      "content": "<p>Fascinating insights. It is good to see that randomness doesn't greatly impact the overall results. </p>",
      "rawMarkdown": "Fascinating insights. It is good to see that randomness doesn't greatly impact the overall results. ",
      "votes": 2
    },
    {
      "id": 1815345,
      "postDate": "2022-06-09T00:04:18.583Z",
      "content": "<p>Impressive posting! Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> :)</p>",
      "rawMarkdown": "Impressive posting! Thank you @ambrosm :)",
      "votes": 2,
      "replies": [
        {
          "id": 1815824,
          "postDate": "2022-06-09T13:51:55.147Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1815273,
      "postDate": "2022-06-08T21:53:48.190Z",
      "content": "<p>Thank you for an interesting experiment.</p>",
      "rawMarkdown": "Thank you for an interesting experiment.",
      "votes": 2
    },
    {
      "id": 1815823,
      "postDate": "2022-06-09T13:51:20.357Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1815215,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-06-08T19:47:03.720000",
      "content": "<p>We all know the best Seed -&gt; 42</p>",
      "votes": 37,
      "replies": [
        {
          "id": 1854720,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-14T00:27:56.537000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1815316,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2022-06-08T23:06:04.127000",
      "content": "<p>Seeing this for gradient boosting makes sense to me, but I wonder if it might be a different story for neural network approaches. Randomness plays a pretty small role in the typical booster -- maybe some column and row subsampling -- but can play a bigger role in neural nets through parameter initialization and the downstream impact on model convergence in a wonky, nonconvex cost space. That's not to say that seed hacking is a good plan for DL, but there's a good reason why e.g. averaging models over multiple random weight initializations is a common ensembling technique.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 1815700,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-06-09T10:27:09.627000",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> Right, we have to be careful not to overgeneralize the result of the experiment.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1815850,
          "author_name": "happycube",
          "author_url": "",
          "post_date": "2022-06-09T14:43:03.223000",
          "content": "<p>In the past using different seeds and stacking the results makes a slight improvement, which can give you a few places even in the silver metal range.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1833492,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2022-06-26T04:16:05.937000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <br>\nA similar discussion happened in an earlier competition- Petfinder where <br>\n<a href=\"https://www.kaggle.com/haoge233\" target=\"_blank\">@haoge233</a> discussed how different seeds affected his score <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/285580\" target=\"_blank\">here</a><br>\n<a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> had raised a similar point, <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/296706\" target=\"_blank\">here</a> and referenced the  paper - <a href=\"https://arxiv.org/pdf/2109.08203.pdf\" target=\"_blank\">https://arxiv.org/pdf/2109.08203.pdf</a>   which suggets </p>\n<blockquote>\n  <p>there are indeed seeds that produce scores sufficiently good to be considered as a significant improvement by the computer vision community. This is a worrying result as the community is currently very much score driven, and yet these can just be artifacts of randomness.</p>\n</blockquote>\n<p>However, this was not useful in the end. As many pointed out and saw post the private LB release and final results.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> noted <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/280302#1551958\" target=\"_blank\">here</a>:</p>\n<blockquote>\n  <p>If using a different seed changes your CV score, that means that the train data is small and/or your model has large variance in accuracy after training.<br>\n  When this happens, you can either try changing your model (lower its variance), or you may have to run many different seeds (or same seed multiple times) and average all the CVs together to get a reliable CV score. </p>\n</blockquote>\n<p>As a result, best is to stick your preferred seed and try to lower variance of the model.<br>\nHope this helps!</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 1815183,
      "author_name": "Carl McBride Ellis",
      "author_url": "",
      "post_date": "2022-06-08T19:04:18.720000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p>You seem to have built up a sort of  \"ergodic hypothesis\" for this competition 😄!<br>\nIn the \"niosy\" post you keep the seed constant and observe the fluctuations over time (<em>i.e.</em> number of iterations), and in this experiment you keep the time constant and create an ensemble of seeds. Both approaches indicating similar expectation values. Wonderful work!</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1815670,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2022-06-09T09:37:20.820000",
      "content": "<p>For fun =)))<br>\n<a href=\"https://arxiv.org/pdf/2109.08203.pdf\" target=\"_blank\">https://arxiv.org/pdf/2109.08203.pdf</a></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1815571,
      "author_name": "AKR",
      "author_url": "",
      "post_date": "2022-06-09T07:07:39.367000",
      "content": "<blockquote>\n  <p><strong>random_state</strong> should only be used for reproducibility and should not give a better model.</p>\n</blockquote>\n<p>I had this notion in my mind, but today I saw it's experimental proof. thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 6,
      "replies": [
        {
          "id": 1941236,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-15T21:26:49.680000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1816682,
      "author_name": "MuhammedSAL98",
      "author_url": "",
      "post_date": "2022-06-10T12:45:29.270000",
      "content": "<p>İ don't think that Seed is one of the important parameters that should be optimized in ML model. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1819228,
          "author_name": "SgangX",
          "author_url": "",
          "post_date": "2022-06-13T14:53:38.170000",
          "content": "<p>The seed actually wouldn't be an important parameters for a ML model. But you will get gold in leaderborard if you improve your LB score 0.001. Sometimes, seed worked.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1823493,
          "author_name": "OmJ",
          "author_url": "",
          "post_date": "2022-06-17T12:32:38.190000",
          "content": "<p>That is true in theory, but in practice changing the seed sometimes can yield slightly better results. That might be related to the nondeterministic nature of some machine learning algorithms!<br>\nBut we all know the best seed is 42 😄. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1838921,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-01T02:15:29.723000",
      "content": "<p>I always wanted to try this. Thanks for sharing your work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1825781,
      "author_name": "Abhinav Kumar",
      "author_url": "",
      "post_date": "2022-06-19T17:46:53.863000",
      "content": "<p>That's Interesting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1822492,
      "author_name": "Nikos Sakellariou",
      "author_url": "",
      "post_date": "2022-06-16T12:45:19.867000",
      "content": "<p>42 indeed is the best seed :P</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1819443,
      "author_name": "Carla Laia",
      "author_url": "",
      "post_date": "2022-06-13T19:11:07.497000",
      "content": "<p>I just choose a random seed. But good to know the seed can sometimes improve score a little bit.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818620,
      "author_name": "SgangX",
      "author_url": "",
      "post_date": "2022-06-13T01:26:20.497000",
      "content": "<p>why 42 is best seed?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1818773,
          "author_name": "林湧森 (Dyson Lin)",
          "author_url": "",
          "post_date": "2022-06-13T06:30:18.390000",
          "content": "<p>The number 42 is especially significant to fans of science fiction novelist Douglas Adams’ “The Hitchhiker’s Guide to the Galaxy,” because that number is the answer given by a supercomputer to “the Ultimate Question of Life, the Universe, and Everything.” <br>\n<a href=\"https://news.mit.edu/2019/answer-life-universe-and-everything-sum-three-cubes-mathematics-0910\" target=\"_blank\">https://news.mit.edu/2019/answer-life-universe-and-everything-sum-three-cubes-mathematics-0910</a></p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 1819222,
          "author_name": "SgangX",
          "author_url": "",
          "post_date": "2022-06-13T14:46:20.287000",
          "content": "<p>Wow! That is an amazing answer. Maybe we are training the things not are models but life.😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1903417,
      "author_name": "Aditya Ray ",
      "author_url": "",
      "post_date": "2022-08-17T12:22:12.333000",
      "content": "<p>Thanks for sharing this; nice to confirm that randomness doesn't have a huge impact on the overall results</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1868705,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-07-24T06:59:43.553000",
      "content": "<p>This is a great effort, thanks for sharing the results! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1823804,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-06-17T16:54:08.867000",
      "content": "<p>Thank you for putting in all of this effort</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1821369,
      "author_name": "huseyincotel",
      "author_url": "",
      "post_date": "2022-06-15T13:28:20.400000",
      "content": "<p>How about using 600 seeds for validation set split and looking for correlation :S</p>\n<p>On more serious notice do you think dropping features with adversarial validation analysis can help with this? I still believe that may be the reason</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1825415,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-06-19T10:16:01.640000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/huseyincot\" target=\"_blank\">@huseyincot</a> Adversarial validation analysis is certainly a good thing to do - I haven't yet started it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1818714,
      "author_name": "pokoni",
      "author_url": "",
      "post_date": "2022-06-13T05:01:36.913000",
      "content": "<p>42 is all u need</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1817729,
      "author_name": "Rohan Dekate",
      "author_url": "",
      "post_date": "2022-06-11T17:43:55.017000",
      "content": "<p>Someone had to do this test! Thanks for the heads up…</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1817703,
      "author_name": "Ultron",
      "author_url": "",
      "post_date": "2022-06-11T17:03:46.180000",
      "content": "<p>What it seems from the Plot is that, for a specific split, the val_score across seeds kinda resemble a Gaussian distribution (close enough). In this case, for A the mean_val_score seems to be somewhere between ~0.792-0.793 and likewise for B. So why not just averaging results across seeds instead of searching the best seed? Just a pretty standard way to make the submissions robust to noise, imo. (As far as ‘seed’ is concerned.)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1818539,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-06-12T20:07:32.700000",
          "content": "<p><a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> Sure, ensembling across seeds is always a good method. With the experiment I wanted to investigate whether it makes sense to select good seeds before ensembling.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1817444,
      "author_name": "Atharv",
      "author_url": "",
      "post_date": "2022-06-11T10:57:34.530000",
      "content": "<p>I think this is the best seed -&gt; 42</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1816973,
      "author_name": "Burrito Dan",
      "author_url": "",
      "post_date": "2022-06-10T18:05:52.903000",
      "content": "<p>Another way to think about this is you will need to win by more than 0.005 on the public leaderboard to be confident it's not just noise.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1816718,
      "author_name": "RITAM UPADHYAY",
      "author_url": "",
      "post_date": "2022-06-10T13:48:35.830000",
      "content": "<p>Thanks for sharing an insightful experiment</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1815642,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-06-09T09:00:21.173000",
      "content": "<p>Fascinating insights. It is good to see that randomness doesn't greatly impact the overall results. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1815345,
      "author_name": "Making TARS",
      "author_url": "",
      "post_date": "2022-06-09T00:04:18.583000",
      "content": "<p>Impressive posting! Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1815824,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-09T13:51:55.147000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1815273,
      "author_name": "Kritdikoon Woraitthinan",
      "author_url": "",
      "post_date": "2022-06-08T21:53:48.190000",
      "content": "<p>Thank you for an interesting experiment.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1815823,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-09T13:51:20.357000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1815153": "Have you ever thought about training your model several times with different seeds and then selecting the seed with the best score for submission? Wouldn't it be nice to increase the lb score so easily? I wanted to know whether this idea works.\n\nFor the experiment, I split the training data into three parts: A training set (80 %) and two validation sets (10 % each, i.e. 45891 customers). (This is the same split as in the [discussion post about the noisy metric](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329088#1811782).) Then I trained 600 instances of my gradient-boosting model with 600 different seeds on this training set. For all 600 instances, I measured the score on both validation sets. If it were possible to select the best seed based on the cv score (or based on the public lb score) and so get a higher private lb score, this process could be simulated with the two validation sets. For the method to work, a good score on validation set A should predict a good score on validation set B. In other words, the scores on validation set A and B would need to be correlated.\n\nA 2d scatterplot of the validation scores immediately destroys every hope: The two scores are fully independent. Whether you select a model with a low score on dataset A (left part of the diagram) or one with a high score on dataset A (right part of the diagram), the expected outcome on dataset B is the same.\n\n![correlation-scatter](https://i.imgur.com/gNW5Fxt.png)\n\nFor a formal statistical test, we can use the function `scipy.stats.pearsonr()` to compute the Pearson correlation coefficient \\\\(r = -0.02\\\\) and the two-sided p-value, which is \\\\(p = 0.65\\\\). A p-value less than \\\\(0.05\\\\) might be considered a significant correlation; \\\\(0.65\\\\) is far above that.\n\nSummary: Don't try to find the best seed - it's impossible!",
    "1815215": "We all know the best Seed -> 42",
    "1815316": "Seeing this for gradient boosting makes sense to me, but I wonder if it might be a different story for neural network approaches. Randomness plays a pretty small role in the typical booster -- maybe some column and row subsampling -- but can play a bigger role in neural nets through parameter initialization and the downstream impact on model convergence in a wonky, nonconvex cost space. That's not to say that seed hacking is a good plan for DL, but there's a good reason why e.g. averaging models over multiple random weight initializations is a common ensembling technique.",
    "1833492": "Dear @ambrosm \nA similar discussion happened in an earlier competition- Petfinder where \n@haoge233 discussed how different seeds affected his score [here](https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/285580)\n@allohvk had raised a similar point, [here](https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/296706) and referenced the  paper - https://arxiv.org/pdf/2109.08203.pdf   which suggets \n>there are indeed seeds that produce scores sufficiently good to be considered as a significant improvement by the computer vision community. This is a worrying result as the community is currently very much score driven, and yet these can just be artifacts of randomness.\n\nHowever, this was not useful in the end. As many pointed out and saw post the private LB release and final results.\n\n@cdeotte noted [here](https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/280302#1551958):\n>If using a different seed changes your CV score, that means that the train data is small and/or your model has large variance in accuracy after training.\n> When this happens, you can either try changing your model (lower its variance), or you may have to run many different seeds (or same seed multiple times) and average all the CVs together to get a reliable CV score. \n\nAs a result, best is to stick your preferred seed and try to lower variance of the model.\nHope this helps!\n",
    "1815183": "Dear @ambrosm \n\nYou seem to have built up a sort of  \"ergodic hypothesis\" for this competition 😄!\nIn the \"niosy\" post you keep the seed constant and observe the fluctuations over time (*i.e.* number of iterations), and in this experiment you keep the time constant and create an ensemble of seeds. Both approaches indicating similar expectation values. Wonderful work!\n\nAll the best,\ncarl",
    "1815670": "For fun =)))\nhttps://arxiv.org/pdf/2109.08203.pdf",
    "1815571": "> **random_state** should only be used for reproducibility and should not give a better model.\n\nI had this notion in my mind, but today I saw it's experimental proof. thanks @ambrosm ",
    "1816682": "İ don't think that Seed is one of the important parameters that should be optimized in ML model. ",
    "1838921": "I always wanted to try this. Thanks for sharing your work!",
    "1825781": "That's Interesting.",
    "1822492": "42 indeed is the best seed :P",
    "1819443": "I just choose a random seed. But good to know the seed can sometimes improve score a little bit.",
    "1818620": "why 42 is best seed?",
    "1903417": "Thanks for sharing this; nice to confirm that randomness doesn't have a huge impact on the overall results",
    "1868705": "This is a great effort, thanks for sharing the results! ",
    "1823804": "Thank you for putting in all of this effort\n",
    "1821369": "How about using 600 seeds for validation set split and looking for correlation :S\n\nOn more serious notice do you think dropping features with adversarial validation analysis can help with this? I still believe that may be the reason",
    "1818714": "42 is all u need",
    "1817729": "Someone had to do this test! Thanks for the heads up...",
    "1817703": "What it seems from the Plot is that, for a specific split, the val_score across seeds kinda resemble a Gaussian distribution (close enough). In this case, for A the mean_val_score seems to be somewhere between ~0.792-0.793 and likewise for B. So why not just averaging results across seeds instead of searching the best seed? Just a pretty standard way to make the submissions robust to noise, imo. (As far as ‘seed’ is concerned.)",
    "1817444": "I think this is the best seed -> 42",
    "1816973": "Another way to think about this is you will need to win by more than 0.005 on the public leaderboard to be confident it's not just noise.",
    "1816718": "Thanks for sharing an insightful experiment",
    "1815642": "Fascinating insights. It is good to see that randomness doesn't greatly impact the overall results. ",
    "1815345": "Impressive posting! Thank you @ambrosm :)",
    "1815273": "Thank you for an interesting experiment.",
    "1815823": ""
  }
}