{
  "id": 200182,
  "title": "How much is randomness affecting...",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200182",
  "author_name": "",
  "post_date": "2020-11-29T08:54:28.007970400Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have been trying to run 10 epochs same fold with all things seeded. But still cuda does have randomness and it is driving me crazy as I now dont know which results to trust. </p>\n<p>According to my cv (cv is stratified fold 0 with seed 42):-</p>\n<p>The same model trained first time scored 0.889 on cv and 0.885 the second time.</p>\n<p>Anyone else seeing results like this? What can be done to minimize this randomness.</p>\n<p>Thank You for your help.</p>",
  "messages": [
    {
      "id": "1095114",
      "postDate": "11/29/2020 08:54:28",
      "content": "<p>I have been trying to run 10 epochs same fold with all things seeded. But still cuda does have randomness and it is driving me crazy as I now dont know which results to trust. </p>\n<p>According to my cv (cv is stratified fold 0 with seed 42):-</p>\n<p>The same model trained first time scored 0.889 on cv and 0.885 the second time.</p>\n<p>Anyone else seeing results like this? What can be done to minimize this randomness.</p>\n<p>Thank You for your help.</p>",
      "rawMarkdown": "I have been trying to run 10 epochs same fold with all things seeded. But still cuda does have randomness and it is driving me crazy as I now dont know which results to trust. \n\nAccording to my cv (cv is stratified fold 0 with seed 42):-\n\nThe same model trained first time scored 0.889 on cv and 0.885 the second time.\n\nAnyone else seeing results like this? What can be done to minimize this randomness.\n\nThank You for your help.",
      "votes": null
    },
    {
      "id": "1095124",
      "postDate": "11/29/2020 09:17:49",
      "content": "<p>I have pretty much deterministic results with the usual seeding </p>\n<pre><code> def seed_everything(seed=42):\n     random.seed(seed)\n     os.environ['PYTHONHASHSEED'] = str(seed)\n     np.random.seed(seed)\n     torch.manual_seed(seed)\n     torch.cuda.manual_seed(seed)\n     torch.cuda.manual_seed_all(seed)\n     torch.backends.cudnn.deterministic = True\n     torch.backends.cudnn.benchmark = False\n</code></pre>",
      "rawMarkdown": "I have pretty much deterministic results with the usual seeding \n```\n\n def seed_everything(seed=42):\n     random.seed(seed)\n     os.environ['PYTHONHASHSEED'] = str(seed)\n     np.random.seed(seed)\n     torch.manual_seed(seed)\n     torch.cuda.manual_seed(seed)\n     torch.cuda.manual_seed_all(seed)\n     torch.backends.cudnn.deterministic = True\n     torch.backends.cudnn.benchmark = False\n\n\n```",
      "votes": null
    },
    {
      "id": "1095143",
      "postDate": "11/29/2020 09:45:17",
      "content": "<p>I use the exact same seeding but the results change on the third place after decimal.</p>",
      "rawMarkdown": "I use the exact same seeding but the results change on the third place after decimal.",
      "votes": null
    },
    {
      "id": "1095291",
      "postDate": "11/29/2020 12:56:35",
      "content": "<p>Do you use pretrained weights to start your training? </p>",
      "rawMarkdown": "Do you use pretrained weights to start your training?",
      "votes": null
    },
    {
      "id": "1095292",
      "postDate": "11/29/2020 12:57:47",
      "content": "<p>Ofcourse I do.</p>",
      "rawMarkdown": "Ofcourse I do.",
      "votes": null
    },
    {
      "id": "1095295",
      "postDate": "11/29/2020 13:06:19",
      "content": "<p>Thats not alway the case, thats why I'm asking. Your both CV results are not very different, it could be the random augmentations, the optimizer method (SGD or MBGD etc.) </p>\n<p>There is also a good amount of noisy data in our train dataset e.g: duplicates, duplicates with different labels, actual cassavas (potatoes) instead of the leaves etc.</p>",
      "rawMarkdown": "Thats not alway the case, thats why I'm asking. Your both CV results are not very different, it could be the random augmentations, the optimizer method (SGD or MBGD etc.) \n\nThere is also a good amount of noisy data in our train dataset e.g: duplicates, duplicates with different labels, actual cassavas (potatoes) instead of the leaves etc.",
      "votes": null
    },
    {
      "id": "1095341",
      "postDate": "11/29/2020 14:00:08",
      "content": "<p>do you re-initialize your weights before the second trial? Maybe try exiting and restarting the notebook so both trials are exactly the same.</p>",
      "rawMarkdown": "do you re-initialize your weights before the second trial? Maybe try exiting and restarting the notebook so both trials are exactly the same.",
      "votes": null
    },
    {
      "id": "1095346",
      "postDate": "11/29/2020 14:05:08",
      "content": "<p>I always restart and run the trials without doing anything which could cause me unstable seed.</p>",
      "rawMarkdown": "I always restart and run the trials without doing anything which could cause me unstable seed.",
      "votes": null
    },
    {
      "id": "1095349",
      "postDate": "11/29/2020 14:07:09",
      "content": "<p>I use Adam optimizer (Is weight decay random or calculated cuz I do use that). Augmentations are not random I have took care of that apparently that seed does seed albumentations library to always give the same result when kernel is restarted.</p>",
      "rawMarkdown": "I use Adam optimizer (Is weight decay random or calculated cuz I do use that). Augmentations are not random I have took care of that apparently that seed does seed albumentations library to always give the same result when kernel is restarted.",
      "votes": null
    },
    {
      "id": "1095449",
      "postDate": "11/29/2020 15:50:11",
      "content": "<p>I think this is totally normal from run to run, averaging multiple run is the only way to make sure. In my opinion if you don't see a significant improvement from your experiment then its probably in the margin of error.</p>",
      "rawMarkdown": "I think this is totally normal from run to run, averaging multiple run is the only way to make sure. In my opinion if you don't see a significant improvement from your experiment then its probably in the margin of error.",
      "votes": null
    },
    {
      "id": "1095456",
      "postDate": "11/29/2020 15:55:21",
      "content": "<p>I was trying to get affect of every augmentation, so the change in score would be slight and hard to tell if its a margin of error. It should run just fine when I average multiple folds.</p>",
      "rawMarkdown": "I was trying to get affect of every augmentation, so the change in score would be slight and hard to tell if its a margin of error. It should run just fine when I average multiple folds.",
      "votes": null
    },
    {
      "id": "1095472",
      "postDate": "11/29/2020 16:15:35",
      "content": "<p>Yeah augmentations are hard to evaluate due to randomness, sometimes your intuition and experience are more beneficial than testing them one by one. (What augs make more sense thats what i mean)</p>",
      "rawMarkdown": "Yeah augmentations are hard to evaluate due to randomness, sometimes your intuition and experience are more beneficial than testing them one by one. (What augs make more sense thats what i mean)",
      "votes": null
    },
    {
      "id": "1095491",
      "postDate": "11/29/2020 16:30:03",
      "content": "<p>I have got all the augmentations in the place the thing was how hard can I take it, I wanted to get every bit of performance augmentations would give which now can not be tested as a single fold 10 epochs takes ~3 hours on my 1080ti and with RANDOMNESS… ;-(</p>",
      "rawMarkdown": "I have got all the augmentations in the place the thing was how hard can I take it, I wanted to get every bit of performance augmentations would give which now can not be tested as a single fold 10 epochs takes ~3 hours on my 1080ti and with RANDOMNESS... ;-(",
      "votes": null
    },
    {
      "id": "1099613",
      "postDate": "12/02/2020 13:32:55",
      "content": "<p>I use the almost the same settings as <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> and I feel that every run is exactly the same with the same seed. </p>",
      "rawMarkdown": "I use the almost the same settings as @serigne and I feel that every run is exactly the same with the same seed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1095124,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "11/29/2020 09:17:49",
      "content": "<p>I have pretty much deterministic results with the usual seeding </p>\n<pre><code> def seed_everything(seed=42):\n     random.seed(seed)\n     os.environ['PYTHONHASHSEED'] = str(seed)\n     np.random.seed(seed)\n     torch.manual_seed(seed)\n     torch.cuda.manual_seed(seed)\n     torch.cuda.manual_seed_all(seed)\n     torch.backends.cudnn.deterministic = True\n     torch.backends.cudnn.benchmark = False\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1095143,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/29/2020 09:45:17",
          "content": "<p>I use the exact same seeding but the results change on the third place after decimal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095291,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/29/2020 12:56:35",
          "content": "<p>Do you use pretrained weights to start your training? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095292,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/29/2020 12:57:47",
          "content": "<p>Ofcourse I do.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095295,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/29/2020 13:06:19",
          "content": "<p>Thats not alway the case, thats why I'm asking. Your both CV results are not very different, it could be the random augmentations, the optimizer method (SGD or MBGD etc.) </p>\n<p>There is also a good amount of noisy data in our train dataset e.g: duplicates, duplicates with different labels, actual cassavas (potatoes) instead of the leaves etc.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1095341,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "11/29/2020 14:00:08",
              "content": "<p>do you re-initialize your weights before the second trial? Maybe try exiting and restarting the notebook so both trials are exactly the same.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1095346,
              "author_name": "harshitsheoran",
              "author_url": "",
              "post_date": "11/29/2020 14:05:08",
              "content": "<p>I always restart and run the trials without doing anything which could cause me unstable seed.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1095349,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/29/2020 14:07:09",
          "content": "<p>I use Adam optimizer (Is weight decay random or calculated cuz I do use that). Augmentations are not random I have took care of that apparently that seed does seed albumentations library to always give the same result when kernel is restarted.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099613,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "12/02/2020 13:32:55",
          "content": "<p>I use the almost the same settings as <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> and I feel that every run is exactly the same with the same seed. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1095449,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "11/29/2020 15:50:11",
      "content": "<p>I think this is totally normal from run to run, averaging multiple run is the only way to make sure. In my opinion if you don't see a significant improvement from your experiment then its probably in the margin of error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1095456,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/29/2020 15:55:21",
          "content": "<p>I was trying to get affect of every augmentation, so the change in score would be slight and hard to tell if its a margin of error. It should run just fine when I average multiple folds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095472,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "11/29/2020 16:15:35",
          "content": "<p>Yeah augmentations are hard to evaluate due to randomness, sometimes your intuition and experience are more beneficial than testing them one by one. (What augs make more sense thats what i mean)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1095491,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/29/2020 16:30:03",
          "content": "<p>I have got all the augmentations in the place the thing was how hard can I take it, I wanted to get every bit of performance augmentations would give which now can not be tested as a single fold 10 epochs takes ~3 hours on my 1080ti and with RANDOMNESS… ;-(</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1095114": "I have been trying to run 10 epochs same fold with all things seeded. But still cuda does have randomness and it is driving me crazy as I now dont know which results to trust. \n\nAccording to my cv (cv is stratified fold 0 with seed 42):-\n\nThe same model trained first time scored 0.889 on cv and 0.885 the second time.\n\nAnyone else seeing results like this? What can be done to minimize this randomness.\n\nThank You for your help.",
    "1095124": "I have pretty much deterministic results with the usual seeding \n```\n\n def seed_everything(seed=42):\n     random.seed(seed)\n     os.environ['PYTHONHASHSEED'] = str(seed)\n     np.random.seed(seed)\n     torch.manual_seed(seed)\n     torch.cuda.manual_seed(seed)\n     torch.cuda.manual_seed_all(seed)\n     torch.backends.cudnn.deterministic = True\n     torch.backends.cudnn.benchmark = False\n\n\n```",
    "1095143": "I use the exact same seeding but the results change on the third place after decimal.",
    "1095291": "Do you use pretrained weights to start your training?",
    "1095292": "Ofcourse I do.",
    "1095295": "Thats not alway the case, thats why I'm asking. Your both CV results are not very different, it could be the random augmentations, the optimizer method (SGD or MBGD etc.) \n\nThere is also a good amount of noisy data in our train dataset e.g: duplicates, duplicates with different labels, actual cassavas (potatoes) instead of the leaves etc.",
    "1095341": "do you re-initialize your weights before the second trial? Maybe try exiting and restarting the notebook so both trials are exactly the same.",
    "1095346": "I always restart and run the trials without doing anything which could cause me unstable seed.",
    "1095349": "I use Adam optimizer (Is weight decay random or calculated cuz I do use that). Augmentations are not random I have took care of that apparently that seed does seed albumentations library to always give the same result when kernel is restarted.",
    "1095449": "I think this is totally normal from run to run, averaging multiple run is the only way to make sure. In my opinion if you don't see a significant improvement from your experiment then its probably in the margin of error.",
    "1095456": "I was trying to get affect of every augmentation, so the change in score would be slight and hard to tell if its a margin of error. It should run just fine when I average multiple folds.",
    "1095472": "Yeah augmentations are hard to evaluate due to randomness, sometimes your intuition and experience are more beneficial than testing them one by one. (What augs make more sense thats what i mean)",
    "1095491": "I have got all the augmentations in the place the thing was how hard can I take it, I wanted to get every bit of performance augmentations would give which now can not be tested as a single fold 10 epochs takes ~3 hours on my 1080ti and with RANDOMNESS... ;-(",
    "1099613": "I use the almost the same settings as @serigne and I feel that every run is exactly the same with the same seed."
  },
  "source": "meta"
}