{
  "id": 217782,
  "title": "Same Code but Very Different Score on Local GPU vs. Colab",
  "url": "/competitions/rfcx-species-audio-detection/discussion/217782",
  "author_name": "",
  "post_date": "2021-02-08T11:17:48.815039700Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I'm experiencing a very weird issue while using Colab.</p>\n<p>Using exact same code, I'm getting way lower LB scores with ckpts from Colab V100 than from my own 2080Ti. While local 2080Ti produces lb ~0.830, Colab produces submissions mostly &lt; 0.80.</p>\n<p>I was wondering if anyone has encountered the issue before or knows how to address it.</p>\n<p>Randomness should be taken cared of as this function is called:</p>\n<pre><code>def seed_everithing(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code></pre>\n<p>and this one during mixup</p>\n<p><code>self.random_state = np.random.RandomState(random_seed)</code></p>\n<p>However, training stats on two devices are similar:<br>\nFor example: </p>\n<pre><code>Local: \nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0521 - LWLRAP:0.7823\nValid Loss:0.0620 - LWLRAP:0.8077\n\nColab\nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0541 - LWLRAP:0.7731\nValid Loss:0.0657 - LWLRAP:0.8067\n</code></pre>\n<p>For reference, my code is based on this public kernel: <br>\n<a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater</a></p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1191293",
      "postDate": "02/08/2021 11:17:48",
      "content": "<p>Hi All,</p>\n<p>I'm experiencing a very weird issue while using Colab.</p>\n<p>Using exact same code, I'm getting way lower LB scores with ckpts from Colab V100 than from my own 2080Ti. While local 2080Ti produces lb ~0.830, Colab produces submissions mostly &lt; 0.80.</p>\n<p>I was wondering if anyone has encountered the issue before or knows how to address it.</p>\n<p>Randomness should be taken cared of as this function is called:</p>\n<pre><code>def seed_everithing(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code></pre>\n<p>and this one during mixup</p>\n<p><code>self.random_state = np.random.RandomState(random_seed)</code></p>\n<p>However, training stats on two devices are similar:<br>\nFor example: </p>\n<pre><code>Local: \nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0521 - LWLRAP:0.7823\nValid Loss:0.0620 - LWLRAP:0.8077\n\nColab\nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0541 - LWLRAP:0.7731\nValid Loss:0.0657 - LWLRAP:0.8067\n</code></pre>\n<p>For reference, my code is based on this public kernel: <br>\n<a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater</a></p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi All,\n\nI'm experiencing a very weird issue while using Colab.\n\nUsing exact same code, I'm getting way lower LB scores with ckpts from Colab V100 than from my own 2080Ti. While local 2080Ti produces lb ~0.830, Colab produces submissions mostly < 0.80.\n\nI was wondering if anyone has encountered the issue before or knows how to address it.\n\nRandomness should be taken cared of as this function is called:\n```\ndef seed_everithing(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n```\nand this one during mixup\n\n`self.random_state = np.random.RandomState(random_seed)`\n\nHowever, training stats on two devices are similar:\nFor example: \n\n```\nLocal: \nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0521 - LWLRAP:0.7823\nValid Loss:0.0620 - LWLRAP:0.8077\n\nColab\nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0541 - LWLRAP:0.7731\nValid Loss:0.0657 - LWLRAP:0.8067\n```\n\n\nFor reference, my code is based on this public kernel: \nhttps://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\n\nThanks!",
      "votes": null
    },
    {
      "id": "1191459",
      "postDate": "02/08/2021 13:16:07",
      "content": "<blockquote>\n  <p>I was wondering if anyone has encountered the issue before or knows how to address it.</p>\n</blockquote>\n<p>I was able to fully eliminate randomness only on my local 2060S, and not on Kaggle P100. Might be an architecture limitaton for earlier GPU generations.</p>\n<p>If your Colab submission generates \"random\" training results which lie close to average for your model, then your local training could be using fully deterministic lucky seed and consistently outperform cloud GPUs.</p>",
      "rawMarkdown": "> I was wondering if anyone has encountered the issue before or knows how to address it.\n\nI was able to fully eliminate randomness only on my local 2060S, and not on Kaggle P100. Might be an architecture limitaton for earlier GPU generations.\n\nIf your Colab submission generates \"random\" training results which lie close to average for your model, then your local training could be using fully deterministic lucky seed and consistently outperform cloud GPUs.",
      "votes": null
    },
    {
      "id": "1193097",
      "postDate": "02/09/2021 13:09:22",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> I see. Thanks for your insights!</p>",
      "rawMarkdown": "fffrrt I see. Thanks for your insights!",
      "votes": null
    },
    {
      "id": "1193098",
      "postDate": "02/09/2021 13:09:58",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> BTW, I was wondering if you are interested in merging for ensemble?</p>",
      "rawMarkdown": "fffrrt BTW, I was wondering if you are interested in merging for ensemble?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1191459,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "02/08/2021 13:16:07",
      "content": "<blockquote>\n  <p>I was wondering if anyone has encountered the issue before or knows how to address it.</p>\n</blockquote>\n<p>I was able to fully eliminate randomness only on my local 2060S, and not on Kaggle P100. Might be an architecture limitaton for earlier GPU generations.</p>\n<p>If your Colab submission generates \"random\" training results which lie close to average for your model, then your local training could be using fully deterministic lucky seed and consistently outperform cloud GPUs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1193097,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/09/2021 13:09:22",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> I see. Thanks for your insights!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1193098,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/09/2021 13:09:58",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> BTW, I was wondering if you are interested in merging for ensemble?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1191293": "Hi All,\n\nI'm experiencing a very weird issue while using Colab.\n\nUsing exact same code, I'm getting way lower LB scores with ckpts from Colab V100 than from my own 2080Ti. While local 2080Ti produces lb ~0.830, Colab produces submissions mostly < 0.80.\n\nI was wondering if anyone has encountered the issue before or knows how to address it.\n\nRandomness should be taken cared of as this function is called:\n```\ndef seed_everithing(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n```\nand this one during mixup\n\n`self.random_state = np.random.RandomState(random_seed)`\n\nHowever, training stats on two devices are similar:\nFor example: \n\n```\nLocal: \nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0521 - LWLRAP:0.7823\nValid Loss:0.0620 - LWLRAP:0.8077\n\nColab\nFold:0, Epoch:48, lr:0.0003\nTrain Loss:0.0541 - LWLRAP:0.7731\nValid Loss:0.0657 - LWLRAP:0.8067\n```\n\n\nFor reference, my code is based on this public kernel: \nhttps://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\n\nThanks!",
    "1191459": "> I was wondering if anyone has encountered the issue before or knows how to address it.\n\nI was able to fully eliminate randomness only on my local 2060S, and not on Kaggle P100. Might be an architecture limitaton for earlier GPU generations.\n\nIf your Colab submission generates \"random\" training results which lie close to average for your model, then your local training could be using fully deterministic lucky seed and consistently outperform cloud GPUs.",
    "1193097": "fffrrt I see. Thanks for your insights!",
    "1193098": "fffrrt BTW, I was wondering if you are interested in merging for ensemble?"
  },
  "source": "meta"
}