{
  "id": 237507,
  "title": "Same Code, Different Performance on Kaggle Kernel vs. Colab",
  "url": "/competitions/birdclef-2021/discussion/237507",
  "author_name": "",
  "post_date": "2021-05-09T05:19:00.906169700Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I took code from this notebook: <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/notebook\" target=\"_blank\">https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/notebook</a> and tried to train for 12 epochs on both colab and kaggle. I made sure same data are used, and code is same except the data downloading part. Random seeds are also set to the same in both setups. </p>\n<p>However, During training, Kaggle and Colab resulted in different slightly different training/validation metrics. During inference and validation on training soundscapes, the performance gap becomes even more significant. The checkpoint from the 12th epoch in Kaggle version got a validation f1 of 0.642736 and lb of 0.59. However, the checkpoint from Colab got a validation f1 of only 0.556972 and lb of 0.55.</p>\n<p>I also experienced this issue in the RFCX competition but didn't figure out why. I was wondering if anyone experienced the same issue or know the solution. Thanks!</p>\n<p>Here's my Kaggle training notebook: <a href=\"https://www.kaggle.com/tonychenxyz/clean-fast-simple-bird-identifier-training-colab\" target=\"_blank\">https://www.kaggle.com/tonychenxyz/clean-fast-simple-bird-identifier-training-colab</a><br>\nHere's my colab notebook: <a href=\"https://colab.research.google.com/drive/1aCP6vMbwiUPwpXC1ASk8KBOB91yXaWca?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1aCP6vMbwiUPwpXC1ASk8KBOB91yXaWca?usp=sharing</a></p>",
  "messages": [
    {
      "id": "1298645",
      "postDate": "05/09/2021 05:19:00",
      "content": "<p>I took code from this notebook: <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/notebook\" target=\"_blank\">https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/notebook</a> and tried to train for 12 epochs on both colab and kaggle. I made sure same data are used, and code is same except the data downloading part. Random seeds are also set to the same in both setups. </p>\n<p>However, During training, Kaggle and Colab resulted in different slightly different training/validation metrics. During inference and validation on training soundscapes, the performance gap becomes even more significant. The checkpoint from the 12th epoch in Kaggle version got a validation f1 of 0.642736 and lb of 0.59. However, the checkpoint from Colab got a validation f1 of only 0.556972 and lb of 0.55.</p>\n<p>I also experienced this issue in the RFCX competition but didn't figure out why. I was wondering if anyone experienced the same issue or know the solution. Thanks!</p>\n<p>Here's my Kaggle training notebook: <a href=\"https://www.kaggle.com/tonychenxyz/clean-fast-simple-bird-identifier-training-colab\" target=\"_blank\">https://www.kaggle.com/tonychenxyz/clean-fast-simple-bird-identifier-training-colab</a><br>\nHere's my colab notebook: <a href=\"https://colab.research.google.com/drive/1aCP6vMbwiUPwpXC1ASk8KBOB91yXaWca?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1aCP6vMbwiUPwpXC1ASk8KBOB91yXaWca?usp=sharing</a></p>",
      "rawMarkdown": "I took code from this notebook: https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/notebook and tried to train for 12 epochs on both colab and kaggle. I made sure same data are used, and code is same except the data downloading part. Random seeds are also set to the same in both setups. \n\nHowever, During training, Kaggle and Colab resulted in different slightly different training/validation metrics. During inference and validation on training soundscapes, the performance gap becomes even more significant. The checkpoint from the 12th epoch in Kaggle version got a validation f1 of 0.642736 and lb of 0.59. However, the checkpoint from Colab got a validation f1 of only 0.556972 and lb of 0.55.\n\nI also experienced this issue in the RFCX competition but didn't figure out why. I was wondering if anyone experienced the same issue or know the solution. Thanks!\n\nHere's my Kaggle training notebook: https://www.kaggle.com/tonychenxyz/clean-fast-simple-bird-identifier-training-colab\nHere's my colab notebook: https://colab.research.google.com/drive/1aCP6vMbwiUPwpXC1ASk8KBOB91yXaWca?usp=sharing",
      "votes": null
    },
    {
      "id": "1298857",
      "postDate": "05/09/2021 09:19:04",
      "content": "<p>Are cuda and pytorch versions the same?  To get these, use this in two cells in your notebook;</p>\n<p><code>! nvidia-smi</code></p>\n<p><code>torch.__version__</code></p>",
      "rawMarkdown": "Are cuda and pytorch versions the same?  To get these, use this in two cells in your notebook;\n\n`! nvidia-smi`\n\n`torch.__version__`",
      "votes": null
    },
    {
      "id": "1298871",
      "postDate": "05/09/2021 09:30:28",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks for pointing out! I just checked and they are indeed different! Kaggle has CUDA 11.0 and torch 1.7.0, while colab has new versions for CUDA 11.2 and torch 1.8.1. I will do some further investigations to figure out whether they caused problems. Thanks again!</p>",
      "rawMarkdown": "cpmpml Thanks for pointing out! I just checked and they are indeed different! Kaggle has CUDA 11.0 and torch 1.7.0, while colab has new versions for CUDA 11.2 and torch 1.8.1. I will do some further investigations to figure out whether they caused problems. Thanks again!",
      "votes": null
    },
    {
      "id": "1299120",
      "postDate": "05/09/2021 13:38:35",
      "content": "<p>If you use pretrained models then you can also check the version you use.</p>",
      "rawMarkdown": "If you use pretrained models then you can also check the version you use.",
      "votes": null
    },
    {
      "id": "1300305",
      "postDate": "05/10/2021 11:37:19",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks! I checked pretrained model versions, and both use v0.0.5 of resnest, and they are downloading pretrained weights from the same url. Tried to just downgrade pytorch to 1.7.0 in colab, still got inferior validation and lb metrics. </p>\n<p>I also tried to downgrade both cuda and torch. I got similar lb score (0.58 vs. 0.59 with kaggle kernel) but lower validation score with train_soundscape (0.584681 vs. 0.642736 with kaggle kernel). </p>\n<p>I did this by reinstalling pytorch with conda specifying both older versions of pytorch and cudatoolkit. One thing to note that is by doing this, I can get same versions from nvcc --version and torch.version.cuda, but nvdia-smi still shows different CUDA versions.</p>",
      "rawMarkdown": "cpmpml Thanks! I checked pretrained model versions, and both use v0.0.5 of resnest, and they are downloading pretrained weights from the same url. Tried to just downgrade pytorch to 1.7.0 in colab, still got inferior validation and lb metrics. \n\nI also tried to downgrade both cuda and torch. I got similar lb score (0.58 vs. 0.59 with kaggle kernel) but lower validation score with train_soundscape (0.584681 vs. 0.642736 with kaggle kernel). \n\nI did this by reinstalling pytorch with conda specifying both older versions of pytorch and cudatoolkit. One thing to note that is by doing this, I can get same versions from nvcc --version and torch.version.cuda, but nvdia-smi still shows different CUDA versions.",
      "votes": null
    },
    {
      "id": "1300343",
      "postDate": "05/10/2021 12:03:15",
      "content": "<p>AFAIK, nvidia-smi shows the cuda version is was compiled with. If you did not reisnstall nvidia smi when you reinstalled cuda then nvidia-smi doe snot display the right cuda version.</p>",
      "rawMarkdown": "AFAIK, nvidia-smi shows the cuda version is was compiled with. If you did not reisnstall nvidia smi when you reinstalled cuda then nvidia-smi doe snot display the right cuda version.",
      "votes": null
    },
    {
      "id": "1300374",
      "postDate": "05/10/2021 12:19:51",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks for the explaination. So that means it should be fine despite nvidia-smi says different things as long as torch is using the same one?</p>",
      "rawMarkdown": "cpmpml Thanks for the explaination. So that means it should be fine despite nvidia-smi says different things as long as torch is using the same one?",
      "votes": null
    },
    {
      "id": "1300413",
      "postDate": "05/10/2021 12:48:50",
      "content": "<p><code>nvcc --version</code> is more reliable AFAIK.</p>",
      "rawMarkdown": "`nvcc --version` is more reliable AFAIK.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1298857,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2021 09:19:04",
      "content": "<p>Are cuda and pytorch versions the same?  To get these, use this in two cells in your notebook;</p>\n<p><code>! nvidia-smi</code></p>\n<p><code>torch.__version__</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1298871,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "05/09/2021 09:30:28",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks for pointing out! I just checked and they are indeed different! Kaggle has CUDA 11.0 and torch 1.7.0, while colab has new versions for CUDA 11.2 and torch 1.8.1. I will do some further investigations to figure out whether they caused problems. Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1299120,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/09/2021 13:38:35",
          "content": "<p>If you use pretrained models then you can also check the version you use.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1300305,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "05/10/2021 11:37:19",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks! I checked pretrained model versions, and both use v0.0.5 of resnest, and they are downloading pretrained weights from the same url. Tried to just downgrade pytorch to 1.7.0 in colab, still got inferior validation and lb metrics. </p>\n<p>I also tried to downgrade both cuda and torch. I got similar lb score (0.58 vs. 0.59 with kaggle kernel) but lower validation score with train_soundscape (0.584681 vs. 0.642736 with kaggle kernel). </p>\n<p>I did this by reinstalling pytorch with conda specifying both older versions of pytorch and cudatoolkit. One thing to note that is by doing this, I can get same versions from nvcc --version and torch.version.cuda, but nvdia-smi still shows different CUDA versions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1300343,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/10/2021 12:03:15",
          "content": "<p>AFAIK, nvidia-smi shows the cuda version is was compiled with. If you did not reisnstall nvidia smi when you reinstalled cuda then nvidia-smi doe snot display the right cuda version.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1300374,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "05/10/2021 12:19:51",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thanks for the explaination. So that means it should be fine despite nvidia-smi says different things as long as torch is using the same one?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1300413,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/10/2021 12:48:50",
          "content": "<p><code>nvcc --version</code> is more reliable AFAIK.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1298645": "I took code from this notebook: https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab/notebook and tried to train for 12 epochs on both colab and kaggle. I made sure same data are used, and code is same except the data downloading part. Random seeds are also set to the same in both setups. \n\nHowever, During training, Kaggle and Colab resulted in different slightly different training/validation metrics. During inference and validation on training soundscapes, the performance gap becomes even more significant. The checkpoint from the 12th epoch in Kaggle version got a validation f1 of 0.642736 and lb of 0.59. However, the checkpoint from Colab got a validation f1 of only 0.556972 and lb of 0.55.\n\nI also experienced this issue in the RFCX competition but didn't figure out why. I was wondering if anyone experienced the same issue or know the solution. Thanks!\n\nHere's my Kaggle training notebook: https://www.kaggle.com/tonychenxyz/clean-fast-simple-bird-identifier-training-colab\nHere's my colab notebook: https://colab.research.google.com/drive/1aCP6vMbwiUPwpXC1ASk8KBOB91yXaWca?usp=sharing",
    "1298857": "Are cuda and pytorch versions the same?  To get these, use this in two cells in your notebook;\n\n`! nvidia-smi`\n\n`torch.__version__`",
    "1298871": "cpmpml Thanks for pointing out! I just checked and they are indeed different! Kaggle has CUDA 11.0 and torch 1.7.0, while colab has new versions for CUDA 11.2 and torch 1.8.1. I will do some further investigations to figure out whether they caused problems. Thanks again!",
    "1299120": "If you use pretrained models then you can also check the version you use.",
    "1300305": "cpmpml Thanks! I checked pretrained model versions, and both use v0.0.5 of resnest, and they are downloading pretrained weights from the same url. Tried to just downgrade pytorch to 1.7.0 in colab, still got inferior validation and lb metrics. \n\nI also tried to downgrade both cuda and torch. I got similar lb score (0.58 vs. 0.59 with kaggle kernel) but lower validation score with train_soundscape (0.584681 vs. 0.642736 with kaggle kernel). \n\nI did this by reinstalling pytorch with conda specifying both older versions of pytorch and cudatoolkit. One thing to note that is by doing this, I can get same versions from nvcc --version and torch.version.cuda, but nvdia-smi still shows different CUDA versions.",
    "1300343": "AFAIK, nvidia-smi shows the cuda version is was compiled with. If you did not reisnstall nvidia smi when you reinstalled cuda then nvidia-smi doe snot display the right cuda version.",
    "1300374": "cpmpml Thanks for the explaination. So that means it should be fine despite nvidia-smi says different things as long as torch is using the same one?",
    "1300413": "`nvcc --version` is more reliable AFAIK."
  },
  "source": "meta"
}