{
  "id": 310795,
  "title": "why  no high LB  pytorch notebooks shared?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/310795",
  "author_name": "",
  "post_date": "2022-03-03T06:07:48.057746700Z",
  "votes": 15,
  "comment_count": 15,
  "views": 0,
  "content": "<p>for my few experiments, It seems that more GPU/TPU RAM needed.</p>\n<p>Or more customized code?</p>",
  "messages": [
    {
      "id": "1710528",
      "postDate": "03/03/2022 06:07:48",
      "content": "<p>for my few experiments, It seems that more GPU/TPU RAM needed.</p>\n<p>Or more customized code?</p>",
      "rawMarkdown": "for my few experiments, It seems that more GPU/TPU RAM needed.\n\nOr more customized code?",
      "votes": null
    },
    {
      "id": "1710634",
      "postDate": "03/03/2022 07:33:20",
      "content": "<p>Pytorch is too slow on Colab/Kaggle GPU compared to TF + TPU</p>",
      "rawMarkdown": "Pytorch is too slow on Colab/Kaggle GPU compared to TF + TPU",
      "votes": null
    },
    {
      "id": "1710689",
      "postDate": "03/03/2022 08:41:57",
      "content": "<p>I guess the reason is not in \"more GPU\" but time</p>",
      "rawMarkdown": "I guess the reason is not in \"more GPU\" but time",
      "votes": null
    },
    {
      "id": "1710846",
      "postDate": "03/03/2022 11:45:15",
      "content": "<p>can you clarify what do you mean by that? i have no experience with modern tf/tpu </p>",
      "rawMarkdown": "can you clarify what do you mean by that? i have no experience with modern tf/tpu",
      "votes": null
    },
    {
      "id": "1710899",
      "postDate": "03/03/2022 13:06:01",
      "content": "<p>I mean you have limit time for GPU/TPU per week in notebook, may be it is not enough for good solution, because we have a lot of data</p>",
      "rawMarkdown": "I mean you have limit time for GPU/TPU per week in notebook, may be it is not enough for good solution, because we have a lot of data",
      "votes": null
    },
    {
      "id": "1711021",
      "postDate": "03/03/2022 15:10:41",
      "content": "<p>Not even inference notebooks are shared. I don't think it is about the training time.</p>",
      "rawMarkdown": "Not even inference notebooks are shared. I don't think it is about the training time.",
      "votes": null
    },
    {
      "id": "1711040",
      "postDate": "03/03/2022 15:29:58",
      "content": "<p>My guess is because of the ease of use of TensorFlow with TPUs (they work out of the box). I've been doing experiments with PyTorch and the time to train on Kaggle or Colab's GPU (Tesla P100) even with efficientnet_b0 on considerable image size is a lot. Plus, you don't have to download the data when using TensorFlow + TPUs as it is streamed directly from GCS. I know there are ways to read TFRecords with PyTorch but it involves extra work and making TPUs work with PyTorch is a nightmare (atleast for me) :(</p>\n<p>And the first high-scoring baseline with inference was shared by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> and people usually settle with the highest voted and first shared baseline</p>",
      "rawMarkdown": "My guess is because of the ease of use of TensorFlow with TPUs (they work out of the box). I've been doing experiments with PyTorch and the time to train on Kaggle or Colab's GPU (Tesla P100) even with efficientnet_b0 on considerable image size is a lot. Plus, you don't have to download the data when using TensorFlow + TPUs as it is streamed directly from GCS. I know there are ways to read TFRecords with PyTorch but it involves extra work and making TPUs work with PyTorch is a nightmare (atleast for me) :(\n\nAnd the first high-scoring baseline with inference was shared by @ks2019 and people usually settle with the highest voted and first shared baseline",
      "votes": null
    },
    {
      "id": "1711173",
      "postDate": "03/03/2022 17:28:17",
      "content": "<p>there are  shared pytorch  inference notebooks. but with small models/low LB.</p>",
      "rawMarkdown": "there are  shared pytorch  inference notebooks. but with small models/low LB.",
      "votes": null
    },
    {
      "id": "1711683",
      "postDate": "03/04/2022 07:49:17",
      "content": "<p>Interesting.  When I use TF on colab, almost no problem.  However when I use pytorch model for training,  gdrive often timeout!!!  </p>",
      "rawMarkdown": "Interesting.  When I use TF on colab, almost no problem.  However when I use pytorch model for training,  gdrive often timeout!!!",
      "votes": null
    },
    {
      "id": "1712224",
      "postDate": "03/04/2022 17:23:33",
      "content": "<p>There are limitation indeed. In my public pytorch notebook (<a href=\"https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling\" target=\"_blank\">https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling</a>) the result for 1 fold was 0.484 . Of course I can improve this result by further tuning but without a impressive GPU system is hard to reach the TF scores.<br>\nAlthough the trick to the competition lies indeed in image pre-processing techniques and post convolutional approaches, from my experience so far with this competition and using ArcFace systems in the convolutional part bigger architecture larger images size and from my results so far larger batch size especially leads to bigger performance (it seems that from my experiences so far larger batch size with ArcFace helps). So even if I architecture some special techniques for pre and post processing, I will still get better results using the TPU system.<br>\nOn my local machine (3x2080TI) I can run B6 <a href=\"https://www.kaggle.com/768x768\" target=\"_blank\">@768x768</a> with batch size 9 (using apex). On TPU you can run B7 <a href=\"https://www.kaggle.com/768x768\" target=\"_blank\">@768x768</a> with batch size 128.  <br>\nMy methodology having limited TPU hours is to make my tests on my local setup using Pytorch and every time I am having a breakthrough to port the solution to TPU</p>",
      "rawMarkdown": "There are limitation indeed. In my public pytorch notebook (https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling) the result for 1 fold was 0.484 . Of course I can improve this result by further tuning but without a impressive GPU system is hard to reach the TF scores.\nAlthough the trick to the competition lies indeed in image pre-processing techniques and post convolutional approaches, from my experience so far with this competition and using ArcFace systems in the convolutional part bigger architecture larger images size and from my results so far larger batch size especially leads to bigger performance (it seems that from my experiences so far larger batch size with ArcFace helps). So even if I architecture some special techniques for pre and post processing, I will still get better results using the TPU system.\nOn my local machine (3x2080TI) I can run B6 @768x768 with batch size 9 (using apex). On TPU you can run B7 @768x768 with batch size 128.  \nMy methodology having limited TPU hours is to make my tests on my local setup using Pytorch and every time I am having a breakthrough to port the solution to TPU",
      "votes": null
    },
    {
      "id": "1712246",
      "postDate": "03/04/2022 18:03:39",
      "content": "<p>thanks for sharing.</p>",
      "rawMarkdown": "thanks for sharing.",
      "votes": null
    },
    {
      "id": "1712487",
      "postDate": "03/05/2022 01:28:57",
      "content": "<p>Does  n_accumulate  work?</p>",
      "rawMarkdown": "Does  n_accumulate  work?",
      "votes": null
    },
    {
      "id": "1712813",
      "postDate": "03/05/2022 11:40:33",
      "content": "<p>Did not try yet in this competition context. Usually gradient accumulations has a bad effect over batch_norm layers and I did not have good experience with this technique. But, it is on my todo experiments list</p>",
      "rawMarkdown": "Did not try yet in this competition context. Usually gradient accumulations has a bad effect over batch_norm layers and I did not have good experience with this technique. But, it is on my todo experiments list",
      "votes": null
    },
    {
      "id": "1714452",
      "postDate": "03/07/2022 03:00:22",
      "content": "<p>TPU v3 x8 ~ 4x V100 32G, thus even a similar torch solution released, most of user could not reproduce it.</p>",
      "rawMarkdown": "TPU v3 x8 ~ 4x V100 32G, thus even a similar torch solution released, most of user could not reproduce it.",
      "votes": null
    },
    {
      "id": "1715014",
      "postDate": "03/07/2022 14:59:58",
      "content": "<p>That makes sense.  Also I find that kaggle TPU version is better than the colab pro TPU.</p>",
      "rawMarkdown": "That makes sense.  Also I find that kaggle TPU version is better than the colab pro TPU.",
      "votes": null
    },
    {
      "id": "1715429",
      "postDate": "03/08/2022 02:01:32",
      "content": "<p>colab one is TPU v2</p>",
      "rawMarkdown": "colab one is TPU v2",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1710634,
      "author_name": "ptran1203",
      "author_url": "",
      "post_date": "03/03/2022 07:33:20",
      "content": "<p>Pytorch is too slow on Colab/Kaggle GPU compared to TF + TPU</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1710689,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "03/03/2022 08:41:57",
      "content": "<p>I guess the reason is not in \"more GPU\" but time</p>",
      "votes": null,
      "replies": [
        {
          "id": 1710846,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "03/03/2022 11:45:15",
          "content": "<p>can you clarify what do you mean by that? i have no experience with modern tf/tpu </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1710899,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "03/03/2022 13:06:01",
          "content": "<p>I mean you have limit time for GPU/TPU per week in notebook, may be it is not enough for good solution, because we have a lot of data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1711021,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "03/03/2022 15:10:41",
          "content": "<p>Not even inference notebooks are shared. I don't think it is about the training time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1711173,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "03/03/2022 17:28:17",
          "content": "<p>there are  shared pytorch  inference notebooks. but with small models/low LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1711040,
      "author_name": "atharvaingle",
      "author_url": "",
      "post_date": "03/03/2022 15:29:58",
      "content": "<p>My guess is because of the ease of use of TensorFlow with TPUs (they work out of the box). I've been doing experiments with PyTorch and the time to train on Kaggle or Colab's GPU (Tesla P100) even with efficientnet_b0 on considerable image size is a lot. Plus, you don't have to download the data when using TensorFlow + TPUs as it is streamed directly from GCS. I know there are ways to read TFRecords with PyTorch but it involves extra work and making TPUs work with PyTorch is a nightmare (atleast for me) :(</p>\n<p>And the first high-scoring baseline with inference was shared by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> and people usually settle with the highest voted and first shared baseline</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1711683,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "03/04/2022 07:49:17",
      "content": "<p>Interesting.  When I use TF on colab, almost no problem.  However when I use pytorch model for training,  gdrive often timeout!!!  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1712224,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "03/04/2022 17:23:33",
      "content": "<p>There are limitation indeed. In my public pytorch notebook (<a href=\"https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling\" target=\"_blank\">https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling</a>) the result for 1 fold was 0.484 . Of course I can improve this result by further tuning but without a impressive GPU system is hard to reach the TF scores.<br>\nAlthough the trick to the competition lies indeed in image pre-processing techniques and post convolutional approaches, from my experience so far with this competition and using ArcFace systems in the convolutional part bigger architecture larger images size and from my results so far larger batch size especially leads to bigger performance (it seems that from my experiences so far larger batch size with ArcFace helps). So even if I architecture some special techniques for pre and post processing, I will still get better results using the TPU system.<br>\nOn my local machine (3x2080TI) I can run B6 <a href=\"https://www.kaggle.com/768x768\" target=\"_blank\">@768x768</a> with batch size 9 (using apex). On TPU you can run B7 <a href=\"https://www.kaggle.com/768x768\" target=\"_blank\">@768x768</a> with batch size 128.  <br>\nMy methodology having limited TPU hours is to make my tests on my local setup using Pytorch and every time I am having a breakthrough to port the solution to TPU</p>",
      "votes": null,
      "replies": [
        {
          "id": 1712246,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "03/04/2022 18:03:39",
          "content": "<p>thanks for sharing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1712487,
          "author_name": "chihantsai",
          "author_url": "",
          "post_date": "03/05/2022 01:28:57",
          "content": "<p>Does  n_accumulate  work?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1712813,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "03/05/2022 11:40:33",
          "content": "<p>Did not try yet in this competition context. Usually gradient accumulations has a bad effect over batch_norm layers and I did not have good experience with this technique. But, it is on my todo experiments list</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1714452,
      "author_name": "steamedsheep",
      "author_url": "",
      "post_date": "03/07/2022 03:00:22",
      "content": "<p>TPU v3 x8 ~ 4x V100 32G, thus even a similar torch solution released, most of user could not reproduce it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1715014,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "03/07/2022 14:59:58",
          "content": "<p>That makes sense.  Also I find that kaggle TPU version is better than the colab pro TPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1715429,
          "author_name": "steamedsheep",
          "author_url": "",
          "post_date": "03/08/2022 02:01:32",
          "content": "<p>colab one is TPU v2</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1710528": "for my few experiments, It seems that more GPU/TPU RAM needed.\n\nOr more customized code?",
    "1710634": "Pytorch is too slow on Colab/Kaggle GPU compared to TF + TPU",
    "1710689": "I guess the reason is not in \"more GPU\" but time",
    "1710846": "can you clarify what do you mean by that? i have no experience with modern tf/tpu",
    "1710899": "I mean you have limit time for GPU/TPU per week in notebook, may be it is not enough for good solution, because we have a lot of data",
    "1711021": "Not even inference notebooks are shared. I don't think it is about the training time.",
    "1711040": "My guess is because of the ease of use of TensorFlow with TPUs (they work out of the box). I've been doing experiments with PyTorch and the time to train on Kaggle or Colab's GPU (Tesla P100) even with efficientnet_b0 on considerable image size is a lot. Plus, you don't have to download the data when using TensorFlow + TPUs as it is streamed directly from GCS. I know there are ways to read TFRecords with PyTorch but it involves extra work and making TPUs work with PyTorch is a nightmare (atleast for me) :(\n\nAnd the first high-scoring baseline with inference was shared by @ks2019 and people usually settle with the highest voted and first shared baseline",
    "1711173": "there are  shared pytorch  inference notebooks. but with small models/low LB.",
    "1711683": "Interesting.  When I use TF on colab, almost no problem.  However when I use pytorch model for training,  gdrive often timeout!!!",
    "1712224": "There are limitation indeed. In my public pytorch notebook (https://www.kaggle.com/vladvdv/pytorch-inference-notebok-arcface-gem-pooling) the result for 1 fold was 0.484 . Of course I can improve this result by further tuning but without a impressive GPU system is hard to reach the TF scores.\nAlthough the trick to the competition lies indeed in image pre-processing techniques and post convolutional approaches, from my experience so far with this competition and using ArcFace systems in the convolutional part bigger architecture larger images size and from my results so far larger batch size especially leads to bigger performance (it seems that from my experiences so far larger batch size with ArcFace helps). So even if I architecture some special techniques for pre and post processing, I will still get better results using the TPU system.\nOn my local machine (3x2080TI) I can run B6 @768x768 with batch size 9 (using apex). On TPU you can run B7 @768x768 with batch size 128.  \nMy methodology having limited TPU hours is to make my tests on my local setup using Pytorch and every time I am having a breakthrough to port the solution to TPU",
    "1712246": "thanks for sharing.",
    "1712487": "Does  n_accumulate  work?",
    "1712813": "Did not try yet in this competition context. Usually gradient accumulations has a bad effect over batch_norm layers and I did not have good experience with this technique. But, it is on my todo experiments list",
    "1714452": "TPU v3 x8 ~ 4x V100 32G, thus even a similar torch solution released, most of user could not reproduce it.",
    "1715014": "That makes sense.  Also I find that kaggle TPU version is better than the colab pro TPU.",
    "1715429": "colab one is TPU v2"
  },
  "source": "meta"
}