{
  "id": 157109,
  "title": "[Torch XLA] Melanoma Crazy Fast",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/157109",
  "author_name": "",
  "post_date": "2020-06-09T10:57:32.099493700Z",
  "votes": 19,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>I have created kernels using TPU with PyTorch models for this competition! </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast\">[Torch XLA] Melanoma Crazy Fast</a></li>\n<li><a href=\"https://www.kaggle.com/shonenkov/inference-melanoma-crazy-fast\">[Inference] Melanoma Crazy Fast</a></li>\n</ul>\n\n<p>Also I have prepared this baseline on Colab! You can make a copy and start research! I recommend to use Colab Pro version with HIGH-RAM mode and TPU Runtime Type. Your speed of research will increase ~5-10 times! Link you can find in training kernel.</p>\n\n<p>Inference kernel has some might helpful information about stable prediction and example with calculation OOF scores. Hope it helps you!</p>\n\n<p>Welcome! </p>\n\n<h3>1. 256x256 image size, resnext50d_32x4d, 1 epoch:</h3>\n\n<ul>\n<li>Kaggle P100:  ~220s (train), ~60s (validation)</li>\n<li>Kaggle TPU:  ~75s (train), ~30s (validation) !Be carefull with BUGS</li>\n<li>Colab TPU:  ~60s (train), ~35s (validation)</li>\n</ul>\n\n<h3>2. 256x256 image size, efficientnet-b5, 1 epoch:</h3>\n\n<ul>\n<li>Kaggle P100:  ~350s (train), ~60s (validation)</li>\n<li>Kaggle TPU:  ~105s (train), ~30s (validation)  !Be carefull with BUGS</li>\n<li>Colab TPU:  ~50s (train), ~25s (validation)</li>\n</ul>\n\n<p>P.S. If you don't understand what effect make these kernels, please, skip!! Please, don't write toxic comments /: it is disheartening. Thank you!</p>",
  "messages": [
    {
      "id": "879264",
      "postDate": "06/09/2020 10:57:32",
      "content": "<p>Hi everyone!</p>\n\n<p>I have created kernels using TPU with PyTorch models for this competition! </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast\">[Torch XLA] Melanoma Crazy Fast</a></li>\n<li><a href=\"https://www.kaggle.com/shonenkov/inference-melanoma-crazy-fast\">[Inference] Melanoma Crazy Fast</a></li>\n</ul>\n\n<p>Also I have prepared this baseline on Colab! You can make a copy and start research! I recommend to use Colab Pro version with HIGH-RAM mode and TPU Runtime Type. Your speed of research will increase ~5-10 times! Link you can find in training kernel.</p>\n\n<p>Inference kernel has some might helpful information about stable prediction and example with calculation OOF scores. Hope it helps you!</p>\n\n<p>Welcome! </p>\n\n<h3>1. 256x256 image size, resnext50d_32x4d, 1 epoch:</h3>\n\n<ul>\n<li>Kaggle P100:  ~220s (train), ~60s (validation)</li>\n<li>Kaggle TPU:  ~75s (train), ~30s (validation) !Be carefull with BUGS</li>\n<li>Colab TPU:  ~60s (train), ~35s (validation)</li>\n</ul>\n\n<h3>2. 256x256 image size, efficientnet-b5, 1 epoch:</h3>\n\n<ul>\n<li>Kaggle P100:  ~350s (train), ~60s (validation)</li>\n<li>Kaggle TPU:  ~105s (train), ~30s (validation)  !Be carefull with BUGS</li>\n<li>Colab TPU:  ~50s (train), ~25s (validation)</li>\n</ul>\n\n<p>P.S. If you don't understand what effect make these kernels, please, skip!! Please, don't write toxic comments /: it is disheartening. Thank you!</p>",
      "rawMarkdown": "Hi everyone!\n\nI have created kernels using TPU with PyTorch models for this competition! \n\n- [[Torch XLA] Melanoma Crazy Fast](https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast)\n- [[Inference] Melanoma Crazy Fast](https://www.kaggle.com/shonenkov/inference-melanoma-crazy-fast)\n\nAlso I have prepared this baseline on Colab! You can make a copy and start research! I recommend to use Colab Pro version with HIGH-RAM mode and TPU Runtime Type. Your speed of research will increase ~5-10 times! Link you can find in training kernel.\n\nInference kernel has some might helpful information about stable prediction and example with calculation OOF scores. Hope it helps you!\n\nWelcome! \n\n### 1. 256x256 image size, resnext50d_32x4d, 1 epoch:\n- Kaggle P100:  ~220s (train), ~60s (validation)\n- Kaggle TPU:  ~75s (train), ~30s (validation) !Be carefull with BUGS\n- Colab TPU:  ~60s (train), ~35s (validation)\n\n### 2. 256x256 image size, efficientnet-b5, 1 epoch:\n- Kaggle P100:  ~350s (train), ~60s (validation)\n- Kaggle TPU:  ~105s (train), ~30s (validation)  !Be carefull with BUGS\n- Colab TPU:  ~50s (train), ~25s (validation)\n\n\nP.S. If you don't understand what effect make these kernels, please, skip!! Please, don't write toxic comments /: it is disheartening. Thank you!",
      "votes": null
    },
    {
      "id": "879994",
      "postDate": "06/09/2020 22:30:35",
      "content": "<p>Why didn't you include Kaggle TPUs in the comparison ?</p>",
      "rawMarkdown": "Why didn't you include Kaggle TPUs in the comparison ?",
      "votes": null
    },
    {
      "id": "880331",
      "postDate": "06/10/2020 07:22:36",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> thank you for reading my kernels! </p>\n\n<p>I really love PyTorch. I wouldn't like to compare with native Tensorflow for TPU. About torch xla:</p>\n\n<p>Kaggle has TPUv3-8 (it is better and faster than Colab TPUv2-8), but training on Kaggle TPU is slower and sometimes impossible! Why? I wrote about \"not enough RAM\" in <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">my previous tpu kernel</a>. Now I would like to add \"not enough CPU\". TPU Kernels of Colab PRO have 36GB RAM and 40 CPU. It allows to use TPU with full power and make faster kaggle TPUv3-8!</p>\n\n<p><code>You should contribute similar CPU and RAM resources for kaggle TPU kernels, thank you in advance!</code></p>\n\n<p>Also I have added <a href=\"https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast?scriptVersionId=35857962\">version 5</a>. You can see exception. Sometimes this exception will appear later, it is really unstable. In Colab TPU I never meet these problems so I have created colab kernels and provide it for other users!</p>\n\n<p>P.S. I have added comparison with Kaggle TPU</p>",
      "rawMarkdown": "mgornergoogle thank you for reading my kernels! \n\nI really love PyTorch. I wouldn't like to compare with native Tensorflow for TPU. About torch xla:\n\nKaggle has TPUv3-8 (it is better and faster than Colab TPUv2-8), but training on Kaggle TPU is slower and sometimes impossible! Why? I wrote about \"not enough RAM\" in [my previous tpu kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta). Now I would like to add \"not enough CPU\". TPU Kernels of Colab PRO have 36GB RAM and 40 CPU. It allows to use TPU with full power and make faster kaggle TPUv3-8!\n\n`You should contribute similar CPU and RAM resources for kaggle TPU kernels, thank you in advance!`\n\nAlso I have added [version 5](https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast?scriptVersionId=35857962). You can see exception. Sometimes this exception will appear later, it is really unstable. In Colab TPU I never meet these problems so I have created colab kernels and provide it for other users!\n\nP.S. I have added comparison with Kaggle TPU",
      "votes": null
    },
    {
      "id": "881450",
      "postDate": "06/11/2020 03:45:02",
      "content": "<p>Thanks so much for the explanation, <a href=\"/shonenkov\">@shonenkov</a>. I'm looking into what we can provide resource wise for TPU notebooks on Kaggle.</p>",
      "rawMarkdown": "Thanks so much for the explanation, @shonenkov. I'm looking into what we can provide resource wise for TPU notebooks on Kaggle.",
      "votes": null
    },
    {
      "id": "883935",
      "postDate": "06/13/2020 05:10:28",
      "content": "<p>+1 for the CPU &amp; RAM resources on TPU Kaggle notebooks. I was super excited to try them out in either PyTorch or PyTorch Lightning, but I've not yet been able to unlock the full potential of a TPU in Kaggle notebook without errors with image data.</p>\n\n<p>I think in the Tweet Sentiment Extraction competition, people had more success with TPUs &amp; PyTorch since preprocessing and loading text is probably less resource intensive.</p>",
      "rawMarkdown": "1 for the CPU &amp; RAM resources on TPU Kaggle notebooks. I was super excited to try them out in either PyTorch or PyTorch Lightning, but I've not yet been able to unlock the full potential of a TPU in Kaggle notebook without errors with image data.\n\nI think in the Tweet Sentiment Extraction competition, people had more success with TPUs &amp; PyTorch since preprocessing and loading text is probably less resource intensive.",
      "votes": null
    },
    {
      "id": "885270",
      "postDate": "06/14/2020 04:28:57",
      "content": "<p>Excellent and very informative post!</p>",
      "rawMarkdown": "Excellent and very informative post!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 879994,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "06/09/2020 22:30:35",
      "content": "<p>Why didn't you include Kaggle TPUs in the comparison ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 880331,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/10/2020 07:22:36",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> thank you for reading my kernels! </p>\n\n<p>I really love PyTorch. I wouldn't like to compare with native Tensorflow for TPU. About torch xla:</p>\n\n<p>Kaggle has TPUv3-8 (it is better and faster than Colab TPUv2-8), but training on Kaggle TPU is slower and sometimes impossible! Why? I wrote about \"not enough RAM\" in <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">my previous tpu kernel</a>. Now I would like to add \"not enough CPU\". TPU Kernels of Colab PRO have 36GB RAM and 40 CPU. It allows to use TPU with full power and make faster kaggle TPUv3-8!</p>\n\n<p><code>You should contribute similar CPU and RAM resources for kaggle TPU kernels, thank you in advance!</code></p>\n\n<p>Also I have added <a href=\"https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast?scriptVersionId=35857962\">version 5</a>. You can see exception. Sometimes this exception will appear later, it is really unstable. In Colab TPU I never meet these problems so I have created colab kernels and provide it for other users!</p>\n\n<p>P.S. I have added comparison with Kaggle TPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 881450,
          "author_name": "mrisdal",
          "author_url": "",
          "post_date": "06/11/2020 03:45:02",
          "content": "<p>Thanks so much for the explanation, <a href=\"/shonenkov\">@shonenkov</a>. I'm looking into what we can provide resource wise for TPU notebooks on Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 883935,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "06/13/2020 05:10:28",
          "content": "<p>+1 for the CPU &amp; RAM resources on TPU Kaggle notebooks. I was super excited to try them out in either PyTorch or PyTorch Lightning, but I've not yet been able to unlock the full potential of a TPU in Kaggle notebook without errors with image data.</p>\n\n<p>I think in the Tweet Sentiment Extraction competition, people had more success with TPUs &amp; PyTorch since preprocessing and loading text is probably less resource intensive.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 885270,
      "author_name": "joydeep1",
      "author_url": "",
      "post_date": "06/14/2020 04:28:57",
      "content": "<p>Excellent and very informative post!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "879264": "Hi everyone!\n\nI have created kernels using TPU with PyTorch models for this competition! \n\n- [[Torch XLA] Melanoma Crazy Fast](https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast)\n- [[Inference] Melanoma Crazy Fast](https://www.kaggle.com/shonenkov/inference-melanoma-crazy-fast)\n\nAlso I have prepared this baseline on Colab! You can make a copy and start research! I recommend to use Colab Pro version with HIGH-RAM mode and TPU Runtime Type. Your speed of research will increase ~5-10 times! Link you can find in training kernel.\n\nInference kernel has some might helpful information about stable prediction and example with calculation OOF scores. Hope it helps you!\n\nWelcome! \n\n### 1. 256x256 image size, resnext50d_32x4d, 1 epoch:\n- Kaggle P100:  ~220s (train), ~60s (validation)\n- Kaggle TPU:  ~75s (train), ~30s (validation) !Be carefull with BUGS\n- Colab TPU:  ~60s (train), ~35s (validation)\n\n### 2. 256x256 image size, efficientnet-b5, 1 epoch:\n- Kaggle P100:  ~350s (train), ~60s (validation)\n- Kaggle TPU:  ~105s (train), ~30s (validation)  !Be carefull with BUGS\n- Colab TPU:  ~50s (train), ~25s (validation)\n\n\nP.S. If you don't understand what effect make these kernels, please, skip!! Please, don't write toxic comments /: it is disheartening. Thank you!",
    "879994": "Why didn't you include Kaggle TPUs in the comparison ?",
    "880331": "mgornergoogle thank you for reading my kernels! \n\nI really love PyTorch. I wouldn't like to compare with native Tensorflow for TPU. About torch xla:\n\nKaggle has TPUv3-8 (it is better and faster than Colab TPUv2-8), but training on Kaggle TPU is slower and sometimes impossible! Why? I wrote about \"not enough RAM\" in [my previous tpu kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta). Now I would like to add \"not enough CPU\". TPU Kernels of Colab PRO have 36GB RAM and 40 CPU. It allows to use TPU with full power and make faster kaggle TPUv3-8!\n\n`You should contribute similar CPU and RAM resources for kaggle TPU kernels, thank you in advance!`\n\nAlso I have added [version 5](https://www.kaggle.com/shonenkov/torch-xla-melanoma-crazy-fast?scriptVersionId=35857962). You can see exception. Sometimes this exception will appear later, it is really unstable. In Colab TPU I never meet these problems so I have created colab kernels and provide it for other users!\n\nP.S. I have added comparison with Kaggle TPU",
    "881450": "Thanks so much for the explanation, @shonenkov. I'm looking into what we can provide resource wise for TPU notebooks on Kaggle.",
    "883935": "1 for the CPU &amp; RAM resources on TPU Kaggle notebooks. I was super excited to try them out in either PyTorch or PyTorch Lightning, but I've not yet been able to unlock the full potential of a TPU in Kaggle notebook without errors with image data.\n\nI think in the Tweet Sentiment Extraction competition, people had more success with TPUs &amp; PyTorch since preprocessing and loading text is probably less resource intensive.",
    "885270": "Excellent and very informative post!"
  },
  "source": "meta"
}