{
  "id": 210889,
  "title": "PYTORCH TPU Starter : Try Pytorch XLA if you are using Pytorch for this competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210889",
  "author_name": "",
  "post_date": "2021-01-12T20:34:57.617929100Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>i understand,not everyone has fancy hardware/GPU and i am one of them :)<br>\nfor people like us,TPU is our last hope and if you are a pytorch lover then pytorch xla is all you need :)<br>\nlast time i used pytorch xla almost 1 year ago in alaska 2 competition,it had a lot of issue related to memory and performance back then but now in 2021 pytorch xla has changed a lot,it is much stronger than it was before(even though TF TPU is better at this moment) </p>\n<p>I have started my first #kaggle notebook of this year Using Pytorch XLA as i am huge fan of pytorch  and that too uses this competitions dataset for experiments</p>\n<p>in this notebook <a href=\"https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease/\" target=\"_blank\">ViT - Pytorch xla (TPU) for leaf disease</a> i tried Visual Transformers/ ViT base model from popular <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">pytorch image models</a><br>\nin version 10 of that notebook i was able to achieve 0.88+ validation accuracy on first fold within 10 epoch that took around 1 hour.</p>\n<p>then in this notebook <a href=\"https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease\" target=\"_blank\">Faster torch tpu baseline for leaf disease</a> i tried to convert <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> san's magnificent kernel titled <strong>Baseline - Modified From Previous Competition</strong> which is a very good boilerplate to use in pretty much all future projects. We all know that this competition is very similar to last pandas challenge where we had label noise and in that competition my team used this notebook of qishen ha san called <strong>Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]</strong> and missed silver by only 1 place in private lb. we tried a lot of different things and lot of different models using that baseline of qishen ha san which helped us get into 51st private lb position. </p>\n<p>i believe we can do very fast experiments with this notebook  <a href=\"https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease\" target=\"_blank\">Faster torch tpu baseline for leaf disease</a>, because<br>\nin draft mode i was able to get ~0.89 validation accuracy after 1st epoch on first fold(which took me only 10 minutes to get that result using ViT Large model)</p>\n<p>In pandas competition our cv lb has good correlation and in this competition we can see cv and public lb has some correlation but how much trust we can put on cv for private lb that's a big question i guess! <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2034058%2F55b579924502626e429f70535aa65ee9%2Foverfit.jpg?generation=1610482797925162&amp;alt=media\" alt=\"\"></p>\n<h1>Problem Faced :</h1>\n<p>i spend endless hour debugging a few things on those 2 tpu kernels that i mentioned above,i am still having trouble saving the best weight file of each fold, if i put xm.save() or xser.save() inside epoch loop for saving each epochs weight file or best weight file of each fold then the kernel commit doesn't finish but in interactive mode everything works fine. i have posted about this problem here <a href=\"https://www.kaggle.com/discussion/210032\" target=\"_blank\">Tpu kernel commit not finishing but worked well during interactive session</a><br>\nanyone faced similar problem? thank you in advance</p>",
  "messages": [
    {
      "id": "1150777",
      "postDate": "01/12/2021 20:34:57",
      "content": "<p>i understand,not everyone has fancy hardware/GPU and i am one of them :)<br>\nfor people like us,TPU is our last hope and if you are a pytorch lover then pytorch xla is all you need :)<br>\nlast time i used pytorch xla almost 1 year ago in alaska 2 competition,it had a lot of issue related to memory and performance back then but now in 2021 pytorch xla has changed a lot,it is much stronger than it was before(even though TF TPU is better at this moment) </p>\n<p>I have started my first #kaggle notebook of this year Using Pytorch XLA as i am huge fan of pytorch  and that too uses this competitions dataset for experiments</p>\n<p>in this notebook <a href=\"https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease/\" target=\"_blank\">ViT - Pytorch xla (TPU) for leaf disease</a> i tried Visual Transformers/ ViT base model from popular <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">pytorch image models</a><br>\nin version 10 of that notebook i was able to achieve 0.88+ validation accuracy on first fold within 10 epoch that took around 1 hour.</p>\n<p>then in this notebook <a href=\"https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease\" target=\"_blank\">Faster torch tpu baseline for leaf disease</a> i tried to convert <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> san's magnificent kernel titled <strong>Baseline - Modified From Previous Competition</strong> which is a very good boilerplate to use in pretty much all future projects. We all know that this competition is very similar to last pandas challenge where we had label noise and in that competition my team used this notebook of qishen ha san called <strong>Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]</strong> and missed silver by only 1 place in private lb. we tried a lot of different things and lot of different models using that baseline of qishen ha san which helped us get into 51st private lb position. </p>\n<p>i believe we can do very fast experiments with this notebook  <a href=\"https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease\" target=\"_blank\">Faster torch tpu baseline for leaf disease</a>, because<br>\nin draft mode i was able to get ~0.89 validation accuracy after 1st epoch on first fold(which took me only 10 minutes to get that result using ViT Large model)</p>\n<p>In pandas competition our cv lb has good correlation and in this competition we can see cv and public lb has some correlation but how much trust we can put on cv for private lb that's a big question i guess! <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2034058%2F55b579924502626e429f70535aa65ee9%2Foverfit.jpg?generation=1610482797925162&amp;alt=media\" alt=\"\"></p>\n<h1>Problem Faced :</h1>\n<p>i spend endless hour debugging a few things on those 2 tpu kernels that i mentioned above,i am still having trouble saving the best weight file of each fold, if i put xm.save() or xser.save() inside epoch loop for saving each epochs weight file or best weight file of each fold then the kernel commit doesn't finish but in interactive mode everything works fine. i have posted about this problem here <a href=\"https://www.kaggle.com/discussion/210032\" target=\"_blank\">Tpu kernel commit not finishing but worked well during interactive session</a><br>\nanyone faced similar problem? thank you in advance</p>",
      "rawMarkdown": "i understand,not everyone has fancy hardware/GPU and i am one of them :)\nfor people like us,TPU is our last hope and if you are a pytorch lover then pytorch xla is all you need :)\nlast time i used pytorch xla almost 1 year ago in alaska 2 competition,it had a lot of issue related to memory and performance back then but now in 2021 pytorch xla has changed a lot,it is much stronger than it was before(even though TF TPU is better at this moment) \n\nI have started my first #kaggle notebook of this year Using Pytorch XLA as i am huge fan of pytorch  and that too uses this competitions dataset for experiments\n\nin this notebook [ViT - Pytorch xla (TPU) for leaf disease](https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease/) i tried Visual Transformers/ ViT base model from popular [pytorch image models](https://github.com/rwightman/pytorch-image-models)\nin version 10 of that notebook i was able to achieve 0.88+ validation accuracy on first fold within 10 epoch that took around 1 hour.\n\nthen in this notebook [Faster torch tpu baseline for leaf disease](https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease) i tried to convert @haqishen san's magnificent kernel titled **Baseline - Modified From Previous Competition** which is a very good boilerplate to use in pretty much all future projects. We all know that this competition is very similar to last pandas challenge where we had label noise and in that competition my team used this notebook of qishen ha san called **Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]** and missed silver by only 1 place in private lb. we tried a lot of different things and lot of different models using that baseline of qishen ha san which helped us get into 51st private lb position. \n\ni believe we can do very fast experiments with this notebook  [Faster torch tpu baseline for leaf disease](https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease), because\nin draft mode i was able to get ~0.89 validation accuracy after 1st epoch on first fold(which took me only 10 minutes to get that result using ViT Large model)\n\nIn pandas competition our cv lb has good correlation and in this competition we can see cv and public lb has some correlation but how much trust we can put on cv for private lb that's a big question i guess! \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2034058%2F55b579924502626e429f70535aa65ee9%2Foverfit.jpg?generation=1610482797925162&alt=media)\n\n# Problem Faced : \ni spend endless hour debugging a few things on those 2 tpu kernels that i mentioned above,i am still having trouble saving the best weight file of each fold, if i put xm.save() or xser.save() inside epoch loop for saving each epochs weight file or best weight file of each fold then the kernel commit doesn't finish but in interactive mode everything works fine. i have posted about this problem here [Tpu kernel commit not finishing but worked well during interactive session](https://www.kaggle.com/discussion/210032)\nanyone faced similar problem? thank you in advance",
      "votes": null
    },
    {
      "id": "1154395",
      "postDate": "01/15/2021 16:06:30",
      "content": "<p>Thank you for the notebook!!</p>",
      "rawMarkdown": "Thank you for the notebook!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1154395,
      "author_name": "anku5hk",
      "author_url": "",
      "post_date": "01/15/2021 16:06:30",
      "content": "<p>Thank you for the notebook!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1150777": "i understand,not everyone has fancy hardware/GPU and i am one of them :)\nfor people like us,TPU is our last hope and if you are a pytorch lover then pytorch xla is all you need :)\nlast time i used pytorch xla almost 1 year ago in alaska 2 competition,it had a lot of issue related to memory and performance back then but now in 2021 pytorch xla has changed a lot,it is much stronger than it was before(even though TF TPU is better at this moment) \n\nI have started my first #kaggle notebook of this year Using Pytorch XLA as i am huge fan of pytorch  and that too uses this competitions dataset for experiments\n\nin this notebook [ViT - Pytorch xla (TPU) for leaf disease](https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease/) i tried Visual Transformers/ ViT base model from popular [pytorch image models](https://github.com/rwightman/pytorch-image-models)\nin version 10 of that notebook i was able to achieve 0.88+ validation accuracy on first fold within 10 epoch that took around 1 hour.\n\nthen in this notebook [Faster torch tpu baseline for leaf disease](https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease) i tried to convert @haqishen san's magnificent kernel titled **Baseline - Modified From Previous Competition** which is a very good boilerplate to use in pretty much all future projects. We all know that this competition is very similar to last pandas challenge where we had label noise and in that competition my team used this notebook of qishen ha san called **Train EfficientNet-B0 w/ 36 tiles_256 [LB0.87]** and missed silver by only 1 place in private lb. we tried a lot of different things and lot of different models using that baseline of qishen ha san which helped us get into 51st private lb position. \n\ni believe we can do very fast experiments with this notebook  [Faster torch tpu baseline for leaf disease](https://www.kaggle.com/mobassir/faster-torch-tpu-baseline-for-leaf-disease), because\nin draft mode i was able to get ~0.89 validation accuracy after 1st epoch on first fold(which took me only 10 minutes to get that result using ViT Large model)\n\nIn pandas competition our cv lb has good correlation and in this competition we can see cv and public lb has some correlation but how much trust we can put on cv for private lb that's a big question i guess! \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2034058%2F55b579924502626e429f70535aa65ee9%2Foverfit.jpg?generation=1610482797925162&alt=media)\n\n# Problem Faced : \ni spend endless hour debugging a few things on those 2 tpu kernels that i mentioned above,i am still having trouble saving the best weight file of each fold, if i put xm.save() or xser.save() inside epoch loop for saving each epochs weight file or best weight file of each fold then the kernel commit doesn't finish but in interactive mode everything works fine. i have posted about this problem here [Tpu kernel commit not finishing but worked well during interactive session](https://www.kaggle.com/discussion/210032)\nanyone faced similar problem? thank you in advance",
    "1154395": "Thank you for the notebook!!"
  },
  "source": "meta"
}