{
  "id": 210293,
  "title": "DeiT training notebook using TPU for Pytorch (with RandAugment)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210293",
  "author_name": "",
  "post_date": "2021-01-10T09:57:30.892551300Z",
  "votes": 28,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>I was inspired by <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">the discussion</a> of other competitions. And I wanted to share <code>DeiT</code> and <code>ViT</code> training code. (I had a problem with my personal GPU, so I also needed TPU code. :) )</p>\n<h1>My training notebook has following:</h1>\n<h3>Models - 4</h3>\n<ul>\n<li><p>DeiT (Data-efficient Image Transformers)<br>\n[<a href=\"https://arxiv.org/abs/2012.12877?fbclid=IwAR2txNeK7FDnm-jRgQV8IrRitr59MJRBi7UxwtA3-R4N77fkVRzNJw2Nkm4\" target=\"_blank\">paper</a>]</p></li>\n<li><p>ViT (Vision Transformer)<br>\n[<a href=\"https://arxiv.org/abs/2010.11929\" target=\"_blank\">paper</a>]</p></li>\n<li><p>Resnext50_32x4d</p></li>\n<li><p>EfficientNet</p></li>\n</ul>\n<h3>Various Loss Functions - 7</h3>\n<ul>\n<li>Updated version - I added <code>TaylorCrossEntropyLoss</code> to original loss function <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208498\" target=\"_blank\">list</a>.</li>\n</ul>\n<h3>Transforms</h3>\n<ul>\n<li>Updated version - <code>RandAugment</code><br>\n(If you want to know how to use AutoAugment, it will help.)</li>\n</ul>\n<h1>End</h1>\n<p>Notebook is <a href=\"https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992\" target=\"_blank\">here</a>.<br>\n<a href=\"https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992\" target=\"_blank\">https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992</a></p>\n<p>I hope it helps.</p>\n<p>p.s. When I write it all down, it doesn't seem like only a DeiT training notebook.</p>",
  "messages": [
    {
      "id": "1147130",
      "postDate": "01/10/2021 09:57:30",
      "content": "<h1>Introduction</h1>\n<p>I was inspired by <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">the discussion</a> of other competitions. And I wanted to share <code>DeiT</code> and <code>ViT</code> training code. (I had a problem with my personal GPU, so I also needed TPU code. :) )</p>\n<h1>My training notebook has following:</h1>\n<h3>Models - 4</h3>\n<ul>\n<li><p>DeiT (Data-efficient Image Transformers)<br>\n[<a href=\"https://arxiv.org/abs/2012.12877?fbclid=IwAR2txNeK7FDnm-jRgQV8IrRitr59MJRBi7UxwtA3-R4N77fkVRzNJw2Nkm4\" target=\"_blank\">paper</a>]</p></li>\n<li><p>ViT (Vision Transformer)<br>\n[<a href=\"https://arxiv.org/abs/2010.11929\" target=\"_blank\">paper</a>]</p></li>\n<li><p>Resnext50_32x4d</p></li>\n<li><p>EfficientNet</p></li>\n</ul>\n<h3>Various Loss Functions - 7</h3>\n<ul>\n<li>Updated version - I added <code>TaylorCrossEntropyLoss</code> to original loss function <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208498\" target=\"_blank\">list</a>.</li>\n</ul>\n<h3>Transforms</h3>\n<ul>\n<li>Updated version - <code>RandAugment</code><br>\n(If you want to know how to use AutoAugment, it will help.)</li>\n</ul>\n<h1>End</h1>\n<p>Notebook is <a href=\"https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992\" target=\"_blank\">here</a>.<br>\n<a href=\"https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992\" target=\"_blank\">https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992</a></p>\n<p>I hope it helps.</p>\n<p>p.s. When I write it all down, it doesn't seem like only a DeiT training notebook.</p>",
      "rawMarkdown": "# Introduction\nI was inspired by [the discussion](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577) of other competitions. And I wanted to share `DeiT` and `ViT` training code. (I had a problem with my personal GPU, so I also needed TPU code. :) )\n\n# My training notebook has following:\n### Models - 4\n- DeiT (Data-efficient Image Transformers)\n[[paper](https://arxiv.org/abs/2012.12877?fbclid=IwAR2txNeK7FDnm-jRgQV8IrRitr59MJRBi7UxwtA3-R4N77fkVRzNJw2Nkm4)]\n\n- ViT (Vision Transformer)\n[[paper](https://arxiv.org/abs/2010.11929)]\n\n- Resnext50_32x4d\n\n- EfficientNet\n\n### Various Loss Functions - 7\n- Updated version - I added `TaylorCrossEntropyLoss` to original loss function [list](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208498).\n\n### Transforms\n- Updated version - `RandAugment`\n  (If you want to know how to use AutoAugment, it will help.)\n\n# End\n\nNotebook is [here](https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992).\nhttps://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992\n\nI hope it helps.\n\np.s. When I write it all down, it doesn't seem like only a DeiT training notebook.",
      "votes": null
    },
    {
      "id": "1150920",
      "postDate": "01/13/2021 00:55:53",
      "content": "<p>You are absolutely amazing👍👍<br>\nHow's your validation performance on DeiT and ViT?  </p>",
      "rawMarkdown": "You are absolutely amazing👍👍\nHow's your validation performance on DeiT and ViT?",
      "votes": null
    },
    {
      "id": "1150971",
      "postDate": "01/13/2021 02:58:02",
      "content": "<p>Thank you!  It's a little lower score compared to my best single model. It seems that it needs to be adjusted a little more for ViT and DeiT in my case. <a href=\"https://www.kaggle.com/billbafare\" target=\"_blank\">@billbafare</a> </p>",
      "rawMarkdown": "Thank you!  It's a little lower score compared to my best single model. It seems that it needs to be adjusted a little more for ViT and DeiT in my case. @billbafare",
      "votes": null
    },
    {
      "id": "1151219",
      "postDate": "01/13/2021 07:35:31",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>,<br>\nWhen you are training with TPU instead of GPU are you getting similar results. When I'm training on tpu I'm getting worser results than when when I trained same model on GPU.</p>",
      "rawMarkdown": "Hello @piantic,\nWhen you are training with TPU instead of GPU are you getting similar results. When I'm training on tpu I'm getting worser results than when when I trained same model on GPU.",
      "votes": null
    },
    {
      "id": "1151229",
      "postDate": "01/13/2021 07:44:24",
      "content": "<p>Good points, the paper mentioned that ViT and DeiT are very sensitive for initial hyperparameters. So, we have to set different CFG setting for TPU or GPU. <a href=\"https://www.kaggle.com/filchy\" target=\"_blank\">@filchy</a> </p>",
      "rawMarkdown": "Good points, the paper mentioned that ViT and DeiT are very sensitive for initial hyperparameters. So, we have to set different CFG setting for TPU or GPU. @filchy",
      "votes": null
    },
    {
      "id": "1151961",
      "postDate": "01/13/2021 17:06:34",
      "content": "<p>Amazing, thank you very much for your answer.</p>",
      "rawMarkdown": "Amazing, thank you very much for your answer.",
      "votes": null
    },
    {
      "id": "1154429",
      "postDate": "01/15/2021 16:21:41",
      "content": "<p>Try training with larger Batch_size(so maybe consider a bit bigger lr) on TPUs, it should give results.</p>",
      "rawMarkdown": "Try training with larger Batch_size(so maybe consider a bit bigger lr) on TPUs, it should give results.",
      "votes": null
    },
    {
      "id": "1155768",
      "postDate": "01/16/2021 16:54:49",
      "content": "<p><a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>  Thank you so for this amazing notebook and for sharing it with us. Which image size did you use? Is it 512*512 or any others? </p>",
      "rawMarkdown": "piantic  Thank you so for this amazing notebook and for sharing it with us. Which image size did you use? Is it 512*512 or any others?",
      "votes": null
    },
    {
      "id": "1155792",
      "postDate": "01/16/2021 17:18:12",
      "content": "<p>I used 224x224 for DeiT and for other models used 384x384 ~ 512x512.<br>\n<a href=\"https://www.kaggle.com/durbin164\" target=\"_blank\">@durbin164</a> </p>",
      "rawMarkdown": "I used 224x224 for DeiT and for other models used 384x384 ~ 512x512.\n@durbin164",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1150920,
      "author_name": "billbafare",
      "author_url": "",
      "post_date": "01/13/2021 00:55:53",
      "content": "<p>You are absolutely amazing👍👍<br>\nHow's your validation performance on DeiT and ViT?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1150971,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "01/13/2021 02:58:02",
          "content": "<p>Thank you!  It's a little lower score compared to my best single model. It seems that it needs to be adjusted a little more for ViT and DeiT in my case. <a href=\"https://www.kaggle.com/billbafare\" target=\"_blank\">@billbafare</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1151219,
      "author_name": "filchy",
      "author_url": "",
      "post_date": "01/13/2021 07:35:31",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>,<br>\nWhen you are training with TPU instead of GPU are you getting similar results. When I'm training on tpu I'm getting worser results than when when I trained same model on GPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1151229,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "01/13/2021 07:44:24",
          "content": "<p>Good points, the paper mentioned that ViT and DeiT are very sensitive for initial hyperparameters. So, we have to set different CFG setting for TPU or GPU. <a href=\"https://www.kaggle.com/filchy\" target=\"_blank\">@filchy</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151961,
          "author_name": "filchy",
          "author_url": "",
          "post_date": "01/13/2021 17:06:34",
          "content": "<p>Amazing, thank you very much for your answer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1154429,
          "author_name": "anku5hk",
          "author_url": "",
          "post_date": "01/15/2021 16:21:41",
          "content": "<p>Try training with larger Batch_size(so maybe consider a bit bigger lr) on TPUs, it should give results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1155768,
      "author_name": "durbin164",
      "author_url": "",
      "post_date": "01/16/2021 16:54:49",
      "content": "<p><a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>  Thank you so for this amazing notebook and for sharing it with us. Which image size did you use? Is it 512*512 or any others? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1155792,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "01/16/2021 17:18:12",
          "content": "<p>I used 224x224 for DeiT and for other models used 384x384 ~ 512x512.<br>\n<a href=\"https://www.kaggle.com/durbin164\" target=\"_blank\">@durbin164</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1147130": "# Introduction\nI was inspired by [the discussion](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577) of other competitions. And I wanted to share `DeiT` and `ViT` training code. (I had a problem with my personal GPU, so I also needed TPU code. :) )\n\n# My training notebook has following:\n### Models - 4\n- DeiT (Data-efficient Image Transformers)\n[[paper](https://arxiv.org/abs/2012.12877?fbclid=IwAR2txNeK7FDnm-jRgQV8IrRitr59MJRBi7UxwtA3-R4N77fkVRzNJw2Nkm4)]\n\n- ViT (Vision Transformer)\n[[paper](https://arxiv.org/abs/2010.11929)]\n\n- Resnext50_32x4d\n\n- EfficientNet\n\n### Various Loss Functions - 7\n- Updated version - I added `TaylorCrossEntropyLoss` to original loss function [list](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208498).\n\n### Transforms\n- Updated version - `RandAugment`\n  (If you want to know how to use AutoAugment, it will help.)\n\n# End\n\nNotebook is [here](https://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992).\nhttps://www.kaggle.com/piantic/cnn-or-transformer-pytorch-xla-tpu-for-cassava?scriptVersionId=51538992\n\nI hope it helps.\n\np.s. When I write it all down, it doesn't seem like only a DeiT training notebook.",
    "1150920": "You are absolutely amazing👍👍\nHow's your validation performance on DeiT and ViT?",
    "1150971": "Thank you!  It's a little lower score compared to my best single model. It seems that it needs to be adjusted a little more for ViT and DeiT in my case. @billbafare",
    "1151219": "Hello @piantic,\nWhen you are training with TPU instead of GPU are you getting similar results. When I'm training on tpu I'm getting worser results than when when I trained same model on GPU.",
    "1151229": "Good points, the paper mentioned that ViT and DeiT are very sensitive for initial hyperparameters. So, we have to set different CFG setting for TPU or GPU. @filchy",
    "1151961": "Amazing, thank you very much for your answer.",
    "1154429": "Try training with larger Batch_size(so maybe consider a bit bigger lr) on TPUs, it should give results.",
    "1155768": "piantic  Thank you so for this amazing notebook and for sharing it with us. Which image size did you use? Is it 512*512 or any others?",
    "1155792": "I used 224x224 for DeiT and for other models used 384x384 ~ 512x512.\n@durbin164"
  },
  "source": "meta"
}