{
  "id": 208366,
  "title": "Training vision transformers with distillation",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/208366",
  "author_name": "Graeme Holliday",
  "post_date": "2021-01-03T06:00:36.545000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I found this paper very interesting: <a href=\"https://arxiv.org/pdf/2012.12877.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.12877.pdf</a><br>\nHowever, I noticed there were not any implementations that easily allowed usage of pretrained models. So, I went ahead and made one!</p>\n<p><a href=\"https://github.com/Graeme22/DistillableVisionTransformer\" target=\"_blank\">https://github.com/Graeme22/DistillableVisionTransformer</a></p>\n<p>I hope this is useful to someone. I haven't tested it for the competition yet but I believe it will make life much easier for those using transformers. Enjoy!</p>",
  "messages": [
    {
      "id": 1136486,
      "postDate": "2021-01-03T06:00:36.547Z",
      "content": "<p>Hi all,</p>\n<p>I found this paper very interesting: <a href=\"https://arxiv.org/pdf/2012.12877.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.12877.pdf</a><br>\nHowever, I noticed there were not any implementations that easily allowed usage of pretrained models. So, I went ahead and made one!</p>\n<p><a href=\"https://github.com/Graeme22/DistillableVisionTransformer\" target=\"_blank\">https://github.com/Graeme22/DistillableVisionTransformer</a></p>\n<p>I hope this is useful to someone. I haven't tested it for the competition yet but I believe it will make life much easier for those using transformers. Enjoy!</p>",
      "rawMarkdown": "Hi all,\n\nI found this paper very interesting: https://arxiv.org/pdf/2012.12877.pdf\nHowever, I noticed there were not any implementations that easily allowed usage of pretrained models. So, I went ahead and made one!\n\nhttps://github.com/Graeme22/DistillableVisionTransformer\n\nI hope this is useful to someone. I haven't tested it for the competition yet but I believe it will make life much easier for those using transformers. Enjoy!",
      "votes": 4
    },
    {
      "id": 1148709,
      "postDate": "2021-01-11T10:31:48.257Z",
      "content": "<p>Sorry, didn't work.<br>\nteacher = ResNext(LB = 0.9), student = ViT_B-16(LB = 0.895)。<br>\nBut the distillation model LB = 0.888.<br>\nSomewhat doubt the paper ------- only tricks😊</p>",
      "rawMarkdown": "Sorry, didn't work.\nteacher = ResNext(LB = 0.9), student = ViT_B-16(LB = 0.895)。\nBut the distillation model LB = 0.888.\nSomewhat doubt the paper ------- only tricks😊",
      "replies": [
        {
          "id": 1149725,
          "postDate": "2021-01-12T04:58:46.120Z",
          "content": "<p>Hi, sorry it didn't provide an improvement for you. As I mentioned, I haven't tested it for this competition, but some things I would try would be larger batch size on TPUs, maybe start the student from scratch instead of pretrained, or mess with the loss hyperparameters. One of the things that is quite important for training transformers is a high volume of data, so perhaps some self-supervised technique on the old competition data could be helpful. Hope that gives you some ideas, good luck!</p>",
          "rawMarkdown": "Hi, sorry it didn't provide an improvement for you. As I mentioned, I haven't tested it for this competition, but some things I would try would be larger batch size on TPUs, maybe start the student from scratch instead of pretrained, or mess with the loss hyperparameters. One of the things that is quite important for training transformers is a high volume of data, so perhaps some self-supervised technique on the old competition data could be helpful. Hope that gives you some ideas, good luck!"
        },
        {
          "id": 1149840,
          "postDate": "2021-01-12T07:33:54.313Z",
          "content": "<p>Totally agree with you about 'start the student from scratch'. It turns out to be a fancy solution.<br>\nFacebook guys save some time, but it is still a lot(300e).<br>\nThank you anyway, the code is great!</p>",
          "rawMarkdown": "Totally agree with you about 'start the student from scratch'. It turns out to be a fancy solution.\nFacebook guys save some time, but it is still a lot(300e).\nThank you anyway, the code is great!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1149724,
      "postDate": "2021-01-12T04:58:13.933Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1148709,
      "author_name": "wangbingo",
      "author_url": "",
      "post_date": "2021-01-11T10:31:48.257000",
      "content": "<p>Sorry, didn't work.<br>\nteacher = ResNext(LB = 0.9), student = ViT_B-16(LB = 0.895)。<br>\nBut the distillation model LB = 0.888.<br>\nSomewhat doubt the paper ------- only tricks😊</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1149725,
          "author_name": "Graeme Holliday",
          "author_url": "",
          "post_date": "2021-01-12T04:58:46.120000",
          "content": "<p>Hi, sorry it didn't provide an improvement for you. As I mentioned, I haven't tested it for this competition, but some things I would try would be larger batch size on TPUs, maybe start the student from scratch instead of pretrained, or mess with the loss hyperparameters. One of the things that is quite important for training transformers is a high volume of data, so perhaps some self-supervised technique on the old competition data could be helpful. Hope that gives you some ideas, good luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1149840,
          "author_name": "wangbingo",
          "author_url": "",
          "post_date": "2021-01-12T07:33:54.313000",
          "content": "<p>Totally agree with you about 'start the student from scratch'. It turns out to be a fancy solution.<br>\nFacebook guys save some time, but it is still a lot(300e).<br>\nThank you anyway, the code is great!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1149724,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-12T04:58:13.933000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1136486": "Hi all,\n\nI found this paper very interesting: https://arxiv.org/pdf/2012.12877.pdf\nHowever, I noticed there were not any implementations that easily allowed usage of pretrained models. So, I went ahead and made one!\n\nhttps://github.com/Graeme22/DistillableVisionTransformer\n\nI hope this is useful to someone. I haven't tested it for the competition yet but I believe it will make life much easier for those using transformers. Enjoy!",
    "1148709": "Sorry, didn't work.\nteacher = ResNext(LB = 0.9), student = ViT_B-16(LB = 0.895)。\nBut the distillation model LB = 0.888.\nSomewhat doubt the paper ------- only tricks😊",
    "1149724": ""
  }
}