{
  "id": 206441,
  "title": "Data-efficient image Transformers (DeiT) - New model by FAIR",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/206441",
  "author_name": "xhlulu",
  "post_date": "2020-12-24T16:33:39.975000",
  "votes": 19,
  "comment_count": 10,
  "views": 0,
  "content": "<p>New paper by Facebook AI Research (FAIR) showing great results for Vision Transformer by only training on 1.2M ImageNet images.</p>\n<ul>\n<li><a href=\"https://ai.facebook.com/blog/data-efficient-image-transformers-a-promising-new-technique-for-image-classification/\" target=\"_blank\">Blog post</a></li>\n<li><a href=\"https://arxiv.org/abs/2012.12877\" target=\"_blank\">Arxiv paper</a></li>\n<li><a href=\"https://github.com/facebookresearch/deit\" target=\"_blank\">Code</a></li>\n</ul>\n<p>Some benchmark on ImageNet:<br>\n<img src=\"https://raw.githubusercontent.com/facebookresearch/deit/main/.github/deit.png\" alt=\"\"></p>\n<p>You can easily install it through torch hub:</p>\n<pre><code>import torch\n# check you have the right version of timm\nimport timm\nassert timm.__version__ == \"0.3.2\"\n\n# now load it with torchhub\nmodel = torch.hub.load(\n    'facebookresearch/deit:main', 'deit_base_patch16_224', \n    pretrained=True\n)\n</code></pre>\n<p>I'm cautiously optimistic about this, As <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> mentioned in <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206345#1124900\" target=\"_blank\">another post</a>:</p>\n<blockquote>\n  <p>Be Carefull when they say they beat EfficientNet . In general, they don't really respect the scaling and training processes in the original EfficientNet paper when doing the comparison. </p>\n  <p>At Facebook, they have already claimed beating Effnet With the Regnet Paper.  But all those who retrained the models (like Ross Wightman with timm) show that the latter is way far from beating Effnet. </p>\n  <p>And if you use pretrained regnet for fine-tuning , you'll see it performs in general poorly than efficientnet, particularly with higher resulutions than 224x224</p>\n</blockquote>\n<p>But it's worth trying out since the code and trained weights are all available!</p>\n<p>Here's an image retrieved from the blog post:</p>\n<p><img src=\"https://scontent.fymy1-2.fna.fbcdn.net/v/t39.2365-6/131568808_300918484690143_5468311438414224765_n.png?_nc_cat=101&amp;ccb=2&amp;_nc_sid=ad8a9d&amp;_nc_ohc=0Bm6ISoMSa4AX8pW3qU&amp;_nc_ht=scontent.fymy1-2.fna&amp;oh=08d7f1568484e12bc4c2a24e824ae0e0&amp;oe=600A4F7A\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1125339,
      "postDate": "2020-12-24T16:33:39.977Z",
      "content": "<p>New paper by Facebook AI Research (FAIR) showing great results for Vision Transformer by only training on 1.2M ImageNet images.</p>\n<ul>\n<li><a href=\"https://ai.facebook.com/blog/data-efficient-image-transformers-a-promising-new-technique-for-image-classification/\" target=\"_blank\">Blog post</a></li>\n<li><a href=\"https://arxiv.org/abs/2012.12877\" target=\"_blank\">Arxiv paper</a></li>\n<li><a href=\"https://github.com/facebookresearch/deit\" target=\"_blank\">Code</a></li>\n</ul>\n<p>Some benchmark on ImageNet:<br>\n<img src=\"https://raw.githubusercontent.com/facebookresearch/deit/main/.github/deit.png\" alt=\"\"></p>\n<p>You can easily install it through torch hub:</p>\n<pre><code>import torch\n# check you have the right version of timm\nimport timm\nassert timm.__version__ == \"0.3.2\"\n\n# now load it with torchhub\nmodel = torch.hub.load(\n    'facebookresearch/deit:main', 'deit_base_patch16_224', \n    pretrained=True\n)\n</code></pre>\n<p>I'm cautiously optimistic about this, As <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> mentioned in <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206345#1124900\" target=\"_blank\">another post</a>:</p>\n<blockquote>\n  <p>Be Carefull when they say they beat EfficientNet . In general, they don't really respect the scaling and training processes in the original EfficientNet paper when doing the comparison. </p>\n  <p>At Facebook, they have already claimed beating Effnet With the Regnet Paper.  But all those who retrained the models (like Ross Wightman with timm) show that the latter is way far from beating Effnet. </p>\n  <p>And if you use pretrained regnet for fine-tuning , you'll see it performs in general poorly than efficientnet, particularly with higher resulutions than 224x224</p>\n</blockquote>\n<p>But it's worth trying out since the code and trained weights are all available!</p>\n<p>Here's an image retrieved from the blog post:</p>\n<p><img src=\"https://scontent.fymy1-2.fna.fbcdn.net/v/t39.2365-6/131568808_300918484690143_5468311438414224765_n.png?_nc_cat=101&amp;ccb=2&amp;_nc_sid=ad8a9d&amp;_nc_ohc=0Bm6ISoMSa4AX8pW3qU&amp;_nc_ht=scontent.fymy1-2.fna&amp;oh=08d7f1568484e12bc4c2a24e824ae0e0&amp;oe=600A4F7A\" alt=\"\"></p>",
      "rawMarkdown": "New paper by Facebook AI Research (FAIR) showing great results for Vision Transformer by only training on 1.2M ImageNet images.\n\n* [Blog post](https://ai.facebook.com/blog/data-efficient-image-transformers-a-promising-new-technique-for-image-classification/)\n* [Arxiv paper](https://arxiv.org/abs/2012.12877)\n* [Code](https://github.com/facebookresearch/deit)\n\nSome benchmark on ImageNet:\n![](https://raw.githubusercontent.com/facebookresearch/deit/main/.github/deit.png)\n\nYou can easily install it through torch hub:\n```\nimport torch\n# check you have the right version of timm\nimport timm\nassert timm.__version__ == \"0.3.2\"\n\n# now load it with torchhub\nmodel = torch.hub.load(\n    'facebookresearch/deit:main', 'deit_base_patch16_224', \n    pretrained=True\n)\n```\n\nI'm cautiously optimistic about this, As @serigne mentioned in [another post](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206345#1124900):\n> Be Carefull when they say they beat EfficientNet . In general, they don't really respect the scaling and training processes in the original EfficientNet paper when doing the comparison. \n> \n> At Facebook, they have already claimed beating Effnet With the Regnet Paper.  But all those who retrained the models (like Ross Wightman with timm) show that the latter is way far from beating Effnet. \n> \n> And if you use pretrained regnet for fine-tuning , you'll see it performs in general poorly than efficientnet, particularly with higher resulutions than 224x224\n\nBut it's worth trying out since the code and trained weights are all available!\n\nHere's an image retrieved from the blog post:\n\n![](https://scontent.fymy1-2.fna.fbcdn.net/v/t39.2365-6/131568808_300918484690143_5468311438414224765_n.png?_nc_cat=101&ccb=2&_nc_sid=ad8a9d&_nc_ohc=0Bm6ISoMSa4AX8pW3qU&_nc_ht=scontent.fymy1-2.fna&oh=08d7f1568484e12bc4c2a24e824ae0e0&oe=600A4F7A)\n",
      "votes": 18
    },
    {
      "id": 1125677,
      "postDate": "2020-12-25T00:15:37.783Z",
      "content": "<p>i will try this.</p>\n<p>i tried Vit before, but couldn't make it work. </p>\n<p>note that efficientnet has inbuilt drop connect. this is one of the reasons why its results is more robust. (or easier for poelp to make it work)</p>",
      "rawMarkdown": "i will try this.\n\ni tried Vit before, but couldn't make it work. \n\nnote that efficientnet has inbuilt drop connect. this is one of the reasons why its results is more robust. (or easier for poelp to make it work)",
      "votes": 4,
      "replies": [
        {
          "id": 1125686,
          "postDate": "2020-12-25T00:44:28.147Z",
          "content": "<p>Funny anecdote about the drop connect: for this competition I was using the qubvel/callidor implementation (the one on <code>keras-applications</code>) but realized that the drop connect rate was incorrect. So I switched to the official tf.keras implementation which has the correct rate, but the performance decrease drastically for larger Effnets (even though I tried various values of <code>drop_connect_rate</code>)</p>",
          "rawMarkdown": "Funny anecdote about the drop connect: for this competition I was using the qubvel/callidor implementation (the one on `keras-applications`) but realized that the drop connect rate was incorrect. So I switched to the official tf.keras implementation which has the correct rate, but the performance decrease drastically for larger Effnets (even though I tried various values of `drop_connect_rate`)",
          "votes": 2
        },
        {
          "id": 1126637,
          "postDate": "2020-12-25T18:37:42.860Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1126038,
      "postDate": "2020-12-25T09:31:53.557Z",
      "content": "<p>Tried it. Cannot break 0.92 :/</p>",
      "rawMarkdown": "Tried it. Cannot break 0.92 :/",
      "votes": 1
    },
    {
      "id": 1127658,
      "postDate": "2020-12-26T18:05:48.443Z",
      "content": "<p>I think self-Attention will replace CNN just like it replaced RNN. But for now, it's not mature enough to outperfom efficiently approaches like Efficientnet.</p>\n<p>Beside of that , I don't like Facebook's guys attitude who copied and paste huge part of <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a> timm , without <a href=\"https://twitter.com/wightmanr/status/1342195464880320512\" target=\"_blank\">crediting him properly</a>  and while putting all the copied code in a <a href=\"https://twitter.com/wightmanr/status/1342143859380219906\" target=\"_blank\">more restrictive license</a></p>",
      "rawMarkdown": "I think self-Attention will replace CNN just like it replaced RNN. But for now, it's not mature enough to outperfom efficiently approaches like Efficientnet.\n\nBeside of that , I don't like Facebook's guys attitude who copied and paste huge part of @rwightman timm , without [crediting him properly](https://twitter.com/wightmanr/status/1342195464880320512)  and while putting all the copied code in a [more restrictive license](https://twitter.com/wightmanr/status/1342143859380219906)",
      "votes": 2
    },
    {
      "id": 1127969,
      "postDate": "2020-12-27T04:33:43.757Z",
      "content": "<p>emmmm，maybe they just used the tricks partial to their model in the evaluation.<br>\nAnyway,it is recognized that efficientnet have a good performance on the majority of task related to image.</p>",
      "rawMarkdown": "emmmm，maybe they just used the tricks partial to their model in the evaluation.\nAnyway,it is recognized that efficientnet have a good performance on the majority of task related to image."
    },
    {
      "id": 1126351,
      "postDate": "2020-12-25T14:38:23.017Z",
      "content": "<p>Here is an easy wrapper for checking out this model in PyTorch<br>\n<a href=\"https://github.com/lucidrains/vit-pytorch\" target=\"_blank\">https://github.com/lucidrains/vit-pytorch</a></p>",
      "rawMarkdown": "Here is an easy wrapper for checking out this model in PyTorch\nhttps://github.com/lucidrains/vit-pytorch"
    },
    {
      "id": 1125879,
      "postDate": "2020-12-25T06:39:00.980Z",
      "content": "<p>Yeah, I just saw this on <a href=\"https://twitter.com/paperswithcode/status/1342100539761434624?s=20\" target=\"_blank\">@paperswithcode</a>'s Twitter Feed. First Transformers replaced RNN's now they are coming after CNN's too 😄. Also, was surprised to see the code in Pytorch most transformer paper nowadays use JAX based ecosystems such as <a href=\"https://github.com/google/trax\" target=\"_blank\">google/Trax</a> or <a href=\"https://github.com/google/flax\" target=\"_blank\">google/Flax</a>. Excited to see more transformers in Kaggle competitions.</p>",
      "rawMarkdown": "Yeah, I just saw this on [@paperswithcode](https://twitter.com/paperswithcode/status/1342100539761434624?s=20)'s Twitter Feed. First Transformers replaced RNN's now they are coming after CNN's too 😄. Also, was surprised to see the code in Pytorch most transformer paper nowadays use JAX based ecosystems such as [google/Trax](https://github.com/google/trax) or [google/Flax](https://github.com/google/flax). Excited to see more transformers in Kaggle competitions.",
      "replies": [
        {
          "id": 1127662,
          "postDate": "2020-12-26T18:11:26.260Z",
          "content": "<p>JAX is from Google. The authors of this paper work at Facebook.  So it's easier/suitable for them to copy the already implemented ViT in Pytorch by <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a>  as basis of their work^^</p>",
          "rawMarkdown": "JAX is from Google. The authors of this paper work at Facebook.  So it's easier/suitable for them to copy the already implemented ViT in Pytorch by @rwightman  as basis of their work^^"
        }
      ]
    },
    {
      "id": 1125344,
      "postDate": "2020-12-24T16:34:23.753Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thoughts on this?</p>",
      "rawMarkdown": "@hengck23 Thoughts on this?"
    }
  ],
  "comments": [
    {
      "id": 1125677,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-25T00:15:37.783000",
      "content": "<p>i will try this.</p>\n<p>i tried Vit before, but couldn't make it work. </p>\n<p>note that efficientnet has inbuilt drop connect. this is one of the reasons why its results is more robust. (or easier for poelp to make it work)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1125686,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-25T00:44:28.147000",
          "content": "<p>Funny anecdote about the drop connect: for this competition I was using the qubvel/callidor implementation (the one on <code>keras-applications</code>) but realized that the drop connect rate was incorrect. So I switched to the official tf.keras implementation which has the correct rate, but the performance decrease drastically for larger Effnets (even though I tried various values of <code>drop_connect_rate</code>)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1126637,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-25T18:37:42.860000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1126038,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-12-25T09:31:53.557000",
      "content": "<p>Tried it. Cannot break 0.92 :/</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1127658,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-12-26T18:05:48.443000",
      "content": "<p>I think self-Attention will replace CNN just like it replaced RNN. But for now, it's not mature enough to outperfom efficiently approaches like Efficientnet.</p>\n<p>Beside of that , I don't like Facebook's guys attitude who copied and paste huge part of <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a> timm , without <a href=\"https://twitter.com/wightmanr/status/1342195464880320512\" target=\"_blank\">crediting him properly</a>  and while putting all the copied code in a <a href=\"https://twitter.com/wightmanr/status/1342143859380219906\" target=\"_blank\">more restrictive license</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1127969,
      "author_name": "Xu Zhiyuan",
      "author_url": "",
      "post_date": "2020-12-27T04:33:43.757000",
      "content": "<p>emmmm，maybe they just used the tricks partial to their model in the evaluation.<br>\nAnyway,it is recognized that efficientnet have a good performance on the majority of task related to image.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1126351,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-12-25T14:38:23.017000",
      "content": "<p>Here is an easy wrapper for checking out this model in PyTorch<br>\n<a href=\"https://github.com/lucidrains/vit-pytorch\" target=\"_blank\">https://github.com/lucidrains/vit-pytorch</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1125879,
      "author_name": "Saurav Maheshkar ☕️",
      "author_url": "",
      "post_date": "2020-12-25T06:39:00.980000",
      "content": "<p>Yeah, I just saw this on <a href=\"https://twitter.com/paperswithcode/status/1342100539761434624?s=20\" target=\"_blank\">@paperswithcode</a>'s Twitter Feed. First Transformers replaced RNN's now they are coming after CNN's too 😄. Also, was surprised to see the code in Pytorch most transformer paper nowadays use JAX based ecosystems such as <a href=\"https://github.com/google/trax\" target=\"_blank\">google/Trax</a> or <a href=\"https://github.com/google/flax\" target=\"_blank\">google/Flax</a>. Excited to see more transformers in Kaggle competitions.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1127662,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-12-26T18:11:26.260000",
          "content": "<p>JAX is from Google. The authors of this paper work at Facebook.  So it's easier/suitable for them to copy the already implemented ViT in Pytorch by <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a>  as basis of their work^^</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1125344,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "2020-12-24T16:34:23.753000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thoughts on this?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125339": "New paper by Facebook AI Research (FAIR) showing great results for Vision Transformer by only training on 1.2M ImageNet images.\n\n* [Blog post](https://ai.facebook.com/blog/data-efficient-image-transformers-a-promising-new-technique-for-image-classification/)\n* [Arxiv paper](https://arxiv.org/abs/2012.12877)\n* [Code](https://github.com/facebookresearch/deit)\n\nSome benchmark on ImageNet:\n![](https://raw.githubusercontent.com/facebookresearch/deit/main/.github/deit.png)\n\nYou can easily install it through torch hub:\n```\nimport torch\n# check you have the right version of timm\nimport timm\nassert timm.__version__ == \"0.3.2\"\n\n# now load it with torchhub\nmodel = torch.hub.load(\n    'facebookresearch/deit:main', 'deit_base_patch16_224', \n    pretrained=True\n)\n```\n\nI'm cautiously optimistic about this, As @serigne mentioned in [another post](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206345#1124900):\n> Be Carefull when they say they beat EfficientNet . In general, they don't really respect the scaling and training processes in the original EfficientNet paper when doing the comparison. \n> \n> At Facebook, they have already claimed beating Effnet With the Regnet Paper.  But all those who retrained the models (like Ross Wightman with timm) show that the latter is way far from beating Effnet. \n> \n> And if you use pretrained regnet for fine-tuning , you'll see it performs in general poorly than efficientnet, particularly with higher resulutions than 224x224\n\nBut it's worth trying out since the code and trained weights are all available!\n\nHere's an image retrieved from the blog post:\n\n![](https://scontent.fymy1-2.fna.fbcdn.net/v/t39.2365-6/131568808_300918484690143_5468311438414224765_n.png?_nc_cat=101&ccb=2&_nc_sid=ad8a9d&_nc_ohc=0Bm6ISoMSa4AX8pW3qU&_nc_ht=scontent.fymy1-2.fna&oh=08d7f1568484e12bc4c2a24e824ae0e0&oe=600A4F7A)\n",
    "1125677": "i will try this.\n\ni tried Vit before, but couldn't make it work. \n\nnote that efficientnet has inbuilt drop connect. this is one of the reasons why its results is more robust. (or easier for poelp to make it work)",
    "1126038": "Tried it. Cannot break 0.92 :/",
    "1127658": "I think self-Attention will replace CNN just like it replaced RNN. But for now, it's not mature enough to outperfom efficiently approaches like Efficientnet.\n\nBeside of that , I don't like Facebook's guys attitude who copied and paste huge part of @rwightman timm , without [crediting him properly](https://twitter.com/wightmanr/status/1342195464880320512)  and while putting all the copied code in a [more restrictive license](https://twitter.com/wightmanr/status/1342143859380219906)",
    "1127969": "emmmm，maybe they just used the tricks partial to their model in the evaluation.\nAnyway,it is recognized that efficientnet have a good performance on the majority of task related to image.",
    "1126351": "Here is an easy wrapper for checking out this model in PyTorch\nhttps://github.com/lucidrains/vit-pytorch",
    "1125879": "Yeah, I just saw this on [@paperswithcode](https://twitter.com/paperswithcode/status/1342100539761434624?s=20)'s Twitter Feed. First Transformers replaced RNN's now they are coming after CNN's too 😄. Also, was surprised to see the code in Pytorch most transformer paper nowadays use JAX based ecosystems such as [google/Trax](https://github.com/google/trax) or [google/Flax](https://github.com/google/flax). Excited to see more transformers in Kaggle competitions.",
    "1125344": "@hengck23 Thoughts on this?"
  }
}