{
  "id": 199276,
  "title": "[place holder] let's try something new ... vision transformer",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199276",
  "author_name": "",
  "post_date": "2020-11-25T06:20:05.967284600Z",
  "votes": 112,
  "comment_count": 41,
  "views": 0,
  "content": "<p>… this post will be updated as my experiments complete …</p>\n<p>some plan I have:</p>\n<ul>\n<li><p>vision transformer : <a href=\"https://openreview.net/pdf?id=YicbFdNTTy\" target=\"_blank\">https://openreview.net/pdf?id=YicbFdNTTy</a><br>\n\"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale\" -ICRL 2021<br>\n<a href=\"https://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1\" target=\"_blank\">https://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1</a><br>\n<a href=\"https://www.youtube.com/watch?v=TrdevFK_am4&amp;t=100s\" target=\"_blank\">https://www.youtube.com/watch?v=TrdevFK_am4&amp;t=100s</a></p></li>\n<li><p>another transformer: <br>\n\"LAMBDANETWORKS: MODELING LONG-RANGE INTERACTIONS WITHOUT ATTENTION\"</p></li>\n<li><p>\"ResNeSt: Split-Attention Networks\"-arvix 2020<br>\n<a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">https://github.com/zhanghang1989/ResNeSt</a><br>\n<a href=\"https://www.youtube.com/watch?v=65MLer7adGo\" target=\"_blank\">https://www.youtube.com/watch?v=65MLer7adGo</a></p></li>\n<li><p>online pseudo label (online semi-supervised)<br>\nthe kaggle wheat detection competition shows that it is possible to learn unseen test data. i think the cassava leaves are quite \"similar\" hence online learning is possible</p></li>\n<li><p>offline semi-supervised learning using unlabelled images from previous kaggle dataset : <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease</a><br>\ne.g.  SimCLRv2, SwAV, MoCo (related : supervised contrastive)</p></li>\n</ul>\n<p>if you have new ideas or papers, pleas let me know!</p>",
  "messages": [
    {
      "id": "1090182",
      "postDate": "11/25/2020 06:20:05",
      "content": "<p>… this post will be updated as my experiments complete …</p>\n<p>some plan I have:</p>\n<ul>\n<li><p>vision transformer : <a href=\"https://openreview.net/pdf?id=YicbFdNTTy\" target=\"_blank\">https://openreview.net/pdf?id=YicbFdNTTy</a><br>\n\"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale\" -ICRL 2021<br>\n<a href=\"https://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1\" target=\"_blank\">https://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1</a><br>\n<a href=\"https://www.youtube.com/watch?v=TrdevFK_am4&amp;t=100s\" target=\"_blank\">https://www.youtube.com/watch?v=TrdevFK_am4&amp;t=100s</a></p></li>\n<li><p>another transformer: <br>\n\"LAMBDANETWORKS: MODELING LONG-RANGE INTERACTIONS WITHOUT ATTENTION\"</p></li>\n<li><p>\"ResNeSt: Split-Attention Networks\"-arvix 2020<br>\n<a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">https://github.com/zhanghang1989/ResNeSt</a><br>\n<a href=\"https://www.youtube.com/watch?v=65MLer7adGo\" target=\"_blank\">https://www.youtube.com/watch?v=65MLer7adGo</a></p></li>\n<li><p>online pseudo label (online semi-supervised)<br>\nthe kaggle wheat detection competition shows that it is possible to learn unseen test data. i think the cassava leaves are quite \"similar\" hence online learning is possible</p></li>\n<li><p>offline semi-supervised learning using unlabelled images from previous kaggle dataset : <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease</a><br>\ne.g.  SimCLRv2, SwAV, MoCo (related : supervised contrastive)</p></li>\n</ul>\n<p>if you have new ideas or papers, pleas let me know!</p>",
      "rawMarkdown": "... this post will be updated as my experiments complete ...\n\nsome plan I have:\n\n- vision transformer : https://openreview.net/pdf?id=YicbFdNTTy\n\"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale\" -ICRL 2021\nhttps://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1\nhttps://www.youtube.com/watch?v=TrdevFK_am4&t=100s\n\n- another transformer: \n\"LAMBDANETWORKS: MODELING LONG-RANGE INTERACTIONS WITHOUT ATTENTION\"\n\n\n- \"ResNeSt: Split-Attention Networks\"-arvix 2020\nhttps://github.com/zhanghang1989/ResNeSt\nhttps://www.youtube.com/watch?v=65MLer7adGo\n\n\n- online pseudo label (online semi-supervised)\nthe kaggle wheat detection competition shows that it is possible to learn unseen test data. i think the cassava leaves are quite \"similar\" hence online learning is possible\n\n\n- offline semi-supervised learning using unlabelled images from previous kaggle dataset : https://www.kaggle.com/c/cassava-disease\ne.g.  SimCLRv2, SwAV, MoCo (related : supervised contrastive)\n\nif you have new ideas or papers, pleas let me know!",
      "votes": null
    },
    {
      "id": "1090222",
      "postDate": "11/25/2020 07:12:45",
      "content": "<p>As is true with most SOTA models these days, PyTorch implementations and pretrained weights are available in Ross Wightman's amazing package:<br>\n<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a></p>",
      "rawMarkdown": "As is true with most SOTA models these days, PyTorch implementations and pretrained weights are available in Ross Wightman's amazing package:\nhttps://github.com/rwightman/pytorch-image-models",
      "votes": null
    },
    {
      "id": "1090312",
      "postDate": "11/25/2020 08:50:13",
      "content": "<p>LambdaResNet is also new <br>\n<a href=\"https://openreview.net/pdf?id=xTJEN-ggl1b\" target=\"_blank\">https://openreview.net/pdf?id=xTJEN-ggl1b</a></p>",
      "rawMarkdown": "LambdaResNet is also new \nhttps://openreview.net/pdf?id=xTJEN-ggl1b",
      "votes": null
    },
    {
      "id": "1090329",
      "postDate": "11/25/2020 09:00:49",
      "content": "<p>others (still reading the paper to see if worth implementing for this challenge):</p>\n<ul>\n<li>\"Circumventing Outliers of AutoAugment with Knowledge Distillation\" Longhui Wei - arvix 2020</li>\n</ul>\n<p>do you have suggestions for adversarial augmentation?</p>\n<p>other interesting work:<br>\n<a href=\"https://www.kaggle.com/kmat2019/cycle-gan-to-enlarge-training-data\" target=\"_blank\">https://www.kaggle.com/kmat2019/cycle-gan-to-enlarge-training-data</a></p>",
      "rawMarkdown": "others (still reading the paper to see if worth implementing for this challenge):\n- \"Circumventing Outliers of AutoAugment with Knowledge Distillation\" Longhui Wei - arvix 2020\n \ndo you have suggestions for adversarial augmentation?\n\nother interesting work:\nhttps://www.kaggle.com/kmat2019/cycle-gan-to-enlarge-training-data",
      "votes": null
    },
    {
      "id": "1090587",
      "postDate": "11/25/2020 12:56:34",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , Since I saw a video about this \"Vision transformers\" I wanted to try it, the same for LambdaNets, not sure if I will have the time to try them, but I am excited to see your results, good luck!</p>",
      "rawMarkdown": "Nice work @hengck23 , Since I saw a video about this \"Vision transformers\" I wanted to try it, the same for LambdaNets, not sure if I will have the time to try them, but I am excited to see your results, good luck!",
      "votes": null
    },
    {
      "id": "1090753",
      "postDate": "11/25/2020 15:06:23",
      "content": "<p>I'm working on Bootstrap your own latent(BYOL). Let's see if it works</p>",
      "rawMarkdown": "I'm working on Bootstrap your own latent(BYOL). Let's see if it works",
      "votes": null
    },
    {
      "id": "1090783",
      "postDate": "11/25/2020 15:26:29",
      "content": "<p>the vision transformer and SimCLR don't seem to be gpu friendly…</p>\n<p>In vision transformer they talk about TPUv3 days and SimCLR requires large batches (4K-8K in their paper) and smaller batches ain't that good. MoCo v2, another semi-supervised method, claims they are much more GPU-friendly and require 8 V100s to work :) - that's exactly 8 more compared to what i have at my disposal.</p>\n<p>But take this with a grain of salt because i haven't tried it myself. TBH, I didn't read those paper fully. I stopped when I saw the computation requirements </p>",
      "rawMarkdown": "the vision transformer and SimCLR don't seem to be gpu friendly...\n\nIn vision transformer they talk about TPUv3 days and SimCLR requires large batches (4K-8K in their paper) and smaller batches ain't that good. MoCo v2, another semi-supervised method, claims they are much more GPU-friendly and require 8 V100s to work :) - that's exactly 8 more compared to what i have at my disposal.\n\nBut take this with a grain of salt because i haven't tried it myself. TBH, I didn't read those paper fully. I stopped when I saw the computation requirements",
      "votes": null
    },
    {
      "id": "1090852",
      "postDate": "11/25/2020 16:14:55",
      "content": "<p>i am thinking of using vision transformer with simple resnet18/34 encoder. i read in another post (<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198219\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198219</a>) that resnet18 can already achieve LB 0.89.</p>\n<p>in my own experiment, resnet34 is only 1% less accurate than efficientnetb4 on CV</p>\n<p>FB SwAV should be the most GPU friendly</p>",
      "rawMarkdown": "i am thinking of using vision transformer with simple resnet18/34 encoder. i read in another post (https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198219) that resnet18 can already achieve LB 0.89.\n\nin my own experiment, resnet34 is only 1% less accurate than efficientnetb4 on CV\n\nFB SwAV should be the most GPU friendly",
      "votes": null
    },
    {
      "id": "1091019",
      "postDate": "11/25/2020 18:20:32",
      "content": "<p>looking forward to see it being used on kaggle, by then i will just lurk</p>",
      "rawMarkdown": "looking forward to see it being used on kaggle, by then i will just lurk",
      "votes": null
    },
    {
      "id": "1091271",
      "postDate": "11/25/2020 22:44:02",
      "content": "<p>Here is one of the two new stuff (the methods). </p>\n<ul>\n<li><a href=\"https://www.nature.com/articles/s41598-020-68453-w\" target=\"_blank\">DA-CapsNet: dual attention mechanism capsule network</a></li>\n<li><a href=\"https://www.sciencedirect.com/science/article/pii/S1361841520302103\" target=\"_blank\">Triple attention learning for classification of 14 thoracic diseases using chest radiography</a></li>\n</ul>",
      "rawMarkdown": "Here is one of the two new stuff (the methods). \n\n- [DA-CapsNet: dual attention mechanism capsule network](https://www.nature.com/articles/s41598-020-68453-w)\n- [Triple attention learning for classification of 14 thoracic diseases using chest radiography](https://www.sciencedirect.com/science/article/pii/S1361841520302103)",
      "votes": null
    },
    {
      "id": "1091697",
      "postDate": "11/26/2020 08:00:46",
      "content": "<p>my prediction:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b30d007f92f02eb3d73d0f52dc897fe%2FSelection_038.png?generation=1606377644246455&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "my prediction:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b30d007f92f02eb3d73d0f52dc897fe%2FSelection_038.png?generation=1606377644246455&alt=media)",
      "votes": null
    },
    {
      "id": "1091728",
      "postDate": "11/26/2020 08:32:02",
      "content": "<p>how to detect duplicates from previous data</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb8b4249b943969f23c45ad34cd982a8%2FSelection_051.png?generation=1606379520098308&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how to detect duplicates from previous data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb8b4249b943969f23c45ad34cd982a8%2FSelection_051.png?generation=1606379520098308&alt=media)",
      "votes": null
    },
    {
      "id": "1091733",
      "postDate": "11/26/2020 08:37:59",
      "content": "<p>interesting, i find that color itself is a very good feature to identify the class. note the similarities between the query and the retrieved nearest neighbor. some kind of metric learning is possible.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7681fd4ae3dd4aa768af5130a2c268a%2FSelection_052.png?generation=1606379876431188&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "interesting, i find that color itself is a very good feature to identify the class. note the similarities between the query and the retrieved nearest neighbor. some kind of metric learning is possible.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7681fd4ae3dd4aa768af5130a2c268a%2FSelection_052.png?generation=1606379876431188&alt=media)",
      "votes": null
    },
    {
      "id": "1091750",
      "postDate": "11/26/2020 08:50:52",
      "content": "<p>after trying to find the intersection between previous-train, previous-test, previous-unlabelled and current-train,current-test (one image), I come to conclude that</p>\n<p>\" the private test set is probably not 100% hidden …  \"</p>\n<p>note: you have actually two test servers. one is from previous challenge and another is current challenge</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7fcd16b83c5ca4d9cf869cb4cbcdf5a1%2FSelection_054.png?generation=1606389845177489&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "after trying to find the intersection between previous-train, previous-test, previous-unlabelled and current-train,current-test (one image), I come to conclude that\n\n\" the private test set is probably not 100% hidden ...  \"\n\nnote: you have actually two test servers. one is from previous challenge and another is current challenge\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7fcd16b83c5ca4d9cf869cb4cbcdf5a1%2FSelection_054.png?generation=1606389845177489&alt=media)",
      "votes": null
    },
    {
      "id": "1091822",
      "postDate": "11/26/2020 10:09:18",
      "content": "<p>Lambda Layer implementation can be found <a href=\"https://github.com/lucidrains/lambda-networks\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "Lambda Layer implementation can be found [here](https://github.com/lucidrains/lambda-networks).",
      "votes": null
    },
    {
      "id": "1091900",
      "postDate": "11/26/2020 11:12:12",
      "content": "<p>Any class mapping from previous cassava disease challenge to this one?</p>",
      "rawMarkdown": "Any class mapping from previous cassava disease challenge to this one?",
      "votes": null
    },
    {
      "id": "1091908",
      "postDate": "11/26/2020 11:16:36",
      "content": "<p>got it. Cassava Brown Streak Disease (CBSD), Cassava Mosaic Disease (CMD), Cassava<br>\nBacterial Blight (CBB) and Cassava Green Mite (CGM) &gt; from their paper to  label maps json file:)</p>",
      "rawMarkdown": "got it. Cassava Brown Streak Disease (CBSD), Cassava Mosaic Disease (CMD), Cassava\nBacterial Blight (CBB) and Cassava Green Mite (CGM) > from their paper to  label maps json file:)",
      "votes": null
    },
    {
      "id": "1091916",
      "postDate": "11/26/2020 11:26:43",
      "content": "<p>2019-2020-duplicate_images (see attachment csv files)</p>",
      "rawMarkdown": "2019-2020-duplicate_images (see attachment csv files)",
      "votes": null
    },
    {
      "id": "1092048",
      "postDate": "11/26/2020 13:55:11",
      "content": "<p><a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> Just to confirm, if I want to use <code>EfficientNets</code>, I can also use Ross Wightman's this package <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch\" target=\"_blank\">here right?</a></p>",
      "rawMarkdown": "tanlikesmath Just to confirm, if I want to use `EfficientNets`, I can also use Ross Wightman's this package [here right?](https://github.com/rwightman/gen-efficientnet-pytorch)",
      "votes": null
    },
    {
      "id": "1092226",
      "postDate": "11/26/2020 16:17:56",
      "content": "<p>Vision Transformer……In the original paper, its batch size is too big to train on my own GPU. The original bs is set to 4096??? WTF</p>",
      "rawMarkdown": "Vision Transformer……In the original paper, its batch size is too big to train on my own GPU. The original bs is set to 4096??? WTF",
      "votes": null
    },
    {
      "id": "1092636",
      "postDate": "11/27/2020 04:08:05",
      "content": "<p>Do you think a vision transformer can outperform an efficient net in this competition?</p>",
      "rawMarkdown": "Do you think a vision transformer can outperform an efficient net in this competition?",
      "votes": null
    },
    {
      "id": "1093004",
      "postDate": "11/27/2020 11:32:05",
      "content": "<p>Thank you for good information.<br>\nCould you provide the reference to online pseudo label used on wheat detection competition? </p>",
      "rawMarkdown": "Thank you for good information.\nCould you provide the reference to online pseudo label used on wheat detection competition?",
      "votes": null
    },
    {
      "id": "1093182",
      "postDate": "11/27/2020 14:20:02",
      "content": "<p>you can make calibration graph like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a5e76d7896ec58e9fb6f165e6f410a7%2FSelection_065.png?generation=1606486788837360&amp;alt=media\" alt=\"\"></p>\n<p>you can even probe or hand label 2019 data … a better solution is to use active learning or human/oracle in the loop method</p>",
      "rawMarkdown": "you can make calibration graph like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a5e76d7896ec58e9fb6f165e6f410a7%2FSelection_065.png?generation=1606486788837360&alt=media)\n\nyou can even probe or hand label 2019 data ... a better solution is to use active learning or human/oracle in the loop method",
      "votes": null
    },
    {
      "id": "1093213",
      "postDate": "11/27/2020 14:45:51",
      "content": "<p>after reading the paper, this is a smart method!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b7394d6675c364d80695b004d905b4%2FSelection_067.png?generation=1606488349313511&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "after reading the paper, this is a smart method!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b7394d6675c364d80695b004d905b4%2FSelection_067.png?generation=1606488349313511&alt=media)",
      "votes": null
    },
    {
      "id": "1093561",
      "postDate": "11/27/2020 20:41:17",
      "content": "<p>Yep, that is what I have been using…</p>",
      "rawMarkdown": "Yep, that is what I have been using...",
      "votes": null
    },
    {
      "id": "1093895",
      "postDate": "11/28/2020 05:26:55",
      "content": "<p>2019 winning solution:<br>\n<a href=\"https://www.kaggle.com/c/cassava-disease/discussion/94114\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease/discussion/94114</a></p>\n<p>but here has more information!!<br>\n<a href=\"https://cloud.tencent.com/developer/article/1453436\" target=\"_blank\">https://cloud.tencent.com/developer/article/1453436</a><br>\n<a href=\"https://zhuanlan.zhihu.com/p/67822883\" target=\"_blank\">https://zhuanlan.zhihu.com/p/67822883</a></p>\n<p>related: <br>\n<a href=\"https://github.com/kwantommy/fgvc6-kaggle-cassava-classification\" target=\"_blank\">https://github.com/kwantommy/fgvc6-kaggle-cassava-classification</a><br>\n<a href=\"https://www.kaggle.com/c/cassava-disease/discussion/94102\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease/discussion/94102</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd2e0829f95affbaba7f97be3a2cd7359%2FSelection_071.png?generation=1606541213610479&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "2019 winning solution:\nhttps://www.kaggle.com/c/cassava-disease/discussion/94114\n\n\nbut here has more information!!\nhttps://cloud.tencent.com/developer/article/1453436\nhttps://zhuanlan.zhihu.com/p/67822883\n\nrelated: \nhttps://github.com/kwantommy/fgvc6-kaggle-cassava-classification\nhttps://www.kaggle.com/c/cassava-disease/discussion/94102\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd2e0829f95affbaba7f97be3a2cd7359%2FSelection_071.png?generation=1606541213610479&alt=media)",
      "votes": null
    },
    {
      "id": "1094102",
      "postDate": "11/28/2020 10:17:35",
      "content": "<p>interesting evaluation on<br>\n\"Evaluating the accuracy of a smartphone-based artificial intelligence system, PlantVillage Nuru, in<br>\ndiagnosing of the viral diseases of cassava\"<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2020.01.26.919449v2.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2020.01.26.919449v2.full.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3e4a60558a615396ca968fc4a19f142%2FSelection_072.png?generation=1606558622866638&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e2ae8827c3313e63ca2c87cd87b12c0%2FSelection_073.png?generation=1606558652268030&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.youtube.com/watch?v=MiGaFll32qM\" target=\"_blank\">https://www.youtube.com/watch?v=MiGaFll32qM</a><br>\n<a href=\"https://www.youtube.com/watch?v=PdQOqoRqywg\" target=\"_blank\">https://www.youtube.com/watch?v=PdQOqoRqywg</a><br>\n<a href=\"https://www.youtube.com/watch?v=71QSLCctMoI\" target=\"_blank\">https://www.youtube.com/watch?v=71QSLCctMoI</a><br>\n… hmm … there is this concept of leaves on top and leaves at bottom</p>\n<p><a href=\"https://news.psu.edu/story/485342/2017/09/29/research/new-mobile-app-diagnoses-crop-diseases-field-and-alerts-rural\" target=\"_blank\">https://news.psu.edu/story/485342/2017/09/29/research/new-mobile-app-diagnoses-crop-diseases-field-and-alerts-rural</a><br>\n<a href=\"https://www.youtube.com/channel/UCwU3ura8EXmHpQqucO74zzA/videos\" target=\"_blank\">https://www.youtube.com/channel/UCwU3ura8EXmHpQqucO74zzA/videos</a><br>\n<a href=\"https://plantvillage.psu.edu/topics/cassava-manioc/infos/diseases_and_pests_description_uses_propagation\" target=\"_blank\">https://plantvillage.psu.edu/topics/cassava-manioc/infos/diseases_and_pests_description_uses_propagation</a></p>",
      "rawMarkdown": "interesting evaluation on\n\"Evaluating the accuracy of a smartphone-based artificial intelligence system, PlantVillage Nuru, in\ndiagnosing of the viral diseases of cassava\"\nhttps://www.biorxiv.org/content/10.1101/2020.01.26.919449v2.full.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3e4a60558a615396ca968fc4a19f142%2FSelection_072.png?generation=1606558622866638&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e2ae8827c3313e63ca2c87cd87b12c0%2FSelection_073.png?generation=1606558652268030&alt=media)\n\nhttps://www.youtube.com/watch?v=MiGaFll32qM\nhttps://www.youtube.com/watch?v=PdQOqoRqywg\nhttps://www.youtube.com/watch?v=71QSLCctMoI\n... hmm ... there is this concept of leaves on top and leaves at bottom\n\nhttps://news.psu.edu/story/485342/2017/09/29/research/new-mobile-app-diagnoses-crop-diseases-field-and-alerts-rural\nhttps://www.youtube.com/channel/UCwU3ura8EXmHpQqucO74zzA/videos\nhttps://plantvillage.psu.edu/topics/cassava-manioc/infos/diseases_and_pests_description_uses_propagation",
      "votes": null
    },
    {
      "id": "1094117",
      "postDate": "11/28/2020 10:30:09",
      "content": "<p>They use TPU v3 on steroid which each core equivalent to GPU V100  ;)</p>\n<p>4096 is likely the total bs of all the 8 cores. </p>",
      "rawMarkdown": "They use TPU v3 on steroid which each core equivalent to GPU V100  ;)\n\n4096 is likely the total bs of all the 8 cores.",
      "votes": null
    },
    {
      "id": "1094131",
      "postDate": "11/28/2020 10:47:06",
      "content": "<p>Nope 1 TPU v3 core falls behind from 25% (in image classification ResNet50) to 60% (in NLP Bert) when compared to 1 A100 ;)</p>",
      "rawMarkdown": "Nope 1 TPU v3 core falls behind from 25% (in image classification ResNet50) to 60% (in NLP Bert) when compared to 1 A100 ;)",
      "votes": null
    },
    {
      "id": "1094146",
      "postDate": "11/28/2020 11:03:18",
      "content": "<p>Sorry ! I meant V100 (TPU v3 with sharded TF datasets)</p>",
      "rawMarkdown": "Sorry ! I meant V100 (TPU v3 with sharded TF datasets)",
      "votes": null
    },
    {
      "id": "1094177",
      "postDate": "11/28/2020 11:32:39",
      "content": "<p>this is based on the bounding box object detection (SSD)</p>\n<p>\"TensorFlow: an ML platform for solving impactful and challenging problems\"<br>\n<a href=\"https://www.frontiersin.org/articles/10.3389/fpls.2019.00272/full\" target=\"_blank\">https://www.frontiersin.org/articles/10.3389/fpls.2019.00272/full</a><br>\n<a href=\"https://www.youtube.com/watch?v=NlpS-DhayQA\" target=\"_blank\">https://www.youtube.com/watch?v=NlpS-DhayQA</a></p>\n<p>related:<br>\n<a href=\"https://openaccess.thecvf.com/content_CVPRW_2020/papers/w5/Tusubira_Improving_In-Field_Cassava_Whitefly_Pest_Surveillance_With_Machine_Learning_CVPRW_2020_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPRW_2020/papers/w5/Tusubira_Improving_In-Field_Cassava_Whitefly_Pest_Surveillance_With_Machine_Learning_CVPRW_2020_paper.pdf</a></p>",
      "rawMarkdown": "this is based on the bounding box object detection (SSD)\n\n\"TensorFlow: an ML platform for solving impactful and challenging problems\"\nhttps://www.frontiersin.org/articles/10.3389/fpls.2019.00272/full\nhttps://www.youtube.com/watch?v=NlpS-DhayQA\n\nrelated:\nhttps://openaccess.thecvf.com/content_CVPRW_2020/papers/w5/Tusubira_Improving_In-Field_Cassava_Whitefly_Pest_Surveillance_With_Machine_Learning_CVPRW_2020_paper.pdf",
      "votes": null
    },
    {
      "id": "1094223",
      "postDate": "11/28/2020 12:25:33",
      "content": "<p>Pretty much you can say that but TPU does have 20% advantage over v100 in image classification and 10% disadvantage in NLP.</p>",
      "rawMarkdown": "Pretty much you can say that but TPU does have 20% advantage over v100 in image classification and 10% disadvantage in NLP.",
      "votes": null
    },
    {
      "id": "1094225",
      "postDate": "11/28/2020 12:26:37",
      "content": "<p>Thanks!, Just the thing I was looking for.</p>\n<p>However I still gotta find the weights if I dont I will make the annotations myself in january.</p>",
      "rawMarkdown": "Thanks!, Just the thing I was looking for.\n\nHowever I still gotta find the weights if I dont I will make the annotations myself in january.",
      "votes": null
    },
    {
      "id": "1094306",
      "postDate": "11/28/2020 13:52:19",
      "content": "<p>tf pretrain model: <a href=\"https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2\" target=\"_blank\">https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2</a></p>",
      "rawMarkdown": "tf pretrain model: https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2",
      "votes": null
    },
    {
      "id": "1096899",
      "postDate": "11/30/2020 21:28:43",
      "content": "<p>i tested some images on tflite model the android app (NURU plantvillage)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7f800ee6b504e46dd508d06cb51bd7a%2FSelection_096.png?generation=1606771653305564&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b30ec19ffdcadb67c3c24754afdc7a8%2FSelection_097.png?generation=1606771696087399&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F53cf28499bfa499581ccb0af40a70702%2FSelection_098.png?generation=1606771715924403&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i tested some images on tflite model the android app (NURU plantvillage)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7f800ee6b504e46dd508d06cb51bd7a%2FSelection_096.png?generation=1606771653305564&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b30ec19ffdcadb67c3c24754afdc7a8%2FSelection_097.png?generation=1606771696087399&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F53cf28499bfa499581ccb0af40a70702%2FSelection_098.png?generation=1606771715924403&alt=media)",
      "votes": null
    },
    {
      "id": "1097247",
      "postDate": "12/01/2020 01:48:41",
      "content": "<p>Results are not good enough, well can not expect anything more from a tflite model.<br>\nThanks for sharing.</p>",
      "rawMarkdown": "Results are not good enough, well can not expect anything more from a tflite model.\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "1097279",
      "postDate": "12/01/2020 02:09:34",
      "content": "<p>fold stratification using clustering or classifier</p>\n<p>in theory, you can train a classifier and give a probability score to a sample.<br>\nalternatively, you can give a similarity score or distance score via clsutering.</p>\n<p>you can use either of these score for  stratified folds.</p>\n<p>if we could see the test data, clustering of 'test+train' data is the best I think</p>",
      "rawMarkdown": "fold stratification using clustering or classifier\n\nin theory, you can train a classifier and give a probability score to a sample.\nalternatively, you can give a similarity score or distance score via clsutering.\n\nyou can use either of these score for  stratified folds.\n\nif we could see the test data, clustering of 'test+train' data is the best I think",
      "votes": null
    },
    {
      "id": "1098904",
      "postDate": "12/01/2020 23:58:29",
      "content": "<p>good timming or did i just crash the server? … turns out that it is rescoring<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa3206c4fd646ad6194e7d912a778403%2FSelection_115.png?generation=1606867107076104&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "good timming or did i just crash the server? ... turns out that it is rescoring\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa3206c4fd646ad6194e7d912a778403%2FSelection_115.png?generation=1606867107076104&alt=media)",
      "votes": null
    },
    {
      "id": "1101148",
      "postDate": "12/03/2020 16:59:03",
      "content": "<p>Wonderful insights! I'd like to add a few more directions for the Vision Transformer to try:</p>\n<ul>\n<li>Multi-head and transfer learning from other tasks (or dataset)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F0e4ef412ce4a47a7cc13b09fea2afe8a%2FWX20201204-005703.png?generation=1607014648287379&amp;alt=media\" alt=\"Image Processing Transformer\"></li>\n</ul>\n<p>As we know, Image Processing Transformer (<a href=\"https://arxiv.org/pdf/2012.00364.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.00364.pdf</a>) has achieved good performance on low-level tasks such as super-resolution, denoising, draining and etc. I wonder if pre-trained models on these tasks can improve the classification performance especially for recognizing low-level features. Also, can we just ensemble different heads to obtain satisfactory results?</p>\n<ul>\n<li>NAS for vision transformers</li>\n</ul>\n<p>Resembling EfficientNet improves the performance of CNNs, network architecture search on vision transformers should improve its performance too. To the best of my knowledge, Google has carried out researches in this field, such as The Evolved Transformer (<a href=\"https://arxiv.org/abs/1901.11117)\" target=\"_blank\">https://arxiv.org/abs/1901.11117)</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F4619ba8e252c1060ffc127ba8eba4a3a%2FWX20201204-011018.png?generation=1607015436770279&amp;alt=media\" alt=\"The Evolved Transformer\"></p>\n<p>I wonder if it works on vision transformers.</p>\n<ul>\n<li>The right way to finetune vision transformers</li>\n</ul>\n<p>In past competitions, I have learned a lot from you kaggle experts about tricks to finetune <code>EfficientNet</code>. I will now keep an eye on this competition to learn how you gays tune <code>Vision Transformer</code>. I have no doubt that semi-supervised learning, such as contrastive loss, can also work for <code>Vision Transformer</code>. I plan to have a try based on your advised to carry out experiments on the previous kaggle dataset : <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease</a></p>",
      "rawMarkdown": "Wonderful insights! I'd like to add a few more directions for the Vision Transformer to try:\n- Multi-head and transfer learning from other tasks (or dataset)\n![Image Processing Transformer](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F0e4ef412ce4a47a7cc13b09fea2afe8a%2FWX20201204-005703.png?generation=1607014648287379&alt=media)\n\nAs we know, Image Processing Transformer (https://arxiv.org/pdf/2012.00364.pdf) has achieved good performance on low-level tasks such as super-resolution, denoising, draining and etc. I wonder if pre-trained models on these tasks can improve the classification performance especially for recognizing low-level features. Also, can we just ensemble different heads to obtain satisfactory results?\n- NAS for vision transformers\n\nResembling EfficientNet improves the performance of CNNs, network architecture search on vision transformers should improve its performance too. To the best of my knowledge, Google has carried out researches in this field, such as The Evolved Transformer (https://arxiv.org/abs/1901.11117).\n![The Evolved Transformer](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F4619ba8e252c1060ffc127ba8eba4a3a%2FWX20201204-011018.png?generation=1607015436770279&alt=media)\n\nI wonder if it works on vision transformers.\n- The right way to finetune vision transformers\n\nIn past competitions, I have learned a lot from you kaggle experts about tricks to finetune `EfficientNet`. I will now keep an eye on this competition to learn how you gays tune `Vision Transformer`. I have no doubt that semi-supervised learning, such as contrastive loss, can also work for `Vision Transformer`. I plan to have a try based on your advised to carry out experiments on the previous kaggle dataset : https://www.kaggle.com/c/cassava-disease",
      "votes": null
    },
    {
      "id": "1103026",
      "postDate": "12/05/2020 15:20:33",
      "content": "<p>Thank you for good information.</p>",
      "rawMarkdown": "Thank you for good information.",
      "votes": null
    },
    {
      "id": "1104957",
      "postDate": "12/07/2020 11:58:44",
      "content": "<p>Thanks for sharing the papers…This is awesome</p>\n<p>interesting to see Attention and transformer based models here too…On a lighter note…Transformers everywhere! Looks like there is no escaping them..Almost all current ongoing competitions use them one way or other..</p>\n<p>But could it be that other models are getting downplayed in all this? just thinking aloud…</p>",
      "rawMarkdown": "Thanks for sharing the papers...This is awesome\n\ninteresting to see Attention and transformer based models here too...On a lighter note...Transformers everywhere! Looks like there is no escaping them..Almost all current ongoing competitions use them one way or other..\n\nBut could it be that other models are getting downplayed in all this? just thinking aloud...",
      "votes": null
    },
    {
      "id": "1177290",
      "postDate": "01/30/2021 07:28:08",
      "content": "<p>We also used VT for another competition. You can check out our implementation: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/rythian47/vision-transformer-goodbye-cnn-training</a></p>",
      "rawMarkdown": "We also used VT for another competition. You can check out our implementation: [https://www.kaggle.com/rythian47/vision-transformer-goodbye-cnn-training](url)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1090222,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "11/25/2020 07:12:45",
      "content": "<p>As is true with most SOTA models these days, PyTorch implementations and pretrained weights are available in Ross Wightman's amazing package:<br>\n<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1092048,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "11/26/2020 13:55:11",
          "content": "<p><a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> Just to confirm, if I want to use <code>EfficientNets</code>, I can also use Ross Wightman's this package <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch\" target=\"_blank\">here right?</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093561,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "11/27/2020 20:41:17",
          "content": "<p>Yep, that is what I have been using…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1090312,
      "author_name": "szacho",
      "author_url": "",
      "post_date": "11/25/2020 08:50:13",
      "content": "<p>LambdaResNet is also new <br>\n<a href=\"https://openreview.net/pdf?id=xTJEN-ggl1b\" target=\"_blank\">https://openreview.net/pdf?id=xTJEN-ggl1b</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1091822,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "11/26/2020 10:09:18",
          "content": "<p>Lambda Layer implementation can be found <a href=\"https://github.com/lucidrains/lambda-networks\" target=\"_blank\">here</a>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1090329,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/25/2020 09:00:49",
      "content": "<p>others (still reading the paper to see if worth implementing for this challenge):</p>\n<ul>\n<li>\"Circumventing Outliers of AutoAugment with Knowledge Distillation\" Longhui Wei - arvix 2020</li>\n</ul>\n<p>do you have suggestions for adversarial augmentation?</p>\n<p>other interesting work:<br>\n<a href=\"https://www.kaggle.com/kmat2019/cycle-gan-to-enlarge-training-data\" target=\"_blank\">https://www.kaggle.com/kmat2019/cycle-gan-to-enlarge-training-data</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1093213,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/27/2020 14:45:51",
          "content": "<p>after reading the paper, this is a smart method!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b7394d6675c364d80695b004d905b4%2FSelection_067.png?generation=1606488349313511&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1090587,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "11/25/2020 12:56:34",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , Since I saw a video about this \"Vision transformers\" I wanted to try it, the same for LambdaNets, not sure if I will have the time to try them, but I am excited to see your results, good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1090753,
      "author_name": "atharvap329",
      "author_url": "",
      "post_date": "11/25/2020 15:06:23",
      "content": "<p>I'm working on Bootstrap your own latent(BYOL). Let's see if it works</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1090783,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "11/25/2020 15:26:29",
      "content": "<p>the vision transformer and SimCLR don't seem to be gpu friendly…</p>\n<p>In vision transformer they talk about TPUv3 days and SimCLR requires large batches (4K-8K in their paper) and smaller batches ain't that good. MoCo v2, another semi-supervised method, claims they are much more GPU-friendly and require 8 V100s to work :) - that's exactly 8 more compared to what i have at my disposal.</p>\n<p>But take this with a grain of salt because i haven't tried it myself. TBH, I didn't read those paper fully. I stopped when I saw the computation requirements </p>",
      "votes": null,
      "replies": [
        {
          "id": 1090852,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/25/2020 16:14:55",
          "content": "<p>i am thinking of using vision transformer with simple resnet18/34 encoder. i read in another post (<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198219\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198219</a>) that resnet18 can already achieve LB 0.89.</p>\n<p>in my own experiment, resnet34 is only 1% less accurate than efficientnetb4 on CV</p>\n<p>FB SwAV should be the most GPU friendly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091019,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "11/25/2020 18:20:32",
          "content": "<p>looking forward to see it being used on kaggle, by then i will just lurk</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091271,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "11/25/2020 22:44:02",
      "content": "<p>Here is one of the two new stuff (the methods). </p>\n<ul>\n<li><a href=\"https://www.nature.com/articles/s41598-020-68453-w\" target=\"_blank\">DA-CapsNet: dual attention mechanism capsule network</a></li>\n<li><a href=\"https://www.sciencedirect.com/science/article/pii/S1361841520302103\" target=\"_blank\">Triple attention learning for classification of 14 thoracic diseases using chest radiography</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091697,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/26/2020 08:00:46",
      "content": "<p>my prediction:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b30d007f92f02eb3d73d0f52dc897fe%2FSelection_038.png?generation=1606377644246455&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091728,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/26/2020 08:32:02",
      "content": "<p>how to detect duplicates from previous data</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb8b4249b943969f23c45ad34cd982a8%2FSelection_051.png?generation=1606379520098308&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1091733,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/26/2020 08:37:59",
          "content": "<p>interesting, i find that color itself is a very good feature to identify the class. note the similarities between the query and the retrieved nearest neighbor. some kind of metric learning is possible.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7681fd4ae3dd4aa768af5130a2c268a%2FSelection_052.png?generation=1606379876431188&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091750,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/26/2020 08:50:52",
          "content": "<p>after trying to find the intersection between previous-train, previous-test, previous-unlabelled and current-train,current-test (one image), I come to conclude that</p>\n<p>\" the private test set is probably not 100% hidden …  \"</p>\n<p>note: you have actually two test servers. one is from previous challenge and another is current challenge</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7fcd16b83c5ca4d9cf869cb4cbcdf5a1%2FSelection_054.png?generation=1606389845177489&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091916,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/26/2020 11:26:43",
          "content": "<p>2019-2020-duplicate_images (see attachment csv files)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093182,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/27/2020 14:20:02",
          "content": "<p>you can make calibration graph like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a5e76d7896ec58e9fb6f165e6f410a7%2FSelection_065.png?generation=1606486788837360&amp;alt=media\" alt=\"\"></p>\n<p>you can even probe or hand label 2019 data … a better solution is to use active learning or human/oracle in the loop method</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091900,
      "author_name": "lukachkhetiani",
      "author_url": "",
      "post_date": "11/26/2020 11:12:12",
      "content": "<p>Any class mapping from previous cassava disease challenge to this one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091908,
          "author_name": "lukachkhetiani",
          "author_url": "",
          "post_date": "11/26/2020 11:16:36",
          "content": "<p>got it. Cassava Brown Streak Disease (CBSD), Cassava Mosaic Disease (CMD), Cassava<br>\nBacterial Blight (CBB) and Cassava Green Mite (CGM) &gt; from their paper to  label maps json file:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092226,
      "author_name": "tenyond",
      "author_url": "",
      "post_date": "11/26/2020 16:17:56",
      "content": "<p>Vision Transformer……In the original paper, its batch size is too big to train on my own GPU. The original bs is set to 4096??? WTF</p>",
      "votes": null,
      "replies": [
        {
          "id": 1094117,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "11/28/2020 10:30:09",
          "content": "<p>They use TPU v3 on steroid which each core equivalent to GPU V100  ;)</p>\n<p>4096 is likely the total bs of all the 8 cores. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094131,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/28/2020 10:47:06",
          "content": "<p>Nope 1 TPU v3 core falls behind from 25% (in image classification ResNet50) to 60% (in NLP Bert) when compared to 1 A100 ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094146,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "11/28/2020 11:03:18",
          "content": "<p>Sorry ! I meant V100 (TPU v3 with sharded TF datasets)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094223,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/28/2020 12:25:33",
          "content": "<p>Pretty much you can say that but TPU does have 20% advantage over v100 in image classification and 10% disadvantage in NLP.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1092636,
      "author_name": "aryankhatana",
      "author_url": "",
      "post_date": "11/27/2020 04:08:05",
      "content": "<p>Do you think a vision transformer can outperform an efficient net in this competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1093004,
      "author_name": "jinkyh",
      "author_url": "",
      "post_date": "11/27/2020 11:32:05",
      "content": "<p>Thank you for good information.<br>\nCould you provide the reference to online pseudo label used on wheat detection competition? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1093895,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/28/2020 05:26:55",
      "content": "<p>2019 winning solution:<br>\n<a href=\"https://www.kaggle.com/c/cassava-disease/discussion/94114\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease/discussion/94114</a></p>\n<p>but here has more information!!<br>\n<a href=\"https://cloud.tencent.com/developer/article/1453436\" target=\"_blank\">https://cloud.tencent.com/developer/article/1453436</a><br>\n<a href=\"https://zhuanlan.zhihu.com/p/67822883\" target=\"_blank\">https://zhuanlan.zhihu.com/p/67822883</a></p>\n<p>related: <br>\n<a href=\"https://github.com/kwantommy/fgvc6-kaggle-cassava-classification\" target=\"_blank\">https://github.com/kwantommy/fgvc6-kaggle-cassava-classification</a><br>\n<a href=\"https://www.kaggle.com/c/cassava-disease/discussion/94102\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease/discussion/94102</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd2e0829f95affbaba7f97be3a2cd7359%2FSelection_071.png?generation=1606541213610479&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1094102,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/28/2020 10:17:35",
      "content": "<p>interesting evaluation on<br>\n\"Evaluating the accuracy of a smartphone-based artificial intelligence system, PlantVillage Nuru, in<br>\ndiagnosing of the viral diseases of cassava\"<br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2020.01.26.919449v2.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2020.01.26.919449v2.full.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3e4a60558a615396ca968fc4a19f142%2FSelection_072.png?generation=1606558622866638&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e2ae8827c3313e63ca2c87cd87b12c0%2FSelection_073.png?generation=1606558652268030&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.youtube.com/watch?v=MiGaFll32qM\" target=\"_blank\">https://www.youtube.com/watch?v=MiGaFll32qM</a><br>\n<a href=\"https://www.youtube.com/watch?v=PdQOqoRqywg\" target=\"_blank\">https://www.youtube.com/watch?v=PdQOqoRqywg</a><br>\n<a href=\"https://www.youtube.com/watch?v=71QSLCctMoI\" target=\"_blank\">https://www.youtube.com/watch?v=71QSLCctMoI</a><br>\n… hmm … there is this concept of leaves on top and leaves at bottom</p>\n<p><a href=\"https://news.psu.edu/story/485342/2017/09/29/research/new-mobile-app-diagnoses-crop-diseases-field-and-alerts-rural\" target=\"_blank\">https://news.psu.edu/story/485342/2017/09/29/research/new-mobile-app-diagnoses-crop-diseases-field-and-alerts-rural</a><br>\n<a href=\"https://www.youtube.com/channel/UCwU3ura8EXmHpQqucO74zzA/videos\" target=\"_blank\">https://www.youtube.com/channel/UCwU3ura8EXmHpQqucO74zzA/videos</a><br>\n<a href=\"https://plantvillage.psu.edu/topics/cassava-manioc/infos/diseases_and_pests_description_uses_propagation\" target=\"_blank\">https://plantvillage.psu.edu/topics/cassava-manioc/infos/diseases_and_pests_description_uses_propagation</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1094177,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/28/2020 11:32:39",
      "content": "<p>this is based on the bounding box object detection (SSD)</p>\n<p>\"TensorFlow: an ML platform for solving impactful and challenging problems\"<br>\n<a href=\"https://www.frontiersin.org/articles/10.3389/fpls.2019.00272/full\" target=\"_blank\">https://www.frontiersin.org/articles/10.3389/fpls.2019.00272/full</a><br>\n<a href=\"https://www.youtube.com/watch?v=NlpS-DhayQA\" target=\"_blank\">https://www.youtube.com/watch?v=NlpS-DhayQA</a></p>\n<p>related:<br>\n<a href=\"https://openaccess.thecvf.com/content_CVPRW_2020/papers/w5/Tusubira_Improving_In-Field_Cassava_Whitefly_Pest_Surveillance_With_Machine_Learning_CVPRW_2020_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPRW_2020/papers/w5/Tusubira_Improving_In-Field_Cassava_Whitefly_Pest_Surveillance_With_Machine_Learning_CVPRW_2020_paper.pdf</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1094225,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "11/28/2020 12:26:37",
          "content": "<p>Thanks!, Just the thing I was looking for.</p>\n<p>However I still gotta find the weights if I dont I will make the annotations myself in january.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096899,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/30/2020 21:28:43",
          "content": "<p>i tested some images on tflite model the android app (NURU plantvillage)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7f800ee6b504e46dd508d06cb51bd7a%2FSelection_096.png?generation=1606771653305564&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b30ec19ffdcadb67c3c24754afdc7a8%2FSelection_097.png?generation=1606771696087399&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F53cf28499bfa499581ccb0af40a70702%2FSelection_098.png?generation=1606771715924403&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1097247,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "12/01/2020 01:48:41",
          "content": "<p>Results are not good enough, well can not expect anything more from a tflite model.<br>\nThanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1094306,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/28/2020 13:52:19",
      "content": "<p>tf pretrain model: <a href=\"https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2\" target=\"_blank\">https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1097279,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/01/2020 02:09:34",
      "content": "<p>fold stratification using clustering or classifier</p>\n<p>in theory, you can train a classifier and give a probability score to a sample.<br>\nalternatively, you can give a similarity score or distance score via clsutering.</p>\n<p>you can use either of these score for  stratified folds.</p>\n<p>if we could see the test data, clustering of 'test+train' data is the best I think</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1098904,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/01/2020 23:58:29",
      "content": "<p>good timming or did i just crash the server? … turns out that it is rescoring<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa3206c4fd646ad6194e7d912a778403%2FSelection_115.png?generation=1606867107076104&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1101148,
      "author_name": "szuzhangzhi",
      "author_url": "",
      "post_date": "12/03/2020 16:59:03",
      "content": "<p>Wonderful insights! I'd like to add a few more directions for the Vision Transformer to try:</p>\n<ul>\n<li>Multi-head and transfer learning from other tasks (or dataset)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F0e4ef412ce4a47a7cc13b09fea2afe8a%2FWX20201204-005703.png?generation=1607014648287379&amp;alt=media\" alt=\"Image Processing Transformer\"></li>\n</ul>\n<p>As we know, Image Processing Transformer (<a href=\"https://arxiv.org/pdf/2012.00364.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.00364.pdf</a>) has achieved good performance on low-level tasks such as super-resolution, denoising, draining and etc. I wonder if pre-trained models on these tasks can improve the classification performance especially for recognizing low-level features. Also, can we just ensemble different heads to obtain satisfactory results?</p>\n<ul>\n<li>NAS for vision transformers</li>\n</ul>\n<p>Resembling EfficientNet improves the performance of CNNs, network architecture search on vision transformers should improve its performance too. To the best of my knowledge, Google has carried out researches in this field, such as The Evolved Transformer (<a href=\"https://arxiv.org/abs/1901.11117)\" target=\"_blank\">https://arxiv.org/abs/1901.11117)</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F4619ba8e252c1060ffc127ba8eba4a3a%2FWX20201204-011018.png?generation=1607015436770279&amp;alt=media\" alt=\"The Evolved Transformer\"></p>\n<p>I wonder if it works on vision transformers.</p>\n<ul>\n<li>The right way to finetune vision transformers</li>\n</ul>\n<p>In past competitions, I have learned a lot from you kaggle experts about tricks to finetune <code>EfficientNet</code>. I will now keep an eye on this competition to learn how you gays tune <code>Vision Transformer</code>. I have no doubt that semi-supervised learning, such as contrastive loss, can also work for <code>Vision Transformer</code>. I plan to have a try based on your advised to carry out experiments on the previous kaggle dataset : <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103026,
      "author_name": "shoiwa",
      "author_url": "",
      "post_date": "12/05/2020 15:20:33",
      "content": "<p>Thank you for good information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1104957,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "12/07/2020 11:58:44",
      "content": "<p>Thanks for sharing the papers…This is awesome</p>\n<p>interesting to see Attention and transformer based models here too…On a lighter note…Transformers everywhere! Looks like there is no escaping them..Almost all current ongoing competitions use them one way or other..</p>\n<p>But could it be that other models are getting downplayed in all this? just thinking aloud…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1177290,
      "author_name": "rythian47",
      "author_url": "",
      "post_date": "01/30/2021 07:28:08",
      "content": "<p>We also used VT for another competition. You can check out our implementation: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/rythian47/vision-transformer-goodbye-cnn-training</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1090182": "... this post will be updated as my experiments complete ...\n\nsome plan I have:\n\n- vision transformer : https://openreview.net/pdf?id=YicbFdNTTy\n\"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale\" -ICRL 2021\nhttps://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1\nhttps://www.youtube.com/watch?v=TrdevFK_am4&t=100s\n\n- another transformer: \n\"LAMBDANETWORKS: MODELING LONG-RANGE INTERACTIONS WITHOUT ATTENTION\"\n\n\n- \"ResNeSt: Split-Attention Networks\"-arvix 2020\nhttps://github.com/zhanghang1989/ResNeSt\nhttps://www.youtube.com/watch?v=65MLer7adGo\n\n\n- online pseudo label (online semi-supervised)\nthe kaggle wheat detection competition shows that it is possible to learn unseen test data. i think the cassava leaves are quite \"similar\" hence online learning is possible\n\n\n- offline semi-supervised learning using unlabelled images from previous kaggle dataset : https://www.kaggle.com/c/cassava-disease\ne.g.  SimCLRv2, SwAV, MoCo (related : supervised contrastive)\n\nif you have new ideas or papers, pleas let me know!",
    "1090222": "As is true with most SOTA models these days, PyTorch implementations and pretrained weights are available in Ross Wightman's amazing package:\nhttps://github.com/rwightman/pytorch-image-models",
    "1090312": "LambdaResNet is also new \nhttps://openreview.net/pdf?id=xTJEN-ggl1b",
    "1090329": "others (still reading the paper to see if worth implementing for this challenge):\n- \"Circumventing Outliers of AutoAugment with Knowledge Distillation\" Longhui Wei - arvix 2020\n \ndo you have suggestions for adversarial augmentation?\n\nother interesting work:\nhttps://www.kaggle.com/kmat2019/cycle-gan-to-enlarge-training-data",
    "1090587": "Nice work @hengck23 , Since I saw a video about this \"Vision transformers\" I wanted to try it, the same for LambdaNets, not sure if I will have the time to try them, but I am excited to see your results, good luck!",
    "1090753": "I'm working on Bootstrap your own latent(BYOL). Let's see if it works",
    "1090783": "the vision transformer and SimCLR don't seem to be gpu friendly...\n\nIn vision transformer they talk about TPUv3 days and SimCLR requires large batches (4K-8K in their paper) and smaller batches ain't that good. MoCo v2, another semi-supervised method, claims they are much more GPU-friendly and require 8 V100s to work :) - that's exactly 8 more compared to what i have at my disposal.\n\nBut take this with a grain of salt because i haven't tried it myself. TBH, I didn't read those paper fully. I stopped when I saw the computation requirements",
    "1090852": "i am thinking of using vision transformer with simple resnet18/34 encoder. i read in another post (https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198219) that resnet18 can already achieve LB 0.89.\n\nin my own experiment, resnet34 is only 1% less accurate than efficientnetb4 on CV\n\nFB SwAV should be the most GPU friendly",
    "1091019": "looking forward to see it being used on kaggle, by then i will just lurk",
    "1091271": "Here is one of the two new stuff (the methods). \n\n- [DA-CapsNet: dual attention mechanism capsule network](https://www.nature.com/articles/s41598-020-68453-w)\n- [Triple attention learning for classification of 14 thoracic diseases using chest radiography](https://www.sciencedirect.com/science/article/pii/S1361841520302103)",
    "1091697": "my prediction:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b30d007f92f02eb3d73d0f52dc897fe%2FSelection_038.png?generation=1606377644246455&alt=media)",
    "1091728": "how to detect duplicates from previous data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb8b4249b943969f23c45ad34cd982a8%2FSelection_051.png?generation=1606379520098308&alt=media)",
    "1091733": "interesting, i find that color itself is a very good feature to identify the class. note the similarities between the query and the retrieved nearest neighbor. some kind of metric learning is possible.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7681fd4ae3dd4aa768af5130a2c268a%2FSelection_052.png?generation=1606379876431188&alt=media)",
    "1091750": "after trying to find the intersection between previous-train, previous-test, previous-unlabelled and current-train,current-test (one image), I come to conclude that\n\n\" the private test set is probably not 100% hidden ...  \"\n\nnote: you have actually two test servers. one is from previous challenge and another is current challenge\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7fcd16b83c5ca4d9cf869cb4cbcdf5a1%2FSelection_054.png?generation=1606389845177489&alt=media)",
    "1091822": "Lambda Layer implementation can be found [here](https://github.com/lucidrains/lambda-networks).",
    "1091900": "Any class mapping from previous cassava disease challenge to this one?",
    "1091908": "got it. Cassava Brown Streak Disease (CBSD), Cassava Mosaic Disease (CMD), Cassava\nBacterial Blight (CBB) and Cassava Green Mite (CGM) > from their paper to  label maps json file:)",
    "1091916": "2019-2020-duplicate_images (see attachment csv files)",
    "1092048": "tanlikesmath Just to confirm, if I want to use `EfficientNets`, I can also use Ross Wightman's this package [here right?](https://github.com/rwightman/gen-efficientnet-pytorch)",
    "1092226": "Vision Transformer……In the original paper, its batch size is too big to train on my own GPU. The original bs is set to 4096??? WTF",
    "1092636": "Do you think a vision transformer can outperform an efficient net in this competition?",
    "1093004": "Thank you for good information.\nCould you provide the reference to online pseudo label used on wheat detection competition?",
    "1093182": "you can make calibration graph like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a5e76d7896ec58e9fb6f165e6f410a7%2FSelection_065.png?generation=1606486788837360&alt=media)\n\nyou can even probe or hand label 2019 data ... a better solution is to use active learning or human/oracle in the loop method",
    "1093213": "after reading the paper, this is a smart method!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b7394d6675c364d80695b004d905b4%2FSelection_067.png?generation=1606488349313511&alt=media)",
    "1093561": "Yep, that is what I have been using...",
    "1093895": "2019 winning solution:\nhttps://www.kaggle.com/c/cassava-disease/discussion/94114\n\n\nbut here has more information!!\nhttps://cloud.tencent.com/developer/article/1453436\nhttps://zhuanlan.zhihu.com/p/67822883\n\nrelated: \nhttps://github.com/kwantommy/fgvc6-kaggle-cassava-classification\nhttps://www.kaggle.com/c/cassava-disease/discussion/94102\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd2e0829f95affbaba7f97be3a2cd7359%2FSelection_071.png?generation=1606541213610479&alt=media)",
    "1094102": "interesting evaluation on\n\"Evaluating the accuracy of a smartphone-based artificial intelligence system, PlantVillage Nuru, in\ndiagnosing of the viral diseases of cassava\"\nhttps://www.biorxiv.org/content/10.1101/2020.01.26.919449v2.full.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc3e4a60558a615396ca968fc4a19f142%2FSelection_072.png?generation=1606558622866638&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e2ae8827c3313e63ca2c87cd87b12c0%2FSelection_073.png?generation=1606558652268030&alt=media)\n\nhttps://www.youtube.com/watch?v=MiGaFll32qM\nhttps://www.youtube.com/watch?v=PdQOqoRqywg\nhttps://www.youtube.com/watch?v=71QSLCctMoI\n... hmm ... there is this concept of leaves on top and leaves at bottom\n\nhttps://news.psu.edu/story/485342/2017/09/29/research/new-mobile-app-diagnoses-crop-diseases-field-and-alerts-rural\nhttps://www.youtube.com/channel/UCwU3ura8EXmHpQqucO74zzA/videos\nhttps://plantvillage.psu.edu/topics/cassava-manioc/infos/diseases_and_pests_description_uses_propagation",
    "1094117": "They use TPU v3 on steroid which each core equivalent to GPU V100  ;)\n\n4096 is likely the total bs of all the 8 cores.",
    "1094131": "Nope 1 TPU v3 core falls behind from 25% (in image classification ResNet50) to 60% (in NLP Bert) when compared to 1 A100 ;)",
    "1094146": "Sorry ! I meant V100 (TPU v3 with sharded TF datasets)",
    "1094177": "this is based on the bounding box object detection (SSD)\n\n\"TensorFlow: an ML platform for solving impactful and challenging problems\"\nhttps://www.frontiersin.org/articles/10.3389/fpls.2019.00272/full\nhttps://www.youtube.com/watch?v=NlpS-DhayQA\n\nrelated:\nhttps://openaccess.thecvf.com/content_CVPRW_2020/papers/w5/Tusubira_Improving_In-Field_Cassava_Whitefly_Pest_Surveillance_With_Machine_Learning_CVPRW_2020_paper.pdf",
    "1094223": "Pretty much you can say that but TPU does have 20% advantage over v100 in image classification and 10% disadvantage in NLP.",
    "1094225": "Thanks!, Just the thing I was looking for.\n\nHowever I still gotta find the weights if I dont I will make the annotations myself in january.",
    "1094306": "tf pretrain model: https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2",
    "1096899": "i tested some images on tflite model the android app (NURU plantvillage)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd7f800ee6b504e46dd508d06cb51bd7a%2FSelection_096.png?generation=1606771653305564&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b30ec19ffdcadb67c3c24754afdc7a8%2FSelection_097.png?generation=1606771696087399&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F53cf28499bfa499581ccb0af40a70702%2FSelection_098.png?generation=1606771715924403&alt=media)",
    "1097247": "Results are not good enough, well can not expect anything more from a tflite model.\nThanks for sharing.",
    "1097279": "fold stratification using clustering or classifier\n\nin theory, you can train a classifier and give a probability score to a sample.\nalternatively, you can give a similarity score or distance score via clsutering.\n\nyou can use either of these score for  stratified folds.\n\nif we could see the test data, clustering of 'test+train' data is the best I think",
    "1098904": "good timming or did i just crash the server? ... turns out that it is rescoring\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa3206c4fd646ad6194e7d912a778403%2FSelection_115.png?generation=1606867107076104&alt=media)",
    "1101148": "Wonderful insights! I'd like to add a few more directions for the Vision Transformer to try:\n- Multi-head and transfer learning from other tasks (or dataset)\n![Image Processing Transformer](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F0e4ef412ce4a47a7cc13b09fea2afe8a%2FWX20201204-005703.png?generation=1607014648287379&alt=media)\n\nAs we know, Image Processing Transformer (https://arxiv.org/pdf/2012.00364.pdf) has achieved good performance on low-level tasks such as super-resolution, denoising, draining and etc. I wonder if pre-trained models on these tasks can improve the classification performance especially for recognizing low-level features. Also, can we just ensemble different heads to obtain satisfactory results?\n- NAS for vision transformers\n\nResembling EfficientNet improves the performance of CNNs, network architecture search on vision transformers should improve its performance too. To the best of my knowledge, Google has carried out researches in this field, such as The Evolved Transformer (https://arxiv.org/abs/1901.11117).\n![The Evolved Transformer](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3262514%2F4619ba8e252c1060ffc127ba8eba4a3a%2FWX20201204-011018.png?generation=1607015436770279&alt=media)\n\nI wonder if it works on vision transformers.\n- The right way to finetune vision transformers\n\nIn past competitions, I have learned a lot from you kaggle experts about tricks to finetune `EfficientNet`. I will now keep an eye on this competition to learn how you gays tune `Vision Transformer`. I have no doubt that semi-supervised learning, such as contrastive loss, can also work for `Vision Transformer`. I plan to have a try based on your advised to carry out experiments on the previous kaggle dataset : https://www.kaggle.com/c/cassava-disease",
    "1103026": "Thank you for good information.",
    "1104957": "Thanks for sharing the papers...This is awesome\n\ninteresting to see Attention and transformer based models here too...On a lighter note...Transformers everywhere! Looks like there is no escaping them..Almost all current ongoing competitions use them one way or other..\n\nBut could it be that other models are getting downplayed in all this? just thinking aloud...",
    "1177290": "We also used VT for another competition. You can check out our implementation: [https://www.kaggle.com/rythian47/vision-transformer-goodbye-cnn-training](url)"
  },
  "source": "meta"
}