{
  "id": 215359,
  "title": "Discussion on what works and what does not",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215359",
  "author_name": "",
  "post_date": "2021-01-29T14:28:11.820731Z",
  "votes": 21,
  "comment_count": 12,
  "views": 0,
  "content": "<p><strong>The following things seem to work:</strong></p>\n<ol>\n<li>Model EfficientNet seems to work better</li>\n<li>25 Epoch and 5 folds is giving good result</li>\n<li>Cutmix is also helping in increasing  the lb and cv.</li>\n<li>Individual folds might give more score but using all folds give more stability.</li>\n<li>Old competition data gives little rise in lb.<br>\nTo get old data in tf records click <a href=\"https://www.kaggle.com/vickygoyal/oldcassavadata\" target=\"_blank\">here</a>. The details of the dataset is  <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490\" target=\"_blank\">here</a></li>\n<li>Ensembling gives rise in lb.</li>\n</ol>",
  "messages": [
    {
      "id": "1176169",
      "postDate": "01/29/2021 14:28:11",
      "content": "<p><strong>The following things seem to work:</strong></p>\n<ol>\n<li>Model EfficientNet seems to work better</li>\n<li>25 Epoch and 5 folds is giving good result</li>\n<li>Cutmix is also helping in increasing  the lb and cv.</li>\n<li>Individual folds might give more score but using all folds give more stability.</li>\n<li>Old competition data gives little rise in lb.<br>\nTo get old data in tf records click <a href=\"https://www.kaggle.com/vickygoyal/oldcassavadata\" target=\"_blank\">here</a>. The details of the dataset is  <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490\" target=\"_blank\">here</a></li>\n<li>Ensembling gives rise in lb.</li>\n</ol>",
      "rawMarkdown": "**The following things seem to work:**\n1. Model EfficientNet seems to work better\n2. 25 Epoch and 5 folds is giving good result\n3. Cutmix is also helping in increasing  the lb and cv.\n4. Individual folds might give more score but using all folds give more stability.\n5. Old competition data gives little rise in lb.\nTo get old data in tf records click [here](https://www.kaggle.com/vickygoyal/oldcassavadata). The details of the dataset is  [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490)\n6. Ensembling gives rise in lb.",
      "votes": null
    },
    {
      "id": "1177434",
      "postDate": "01/30/2021 09:31:03",
      "content": "<p>I have yet to try cut mix,  I am using early stopping, never getting past 20 epochs as it stops early. </p>\n<p>What weights do you use (imagenet/noisy-student)?<br>\nI saw people saying bitempered logistic loss was giving good results but that is not the case for me. What loss do you use?</p>",
      "rawMarkdown": "I have yet to try cut mix,  I am using early stopping, never getting past 20 epochs as it stops early. \n\nWhat weights do you use (imagenet/noisy-student)?\nI saw people saying bitempered logistic loss was giving good results but that is not the case for me. What loss do you use?",
      "votes": null
    },
    {
      "id": "1177483",
      "postDate": "01/30/2021 10:03:18",
      "content": "<p>I used noisy-student and same old categorical_crossentropy loss.</p>",
      "rawMarkdown": "I used noisy-student and same old categorical_crossentropy loss.",
      "votes": null
    },
    {
      "id": "1177879",
      "postDate": "01/30/2021 15:03:02",
      "content": "<p>I want to ask, what image size you use? I used 448x448 (initial results were 384x384) which increased my CV, however I  didn't try 512x512 since I ran out of memory (with batch size of 32). Maybe i need to lower the batch size, but i don't know if it is good, since we have noisy labels?<br>\nCheers.</p>",
      "rawMarkdown": "I want to ask, what image size you use? I used 448x448 (initial results were 384x384) which increased my CV, however I  didn't try 512x512 since I ran out of memory (with batch size of 32). Maybe i need to lower the batch size, but i don't know if it is good, since we have noisy labels?\nCheers.",
      "votes": null
    },
    {
      "id": "1178096",
      "postDate": "01/30/2021 16:40:57",
      "content": "<p>I used 512*512 using TPU. On GPU you will run out of memory in that. Image size of 448 is also fine.  Image size between 384-512 are good enough for the training.</p>",
      "rawMarkdown": "I used 512*512 using TPU. On GPU you will run out of memory in that. Image size of 448 is also fine.  Image size between 384-512 are good enough for the training.",
      "votes": null
    },
    {
      "id": "1179889",
      "postDate": "02/01/2021 00:03:43",
      "content": "<p>All of what you mentioned also do work in my case</p>",
      "rawMarkdown": "All of what you mentioned also do work in my case",
      "votes": null
    },
    {
      "id": "1181352",
      "postDate": "02/01/2021 20:02:26",
      "content": "<p>Removing the noise from the dataset reduced my lb score which tells that test data is also noisy on similar lines.</p>",
      "rawMarkdown": "Removing the noise from the dataset reduced my lb score which tells that test data is also noisy on similar lines.",
      "votes": null
    },
    {
      "id": "1181598",
      "postDate": "02/02/2021 02:18:52",
      "content": "<p>25 Epoch and 5 folds mean every fold train  5 Epoch ? </p>",
      "rawMarkdown": "25 Epoch and 5 folds mean every fold train  5 Epoch ?",
      "votes": null
    },
    {
      "id": "1181750",
      "postDate": "02/02/2021 05:12:08",
      "content": "<p>Dataset divided into 5 folds and each fold trained for 25 epoch.</p>",
      "rawMarkdown": "Dataset divided into 5 folds and each fold trained for 25 epoch.",
      "votes": null
    },
    {
      "id": "1181999",
      "postDate": "02/02/2021 08:47:16",
      "content": "<p>Oh thanks.</p>",
      "rawMarkdown": "Oh thanks.",
      "votes": null
    },
    {
      "id": "1199495",
      "postDate": "02/13/2021 20:56:25",
      "content": "<p>For bi tempered loss refer to link <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "For bi tempered loss refer to link [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911)",
      "votes": null
    },
    {
      "id": "1199507",
      "postDate": "02/13/2021 21:09:15",
      "content": "<p>In my modest opinion, to achieve a reasonable (not the best, which is beyond my reach) score at the LB (above 0.902) it takes:</p>\n<ul>\n<li>to train a couple, three or four CV models with pretrained weights from: ResNet variants (ssl|swsl)<em>resnext(50|101)_32x(4|8|16), (tf| )_efficientnet</em>(b3|b4|b4)_ns) with 2020 dataset, maybe upsized with 2018 training dataset, and perhaps with extra and test images from 2018 competition</li>\n<li>Trying different ensemble them in differente ways: excluding worst models, combining good predictors for classes 0,1,2 and 4, etc.</li>\n<li>hyperparameter tuning transformations and number of iterations for TTA. In my experience, very light TTA works best (random crop, H and V flips).</li>\n</ul>",
      "rawMarkdown": "In my modest opinion, to achieve a reasonable (not the best, which is beyond my reach) score at the LB (above 0.902) it takes:\n- to train a couple, three or four CV models with pretrained weights from: ResNet variants (ssl|swsl)_resnext(50|101)_32x(4|8|16), (tf| )_efficientnet_(b3|b4|b4)_ns) with 2020 dataset, maybe upsized with 2018 training dataset, and perhaps with extra and test images from 2018 competition\n- Trying different ensemble them in differente ways: excluding worst models, combining good predictors for classes 0,1,2 and 4, etc.\n- hyperparameter tuning transformations and number of iterations for TTA. In my experience, very light TTA works best (random crop, H and V flips).",
      "votes": null
    },
    {
      "id": "1199518",
      "postDate": "02/13/2021 21:18:56",
      "content": "<p>Thanks, I did what u mentioned in point 1 and partially in point 2 but i have not done tta, will try it.</p>",
      "rawMarkdown": "Thanks, I did what u mentioned in point 1 and partially in point 2 but i have not done tta, will try it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1177434,
      "author_name": "mohneesh7",
      "author_url": "",
      "post_date": "01/30/2021 09:31:03",
      "content": "<p>I have yet to try cut mix,  I am using early stopping, never getting past 20 epochs as it stops early. </p>\n<p>What weights do you use (imagenet/noisy-student)?<br>\nI saw people saying bitempered logistic loss was giving good results but that is not the case for me. What loss do you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1177483,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "01/30/2021 10:03:18",
          "content": "<p>I used noisy-student and same old categorical_crossentropy loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1177879,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "01/30/2021 15:03:02",
      "content": "<p>I want to ask, what image size you use? I used 448x448 (initial results were 384x384) which increased my CV, however I  didn't try 512x512 since I ran out of memory (with batch size of 32). Maybe i need to lower the batch size, but i don't know if it is good, since we have noisy labels?<br>\nCheers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1178096,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "01/30/2021 16:40:57",
          "content": "<p>I used 512*512 using TPU. On GPU you will run out of memory in that. Image size of 448 is also fine.  Image size between 384-512 are good enough for the training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1179889,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "02/01/2021 00:03:43",
      "content": "<p>All of what you mentioned also do work in my case</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1181352,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/01/2021 20:02:26",
      "content": "<p>Removing the noise from the dataset reduced my lb score which tells that test data is also noisy on similar lines.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1181598,
      "author_name": "huihui2013bupt",
      "author_url": "",
      "post_date": "02/02/2021 02:18:52",
      "content": "<p>25 Epoch and 5 folds mean every fold train  5 Epoch ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1181750,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "02/02/2021 05:12:08",
          "content": "<p>Dataset divided into 5 folds and each fold trained for 25 epoch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1181999,
      "author_name": "huihui2013bupt",
      "author_url": "",
      "post_date": "02/02/2021 08:47:16",
      "content": "<p>Oh thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1199495,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/13/2021 20:56:25",
      "content": "<p>For bi tempered loss refer to link <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1199507,
      "author_name": "jcesquiveld",
      "author_url": "",
      "post_date": "02/13/2021 21:09:15",
      "content": "<p>In my modest opinion, to achieve a reasonable (not the best, which is beyond my reach) score at the LB (above 0.902) it takes:</p>\n<ul>\n<li>to train a couple, three or four CV models with pretrained weights from: ResNet variants (ssl|swsl)<em>resnext(50|101)_32x(4|8|16), (tf| )_efficientnet</em>(b3|b4|b4)_ns) with 2020 dataset, maybe upsized with 2018 training dataset, and perhaps with extra and test images from 2018 competition</li>\n<li>Trying different ensemble them in differente ways: excluding worst models, combining good predictors for classes 0,1,2 and 4, etc.</li>\n<li>hyperparameter tuning transformations and number of iterations for TTA. In my experience, very light TTA works best (random crop, H and V flips).</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1199518,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "02/13/2021 21:18:56",
          "content": "<p>Thanks, I did what u mentioned in point 1 and partially in point 2 but i have not done tta, will try it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1176169": "**The following things seem to work:**\n1. Model EfficientNet seems to work better\n2. 25 Epoch and 5 folds is giving good result\n3. Cutmix is also helping in increasing  the lb and cv.\n4. Individual folds might give more score but using all folds give more stability.\n5. Old competition data gives little rise in lb.\nTo get old data in tf records click [here](https://www.kaggle.com/vickygoyal/oldcassavadata). The details of the dataset is  [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490)\n6. Ensembling gives rise in lb.",
    "1177434": "I have yet to try cut mix,  I am using early stopping, never getting past 20 epochs as it stops early. \n\nWhat weights do you use (imagenet/noisy-student)?\nI saw people saying bitempered logistic loss was giving good results but that is not the case for me. What loss do you use?",
    "1177483": "I used noisy-student and same old categorical_crossentropy loss.",
    "1177879": "I want to ask, what image size you use? I used 448x448 (initial results were 384x384) which increased my CV, however I  didn't try 512x512 since I ran out of memory (with batch size of 32). Maybe i need to lower the batch size, but i don't know if it is good, since we have noisy labels?\nCheers.",
    "1178096": "I used 512*512 using TPU. On GPU you will run out of memory in that. Image size of 448 is also fine.  Image size between 384-512 are good enough for the training.",
    "1179889": "All of what you mentioned also do work in my case",
    "1181352": "Removing the noise from the dataset reduced my lb score which tells that test data is also noisy on similar lines.",
    "1181598": "25 Epoch and 5 folds mean every fold train  5 Epoch ?",
    "1181750": "Dataset divided into 5 folds and each fold trained for 25 epoch.",
    "1181999": "Oh thanks.",
    "1199495": "For bi tempered loss refer to link [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911)",
    "1199507": "In my modest opinion, to achieve a reasonable (not the best, which is beyond my reach) score at the LB (above 0.902) it takes:\n- to train a couple, three or four CV models with pretrained weights from: ResNet variants (ssl|swsl)_resnext(50|101)_32x(4|8|16), (tf| )_efficientnet_(b3|b4|b4)_ns) with 2020 dataset, maybe upsized with 2018 training dataset, and perhaps with extra and test images from 2018 competition\n- Trying different ensemble them in differente ways: excluding worst models, combining good predictors for classes 0,1,2 and 4, etc.\n- hyperparameter tuning transformations and number of iterations for TTA. In my experience, very light TTA works best (random crop, H and V flips).",
    "1199518": "Thanks, I did what u mentioned in point 1 and partially in point 2 but i have not done tta, will try it."
  },
  "source": "meta"
}