{
  "id": 203594,
  "title": "Sharing some improvements and experiments",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203594",
  "author_name": "",
  "post_date": "2020-12-15T22:02:11.028161500Z",
  "votes": 180,
  "comment_count": 71,
  "views": 0,
  "content": "<p>Recently I have been doing many experiments to improve my first notebook <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training\" target=\"_blank\">Cassava Leaf Disease - TPU Tensorflow - Training\n</a>, and wanted to share my results and maybe have some discussions about what some of you think about.</p>\n<p>As part of the <code>TPU star</code> program, I got access to <code>TPU-v2 Pods</code> for a few weeks and wanted to make sure I was optimizing the runtime, so I also improved the Tensorflow pipeline.</p>\n<p>Here is the link for the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">training notebook</a> and here is the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference\" target=\"_blank\">inference notebook</a>.</p>\n<p>5-Fold CV was ranging between <code>0.890</code> and <code>0.898</code></p>\n<hr>\n<h4>Improvements</h4>\n<ul>\n<li><strong>Custom training loop</strong>: Using a custom training loop greatly improves the training time and resource usage, if you are using Tensorflow I highly recommend to use it, your training time will be a lot less, it also makes it easier to do some nice tricks during training.</li>\n<li><strong>Maximize MXU and minimize Idle time</strong>: I have made a few adjustments to the Tensorflow pipeline to improve performance.</li>\n</ul>\n<hr>\n<h4>Experiments</h4>\n<p><strong>Small improvements using external data (2019 competition)</strong>: to be honest, I was expecting better results using data from the 2019 competition, I remember that at the <code>APTOS</code> competition data from the previous competition was key, at the <code>Melanoma</code> competition it also helped more.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F0568c1bf1afe589c7ac5a3e058637053%2FScreenshot%20from%202020-12-15%2019-03-28.png?generation=1608069848078572&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>Small improvements from using CCE label smoothing</strong>: This was expected since the competition has some noise, as reported by other users, the thing was that here I got slighter better results by using <code>label smoothing of 0.3</code>, for me this value is a little high.</p>\n<hr>\n<p><strong>Small improvements from using CutOut</strong>: Using <code>CutOut</code> improved generalization, I am using a custom <code>CutOut</code> function and have not tweaked its parameters a lot.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2Fd15c79d66df79f98360a231be5f6668c%2FScreenshot%20from%202020-12-15%2018-59-53.png?generation=1608069787618497&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>Small improvements from oversampling classes 0, 1, 2, and 4</strong>:  Oversampling the dataset with the minority classes improved the training, interestingly this worked a lot better than using class weights.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F280fbb4d2f8dac94a356112f823bdf62%2FScreenshot%20from%202020-12-15%2019-00-04.png?generation=1608069808112772&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>Small improvements from keeping batch normalization layers frozen</strong>:  This was curious for me, I tried to do the same during the <code>Melanoma</code> competition,  but there, it did not work, maybe here the plant's images are closer to the <code>Imagenet</code> data?</p>\n<p>Keeping the <code>Batch Normalization</code> layers frozen during fine-tuning is recommended on the <a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">Image classification via fine-tuning with EfficientNet</a> at the Keras repository.</p>\n<hr>\n<p><strong>No relevant improvements by using class weights</strong>: I have tried generating class weights in different ways, but got no improvements from it, did not matter using it within a custom training loop or the regular training.</p>\n<hr>\n<p><strong>No relevant improvements by using MixUp</strong>: I was not expecting it to work, but it could also be a problem with my implementation.</p>\n<hr>\n<p><strong>No relevant improvements by using different backbones</strong>: I have tried the other <code>EfficientNet</code> backbones and some others as well, but for me, I got better results with <code>EfficientNet</code> <code>B3</code> to <code>B6</code>.</p>\n<hr>\n<p><strong>Worse performance by using different image resolution even the default EfficientNet input size</strong>: At my experiments, the best results came from image resolution <code>512x512</code>.</p>\n<hr>\n<p><strong>Was not able to make progressive unfreezing work</strong>: The concept is cool, and I tried to unfreeze a block of <code>EfficientNet</code> at a time but did not get improvements.</p>\n<hr>\n<p><strong>Changing the learning rate batch-wise seems more efficient than epoch wise, especially for the warm-up phase</strong>: The thing with changing the learning rate epoch-wise is that the model will train with a single learning rate for the complete epoch, changing it batch-wise gives a \"continuous\" change of the learning rate, this is interesting for the warm-up phase.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F9b2c734d02550c9e43c07e4471db12d8%2FScreenshot%20from%202020-12-15%2019-00-13.png?generation=1608069761044709&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>I will keep doing some more experiments during the week and will update this thread with the results.</p>",
  "messages": [
    {
      "id": "1113996",
      "postDate": "12/15/2020 22:02:11",
      "content": "<p>Recently I have been doing many experiments to improve my first notebook <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training\" target=\"_blank\">Cassava Leaf Disease - TPU Tensorflow - Training\n</a>, and wanted to share my results and maybe have some discussions about what some of you think about.</p>\n<p>As part of the <code>TPU star</code> program, I got access to <code>TPU-v2 Pods</code> for a few weeks and wanted to make sure I was optimizing the runtime, so I also improved the Tensorflow pipeline.</p>\n<p>Here is the link for the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">training notebook</a> and here is the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference\" target=\"_blank\">inference notebook</a>.</p>\n<p>5-Fold CV was ranging between <code>0.890</code> and <code>0.898</code></p>\n<hr>\n<h4>Improvements</h4>\n<ul>\n<li><strong>Custom training loop</strong>: Using a custom training loop greatly improves the training time and resource usage, if you are using Tensorflow I highly recommend to use it, your training time will be a lot less, it also makes it easier to do some nice tricks during training.</li>\n<li><strong>Maximize MXU and minimize Idle time</strong>: I have made a few adjustments to the Tensorflow pipeline to improve performance.</li>\n</ul>\n<hr>\n<h4>Experiments</h4>\n<p><strong>Small improvements using external data (2019 competition)</strong>: to be honest, I was expecting better results using data from the 2019 competition, I remember that at the <code>APTOS</code> competition data from the previous competition was key, at the <code>Melanoma</code> competition it also helped more.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F0568c1bf1afe589c7ac5a3e058637053%2FScreenshot%20from%202020-12-15%2019-03-28.png?generation=1608069848078572&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>Small improvements from using CCE label smoothing</strong>: This was expected since the competition has some noise, as reported by other users, the thing was that here I got slighter better results by using <code>label smoothing of 0.3</code>, for me this value is a little high.</p>\n<hr>\n<p><strong>Small improvements from using CutOut</strong>: Using <code>CutOut</code> improved generalization, I am using a custom <code>CutOut</code> function and have not tweaked its parameters a lot.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2Fd15c79d66df79f98360a231be5f6668c%2FScreenshot%20from%202020-12-15%2018-59-53.png?generation=1608069787618497&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>Small improvements from oversampling classes 0, 1, 2, and 4</strong>:  Oversampling the dataset with the minority classes improved the training, interestingly this worked a lot better than using class weights.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F280fbb4d2f8dac94a356112f823bdf62%2FScreenshot%20from%202020-12-15%2019-00-04.png?generation=1608069808112772&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>Small improvements from keeping batch normalization layers frozen</strong>:  This was curious for me, I tried to do the same during the <code>Melanoma</code> competition,  but there, it did not work, maybe here the plant's images are closer to the <code>Imagenet</code> data?</p>\n<p>Keeping the <code>Batch Normalization</code> layers frozen during fine-tuning is recommended on the <a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">Image classification via fine-tuning with EfficientNet</a> at the Keras repository.</p>\n<hr>\n<p><strong>No relevant improvements by using class weights</strong>: I have tried generating class weights in different ways, but got no improvements from it, did not matter using it within a custom training loop or the regular training.</p>\n<hr>\n<p><strong>No relevant improvements by using MixUp</strong>: I was not expecting it to work, but it could also be a problem with my implementation.</p>\n<hr>\n<p><strong>No relevant improvements by using different backbones</strong>: I have tried the other <code>EfficientNet</code> backbones and some others as well, but for me, I got better results with <code>EfficientNet</code> <code>B3</code> to <code>B6</code>.</p>\n<hr>\n<p><strong>Worse performance by using different image resolution even the default EfficientNet input size</strong>: At my experiments, the best results came from image resolution <code>512x512</code>.</p>\n<hr>\n<p><strong>Was not able to make progressive unfreezing work</strong>: The concept is cool, and I tried to unfreeze a block of <code>EfficientNet</code> at a time but did not get improvements.</p>\n<hr>\n<p><strong>Changing the learning rate batch-wise seems more efficient than epoch wise, especially for the warm-up phase</strong>: The thing with changing the learning rate epoch-wise is that the model will train with a single learning rate for the complete epoch, changing it batch-wise gives a \"continuous\" change of the learning rate, this is interesting for the warm-up phase.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F9b2c734d02550c9e43c07e4471db12d8%2FScreenshot%20from%202020-12-15%2019-00-13.png?generation=1608069761044709&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>I will keep doing some more experiments during the week and will update this thread with the results.</p>",
      "rawMarkdown": "Recently I have been doing many experiments to improve my first notebook [Cassava Leaf Disease - TPU Tensorflow - Training\n](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training), and wanted to share my results and maybe have some discussions about what some of you think about.\n\nAs part of the `TPU star` program, I got access to `TPU-v2 Pods` for a few weeks and wanted to make sure I was optimizing the runtime, so I also improved the Tensorflow pipeline.\n\nHere is the link for the [training notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods) and here is the [inference notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference).\n\n5-Fold CV was ranging between `0.890` and `0.898`\n\n___\n\n#### Improvements\n\n- **Custom training loop**: Using a custom training loop greatly improves the training time and resource usage, if you are using Tensorflow I highly recommend to use it, your training time will be a lot less, it also makes it easier to do some nice tricks during training.\n- **Maximize MXU and minimize Idle time**: I have made a few adjustments to the Tensorflow pipeline to improve performance.\n\n___\n\n#### Experiments\n\n**Small improvements using external data (2019 competition)**: to be honest, I was expecting better results using data from the 2019 competition, I remember that at the `APTOS` competition data from the previous competition was key, at the `Melanoma` competition it also helped more.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F0568c1bf1afe589c7ac5a3e058637053%2FScreenshot%20from%202020-12-15%2019-03-28.png?generation=1608069848078572&alt=media)\n___\n**Small improvements from using CCE label smoothing**: This was expected since the competition has some noise, as reported by other users, the thing was that here I got slighter better results by using `label smoothing of 0.3`, for me this value is a little high.\n___\n**Small improvements from using CutOut**: Using `CutOut` improved generalization, I am using a custom `CutOut` function and have not tweaked its parameters a lot.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2Fd15c79d66df79f98360a231be5f6668c%2FScreenshot%20from%202020-12-15%2018-59-53.png?generation=1608069787618497&alt=media)\n___\n\n**Small improvements from oversampling classes 0, 1, 2, and 4**:  Oversampling the dataset with the minority classes improved the training, interestingly this worked a lot better than using class weights.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F280fbb4d2f8dac94a356112f823bdf62%2FScreenshot%20from%202020-12-15%2019-00-04.png?generation=1608069808112772&alt=media)\n___\n**Small improvements from keeping batch normalization layers frozen**:  This was curious for me, I tried to do the same during the `Melanoma` competition,  but there, it did not work, maybe here the plant's images are closer to the `Imagenet` data?\n\nKeeping the `Batch Normalization` layers frozen during fine-tuning is recommended on the [Image classification via fine-tuning with EfficientNet](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/) at the Keras repository.\n___\n**No relevant improvements by using class weights**: I have tried generating class weights in different ways, but got no improvements from it, did not matter using it within a custom training loop or the regular training.\n___\n**No relevant improvements by using MixUp**: I was not expecting it to work, but it could also be a problem with my implementation.\n___\n**No relevant improvements by using different backbones**: I have tried the other `EfficientNet` backbones and some others as well, but for me, I got better results with `EfficientNet` `B3` to `B6`.\n___\n**Worse performance by using different image resolution even the default EfficientNet input size**: At my experiments, the best results came from image resolution `512x512`.\n___\n**Was not able to make progressive unfreezing work**: The concept is cool, and I tried to unfreeze a block of `EfficientNet` at a time but did not get improvements.\n___\n**Changing the learning rate batch-wise seems more efficient than epoch wise, especially for the warm-up phase**: The thing with changing the learning rate epoch-wise is that the model will train with a single learning rate for the complete epoch, changing it batch-wise gives a \"continuous\" change of the learning rate, this is interesting for the warm-up phase.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F9b2c734d02550c9e43c07e4471db12d8%2FScreenshot%20from%202020-12-15%2019-00-13.png?generation=1608069761044709&alt=media)\n___\n\nI will keep doing some more experiments during the week and will update this thread with the results.",
      "votes": null
    },
    {
      "id": "1115081",
      "postDate": "12/16/2020 01:04:44",
      "content": "<p>Hey, firstly thanks for sharing as this is a very nicely written post. My goal for this competition was to learn how to build a custom TPU training loop, but all my efforts lead to a much slower training time, even if i don't do augmentations inside the training loop. How much faster have you been able to achieve comparing to vanilla keras?</p>",
      "rawMarkdown": "Hey, firstly thanks for sharing as this is a very nicely written post. My goal for this competition was to learn how to build a custom TPU training loop, but all my efforts lead to a much slower training time, even if i don't do augmentations inside the training loop. How much faster have you been able to achieve comparing to vanilla keras?",
      "votes": null
    },
    {
      "id": "1115113",
      "postDate": "12/16/2020 01:59:01",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> , you're welcome.</p>\n<p>Custom training loops can be tricky for the first time, but after some hacking you get comfortable with it, I recommend you to fork my kernel and try to play a little.</p>\n<p>About the training time, if you look at the <code>version 22</code> of <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training#Model-evaluation\" target=\"_blank\">this kernel</a><br>\nI am using EfficientNetB4 and the training time is something like this:</p>\n<pre><code>Epoch 1/30\n133/133 - 106s - loss: 1.7036 - sparse_categorical_accuracy: 0.1527 - val_loss: 1.5900 - val_sparse_categorical_accuracy: 0.2801 - lr: 1.0000e-08\nEpoch 2/30\n133/133 - 71s - loss: 0.8568 - sparse_categorical_accuracy: 0.6930 - val_loss: 0.5133 - val_sparse_categorical_accuracy: 0.8229 - lr: 8.0007e-05\n</code></pre>\n<p>On the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">kernel</a> with custom training loop I also use EfficientNetB4 and the training time is:</p>\n<pre><code>EPOCH 1/10\ntime: 296.3s loss: 1.2660 accuracy: 0.6615 val_loss: 0.6311 val_accuracy: 0.8638 lr: 0.00032\nEPOCH 2/10\ntime: 38.9s loss: 1.1040 accuracy: 0.8217 val_loss: 0.5619 val_accuracy: 0.8760 lr: 0.0003104\n</code></pre>\n<p>As you can see the first epoch is slower, but after that, the epochs are much faster.<br>\nNote that the first uses TPUv3 and the second uses the TPUv2 Pod, I will try to run the second kernel with a regular Keras fitting, then I will post here the results.</p>",
      "rawMarkdown": "Hey @capiru , you're welcome.\n\nCustom training loops can be tricky for the first time, but after some hacking you get comfortable with it, I recommend you to fork my kernel and try to play a little.\n\nAbout the training time, if you look at the `version 22` of [this kernel](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training#Model-evaluation)\nI am using EfficientNetB4 and the training time is something like this:\n```\nEpoch 1/30\n133/133 - 106s - loss: 1.7036 - sparse_categorical_accuracy: 0.1527 - val_loss: 1.5900 - val_sparse_categorical_accuracy: 0.2801 - lr: 1.0000e-08\nEpoch 2/30\n133/133 - 71s - loss: 0.8568 - sparse_categorical_accuracy: 0.6930 - val_loss: 0.5133 - val_sparse_categorical_accuracy: 0.8229 - lr: 8.0007e-05\n```\n\nOn the [kernel](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods) with custom training loop I also use EfficientNetB4 and the training time is:\n```\nEPOCH 1/10\ntime: 296.3s loss: 1.2660 accuracy: 0.6615 val_loss: 0.6311 val_accuracy: 0.8638 lr: 0.00032\nEPOCH 2/10\ntime: 38.9s loss: 1.1040 accuracy: 0.8217 val_loss: 0.5619 val_accuracy: 0.8760 lr: 0.0003104\n```\n\nAs you can see the first epoch is slower, but after that, the epochs are much faster.\nNote that the first uses TPUv3 and the second uses the TPUv2 Pod, I will try to run the second kernel with a regular Keras fitting, then I will post here the results.",
      "votes": null
    },
    {
      "id": "1115218",
      "postDate": "12/16/2020 05:09:51",
      "content": "<p>Thank you! Finally someone raised the issue for Batch Normalization layers but the point here I feel is, if you are doing finetuning then only freezing of BN matters. If you are doing training from scratch then you must not freeze BN layers. I have a complete thread <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203095\" target=\"_blank\">here</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F29c16e38ae51f4bb709114afca3e258b%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095363441873&amp;alt=media\" alt=\"pic\"></p>",
      "rawMarkdown": "Thank you! Finally someone raised the issue for Batch Normalization layers but the point here I feel is, if you are doing finetuning then only freezing of BN matters. If you are doing training from scratch then you must not freeze BN layers. I have a complete thread [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203095)\n\n![pic](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F29c16e38ae51f4bb709114afca3e258b%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095363441873&alt=media)",
      "votes": null
    },
    {
      "id": "1115267",
      "postDate": "12/16/2020 06:37:24",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": null
    },
    {
      "id": "1115613",
      "postDate": "12/16/2020 12:16:55",
      "content": "<p>You're welcome <a href=\"https://www.kaggle.com/slowlearnermack\" target=\"_blank\">@slowlearnermack</a> </p>",
      "rawMarkdown": "You're welcome @slowlearnermack",
      "votes": null
    },
    {
      "id": "1115631",
      "postDate": "12/16/2020 12:32:24",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> those are some interesting points, here is what I think.</p>\n<p><strong>What <code>Batch Normalization</code> actually do?</strong>: In simple words, it would be what you pointed, \"Learn the mean and variance of the input data\".</p>\n<p>So, if you are training from scratch you should keep them trainable so they can learn the data distribution.</p>\n<p>But for fine-tunning, I can see two cases:</p>\n<ol>\n<li><strong>Fine-tuning for a domain close to the pre-train domain</strong> (e.g. pre-train on <code>Imagenet</code> but fine-tune to <code>Cat vs Dog</code>), those two domains are very close, so it is intuitive to think that the distribution learned by the pre-train step will be useful for the fine-tune task.</li>\n<li><strong>Fine-tuning for a domain very different to the pre-train domain</strong> (e.g. pre-train on <code>Imagenet</code> but fine-tune to <code>Melanoma classification</code>), there are no images similar to Melanomas on the <code>Imagenet</code> data, so it would be intuitive to think that the data distribution (mean and variance) from <code>Imagenet</code> won't help as much as the <code>Cat vs Dog</code> task.</li>\n</ol>\n<p>Here in this competition, the task domain may be closer to the <code>Imagenet</code>, and keeping the <code>Batch Normalization</code> layers frozen may help.</p>\n<p>As I mentioned, at the Melanoma competition I tried the same but it did not work, and here there was an improvement but it was not too big.</p>",
      "rawMarkdown": "Hey @harveenchadha those are some interesting points, here is what I think.\n\n**What `Batch Normalization` actually do?**: In simple words, it would be what you pointed, \"Learn the mean and variance of the input data\".\n\nSo, if you are training from scratch you should keep them trainable so they can learn the data distribution.\n\nBut for fine-tunning, I can see two cases:\n1. **Fine-tuning for a domain close to the pre-train domain** (e.g. pre-train on `Imagenet` but fine-tune to `Cat vs Dog`), those two domains are very close, so it is intuitive to think that the distribution learned by the pre-train step will be useful for the fine-tune task.\n2. **Fine-tuning for a domain very different to the pre-train domain** (e.g. pre-train on `Imagenet` but fine-tune to `Melanoma classification`), there are no images similar to Melanomas on the `Imagenet` data, so it would be intuitive to think that the data distribution (mean and variance) from `Imagenet` won't help as much as the `Cat vs Dog` task.\n\nHere in this competition, the task domain may be closer to the `Imagenet`, and keeping the `Batch Normalization` layers frozen may help.\n\nAs I mentioned, at the Melanoma competition I tried the same but it did not work, and here there was an improvement but it was not too big.",
      "votes": null
    },
    {
      "id": "1115655",
      "postDate": "12/16/2020 12:54:43",
      "content": "<p>I think whatever you just said makes a lot of sense. To check if the input data's mean and variance is same as imagenet, we can calculate that as well before making a decision. Right?</p>",
      "rawMarkdown": "I think whatever you just said makes a lot of sense. To check if the input data's mean and variance is same as imagenet, we can calculate that as well before making a decision. Right?",
      "votes": null
    },
    {
      "id": "1115877",
      "postDate": "12/16/2020 16:48:53",
      "content": "<p>Nice observations! I wonder how you did oversampling here? Using class weights &gt; 1 or just repeating the images in the dataset? I've never done anything like that, so sorry if this is too naive question. Thanks</p>",
      "rawMarkdown": "Nice observations! I wonder how you did oversampling here? Using class weights > 1 or just repeating the images in the dataset? I've never done anything like that, so sorry if this is too naive question. Thanks",
      "votes": null
    },
    {
      "id": "1115952",
      "postDate": "12/16/2020 17:35:12",
      "content": "<p><a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> this may be a good idea, but in practice, it may not be this simple.</p>\n<p>It may be a little tricky to figure out if the distribution from one dataset is close enough to 'Imagenet'.</p>\n<p>I would suggest just doing an exploration over the dataset, and see if the images are similar to <code>Imagenet</code>, then it is probably worth it to just try both options anyway and see if there is any improvement 😄.</p>",
      "rawMarkdown": "harveenchadha this may be a good idea, but in practice, it may not be this simple.\n\nIt may be a little tricky to figure out if the distribution from one dataset is close enough to 'Imagenet'.\n\nI would suggest just doing an exploration over the dataset, and see if the images are similar to `Imagenet`, then it is probably worth it to just try both options anyway and see if there is any improvement 😄.",
      "votes": null
    },
    {
      "id": "1116005",
      "postDate": "12/16/2020 18:24:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> I actually did both options, but <code>oversampling</code> refers to <code>increasing the data</code>, in the notebook at the beginning of the training loop you can see that I am <code>adding more data</code> to the minority classes.</p>\n<p>The second option is just <code>increasing class weights</code> for the minority classes, I got no improvements by doing it here, but in theory, the results should be similar.</p>\n<p>Both are very easy to implement, especially if you are using a regular Keras <code>.fit</code> training, but you can feel free to fork the kernel and do some experimentations.</p>",
      "rawMarkdown": "Hi @kaushal2896 I actually did both options, but `oversampling` refers to `increasing the data`, in the notebook at the beginning of the training loop you can see that I am `adding more data` to the minority classes.\n\nThe second option is just `increasing class weights` for the minority classes, I got no improvements by doing it here, but in theory, the results should be similar.\n\nBoth are very easy to implement, especially if you are using a regular Keras `.fit` training, but you can feel free to fork the kernel and do some experimentations.",
      "votes": null
    },
    {
      "id": "1116026",
      "postDate": "12/16/2020 18:51:43",
      "content": "<p>Okay. Thanks for the information!</p>",
      "rawMarkdown": "Okay. Thanks for the information!",
      "votes": null
    },
    {
      "id": "1116881",
      "postDate": "12/17/2020 14:38:03",
      "content": "<p>Thanks for sharing!!! Are you using the crossentropy loss or some custom loss function?</p>",
      "rawMarkdown": "Thanks for sharing!!! Are you using the crossentropy loss or some custom loss function?",
      "votes": null
    },
    {
      "id": "1116915",
      "postDate": "12/17/2020 15:01:16",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> <br>\nPlease I am new to Kaggle competitions, I have been trying to submit but having difficulties. It says I can't submit with the internet turned on, but when I turn off internet it does not generate output.</p>\n<p>Any help will be appreciated.</p>",
      "rawMarkdown": "Hello @dimitreoliveira \nPlease I am new to Kaggle competitions, I have been trying to submit but having difficulties. It says I can't submit with the internet turned on, but when I turn off internet it does not generate output.\n\nAny help will be appreciated.",
      "votes": null
    },
    {
      "id": "1116963",
      "postDate": "12/17/2020 15:40:16",
      "content": "<p>hello<br>\nyou can make two notebooks; one for training and other for inference, while training you can use the internet and save your weights, later you can use these weights in offline mode for predicting on the test. </p>",
      "rawMarkdown": "hello\nyou can make two notebooks; one for training and other for inference, while training you can use the internet and save your weights, later you can use these weights in offline mode for predicting on the test.",
      "votes": null
    },
    {
      "id": "1116979",
      "postDate": "12/17/2020 16:01:22",
      "content": "<p>When I eventually got a submission file submitted. It says submission file not found. </p>",
      "rawMarkdown": "When I eventually got a submission file submitted. It says submission file not found.",
      "votes": null
    },
    {
      "id": "1116981",
      "postDate": "12/17/2020 16:02:40",
      "content": "<p>See this it’s frustrating 😭</p>",
      "rawMarkdown": "See this it’s frustrating 😭",
      "votes": null
    },
    {
      "id": "1117238",
      "postDate": "12/17/2020 20:27:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> , I am using basic <code>Categorical crossentropy</code> with <code>label smoothing</code>, it is defined at the training loop by <code>loss_fn = losses.categorical_crossentropy</code>.</p>",
      "rawMarkdown": "Hi @mrinath , I am using basic `Categorical crossentropy` with `label smoothing`, it is defined at the training loop by `loss_fn = losses.categorical_crossentropy`.",
      "votes": null
    },
    {
      "id": "1117240",
      "postDate": "12/17/2020 20:29:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jbolaifa\" target=\"_blank\">@jbolaifa</a> , as pointed by <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> , for this competition you need a notebook that does not use the internet to make the submission.</p>\n<p>If you just want to try making the submission you can fork this <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference\" target=\"_blank\">notebook</a> and submit.</p>",
      "rawMarkdown": "Hi @jbolaifa , as pointed by @mrinath , for this competition you need a notebook that does not use the internet to make the submission.\n\nIf you just want to try making the submission you can fork this [notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference) and submit.",
      "votes": null
    },
    {
      "id": "1119910",
      "postDate": "12/20/2020 12:52:15",
      "content": "<p>Great job, thanks for sharing! </p>",
      "rawMarkdown": "Great job, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1120046",
      "postDate": "12/20/2020 15:04:47",
      "content": "<p>Thanks for sharing, appreciate it!</p>",
      "rawMarkdown": "Thanks for sharing, appreciate it!",
      "votes": null
    },
    {
      "id": "1120625",
      "postDate": "12/21/2020 01:44:16",
      "content": "<p>Hello, I would like to ask whether the use of the 2019 data set has improved LB or CV?<br>\nI have used it before, but there seems to be no change, do I need to take more processing？</p>",
      "rawMarkdown": "Hello, I would like to ask whether the use of the 2019 data set has improved LB or CV?\nI have used it before, but there seems to be no change, do I need to take more processing？",
      "votes": null
    },
    {
      "id": "1120639",
      "postDate": "12/21/2020 02:11:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sjtuyxc\" target=\"_blank\">@sjtuyxc</a> , I have experimented with 2019 data, and the training kernel uses it, but as I mentioned in the <code>Experiments</code> section: </p>\n<pre><code>Small improvements using external data (2019 competition): to be honest, I was expecting better results using data from the 2019 competition, I remember that at the APTOS competition data from the previous competition was key, at the Melanoma competition it also helped more.\n</code></pre>\n<p>I decided to leave the <code>2019</code> for training since it has similar quality and may help with generalization, but in terms of score, the improvements were very minimal, so far I have not seen people reporting great improvements using the external data.</p>",
      "rawMarkdown": "Hi @sjtuyxc , I have experimented with 2019 data, and the training kernel uses it, but as I mentioned in the `Experiments` section: \n\n```\nSmall improvements using external data (2019 competition): to be honest, I was expecting better results using data from the 2019 competition, I remember that at the APTOS competition data from the previous competition was key, at the Melanoma competition it also helped more.\n```\n\nI decided to leave the `2019` for training since it has similar quality and may help with generalization, but in terms of score, the improvements were very minimal, so far I have not seen people reporting great improvements using the external data.",
      "votes": null
    },
    {
      "id": "1120960",
      "postDate": "12/21/2020 08:45:25",
      "content": "<p>Hi, I'm interested in this discussion too. I have experimented same small change or no change at all in both CV and LB score.</p>\n<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> thanks for your notebooks and work. Actually I'm using your tfrecords for 2019 data, but I cannot find the notebook to generate them. Did you check for duplicates?</p>\n<p>I was wondering if there is some way of using 2019 data in this competition also taking into account of unlabeled images and maybe go for a semi-supervised approach.</p>",
      "rawMarkdown": "Hi, I'm interested in this discussion too. I have experimented same small change or no change at all in both CV and LB score.\n\n@dimitreoliveira thanks for your notebooks and work. Actually I'm using your tfrecords for 2019 data, but I cannot find the notebook to generate them. Did you check for duplicates?\n\nI was wondering if there is some way of using 2019 data in this competition also taking into account of unlabeled images and maybe go for a semi-supervised approach.",
      "votes": null
    },
    {
      "id": "1120968",
      "postDate": "12/21/2020 08:50:02",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> did you used any custom loss on TPU?</p>\n<p>Actually I'm trying to implement my own loss function (that works fine if I run training on GPU) but experiencing some weird problem when training on TPU.</p>\n<p>My loss function uses tf.cond and tf.while_loop, but everything should be TPU compatible, right? Anyway I'm receiving some compilation error from TPU without any clear explanation, something like:</p>\n<p><code>FakeParam OpKernel not implemented on XLA_JIT</code></p>\n<p>Did you ever experienced a similar error?</p>",
      "rawMarkdown": "dimitreoliveira did you used any custom loss on TPU?\n\nActually I'm trying to implement my own loss function (that works fine if I run training on GPU) but experiencing some weird problem when training on TPU.\n\nMy loss function uses tf.cond and tf.while_loop, but everything should be TPU compatible, right? Anyway I'm receiving some compilation error from TPU without any clear explanation, something like:\n\n`FakeParam OpKernel not implemented on XLA_JIT`\n\nDid you ever experienced a similar error?",
      "votes": null
    },
    {
      "id": "1121219",
      "postDate": "12/21/2020 13:05:27",
      "content": "<p>Thanks~I got it！</p>",
      "rawMarkdown": "Thanks~I got it！",
      "votes": null
    },
    {
      "id": "1121283",
      "postDate": "12/21/2020 14:18:19",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> I used a \"custom loss\" to apply class weights to <code>categorical cross-entropy</code> and it worked, by your error, I guess you might be using some function that can't be compiled by the <code>XLA</code> compiler (TPU uses XLA), also note that some functions are not available to TPUs, for example, some paddings.</p>",
      "rawMarkdown": "Hey @lazcoder I used a \"custom loss\" to apply class weights to `categorical cross-entropy` and it worked, by your error, I guess you might be using some function that can't be compiled by the `XLA` compiler (TPU uses XLA), also note that some functions are not available to TPUs, for example, some paddings.",
      "votes": null
    },
    {
      "id": "1121287",
      "postDate": "12/21/2020 14:22:38",
      "content": "<p>You're welcome <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , I have not made the notebook to create the external data public because it is essentially the same code to create the regular <code>TFRecords</code>, but if you want to check out the code to evaluate or see how it is done I can make it public.</p>\n<p>About the unlabeled data, I was also thinking to use them as <code>semi-supervised</code>, as you said, but I had no time yet, it is probably worth it to try.</p>",
      "rawMarkdown": "You're welcome @lazcoder , I have not made the notebook to create the external data public because it is essentially the same code to create the regular `TFRecords`, but if you want to check out the code to evaluate or see how it is done I can make it public.\n\nAbout the unlabeled data, I was also thinking to use them as `semi-supervised`, as you said, but I had no time yet, it is probably worth it to try.",
      "votes": null
    },
    {
      "id": "1121380",
      "postDate": "12/21/2020 15:53:12",
      "content": "<p>thanks for sharing! Have you tried using a customized cost function? </p>",
      "rawMarkdown": "thanks for sharing! Have you tried using a customized cost function?",
      "votes": null
    },
    {
      "id": "1121701",
      "postDate": "12/21/2020 20:39:21",
      "content": "<p>Okay, thanks <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>.</p>\n<p>Did you used all the dataset from 2019 competition (i.e. both training and test set)?</p>",
      "rawMarkdown": "Okay, thanks @dimitreoliveira.\n\nDid you used all the dataset from 2019 competition (i.e. both training and test set)?",
      "votes": null
    },
    {
      "id": "1121766",
      "postDate": "12/21/2020 22:24:32",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , I have used only the training data.</p>",
      "rawMarkdown": "Hey @lazcoder , I have used only the training data.",
      "votes": null
    },
    {
      "id": "1121770",
      "postDate": "12/21/2020 22:26:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/denizyigit\" target=\"_blank\">@denizyigit</a> , For this notebook, the only custom loss was <code>categorical cross-entropy</code> with class weights, but I am doing some experiments with <code>supervised contrastive learning</code>, I will release the notebook in a day or two.</p>\n<p>Have you tried with success any custom loss?</p>",
      "rawMarkdown": "Hi @denizyigit , For this notebook, the only custom loss was `categorical cross-entropy` with class weights, but I am doing some experiments with `supervised contrastive learning`, I will release the notebook in a day or two.\n\nHave you tried with success any custom loss?",
      "votes": null
    },
    {
      "id": "1121812",
      "postDate": "12/21/2020 23:36:11",
      "content": "<p>I'm trying to implement bi-tempered logistic loss as from <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">this discussion</a> it seems to give better performances on noisy datasets. In my experiments anyway I don't get so much improvements. </p>\n<p>I adapted <a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">pytorch implementation</a> from, but the <a href=\"https://github.com/google/bi-tempered-loss\" target=\"_blank\">original version</a> is in tensorflow</p>\n<p>My version seems slightly faster, but when changing temperatures the training get stuck and no error is thrown 👍</p>",
      "rawMarkdown": "I'm trying to implement bi-tempered logistic loss as from [this discussion](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017) it seems to give better performances on noisy datasets. In my experiments anyway I don't get so much improvements. \n\nI adapted [pytorch implementation](https://github.com/mlpanda/bi-tempered-loss-pytorch) from, but the [original version](https://github.com/google/bi-tempered-loss) is in tensorflow\n\nMy version seems slightly faster, but when changing temperatures the training get stuck and no error is thrown 👍",
      "votes": null
    },
    {
      "id": "1121839",
      "postDate": "12/22/2020 00:23:52",
      "content": "<p>Nice <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , about the behaviour, if the training freezes maybe your implementation is turning the gradients to 0.</p>",
      "rawMarkdown": "Nice @lazcoder , about the behaviour, if the training freezes maybe your implementation is turning the gradients to 0.",
      "votes": null
    },
    {
      "id": "1122210",
      "postDate": "12/22/2020 09:10:58",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> is there any way to check for 0 gradients? On CPU/GPU usually training stops throwing an error by setting gradients to None and not to 0. Is there any way to force none gradients on TPU and avoid stucking loops?</p>",
      "rawMarkdown": "dimitreoliveira is there any way to check for 0 gradients? On CPU/GPU usually training stops throwing an error by setting gradients to None and not to 0. Is there any way to force none gradients on TPU and avoid stucking loops?",
      "votes": null
    },
    {
      "id": "1122214",
      "postDate": "12/22/2020 09:12:24",
      "content": "<p>Anyway <a href=\"https://www.kaggle.com/denizyigit\" target=\"_blank\">@denizyigit</a> have a look at <a href=\"https://www.kaggle.com/nasirkhalid24/loss-functions-to-help-with-noisy-labelled-data\" target=\"_blank\">this nootebook</a> for some interesting loss functions implemented in tensorflow (it also includes tf version of bi-tempered logistic loss i was talking about)</p>",
      "rawMarkdown": "Anyway @denizyigit have a look at [this nootebook](https://www.kaggle.com/nasirkhalid24/loss-functions-to-help-with-noisy-labelled-data) for some interesting loss functions implemented in tensorflow (it also includes tf version of bi-tempered logistic loss i was talking about)",
      "votes": null
    },
    {
      "id": "1122221",
      "postDate": "12/22/2020 09:16:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>,</p>\n<p>This is a more tensorflow-related question rather than TPU specific.</p>\n<p>Do you have any clue onhow to use a custom gradient computation within your training loop?<br>\nI mean, if i define a custom loss and a custom gradient function and use it with standard custom training loop (tf.gradientTape), I get an error that says: \"gradients cannot flow correctly\".</p>\n<p>Do you have any idea on how to use a custom implementation of gradient computation?</p>",
      "rawMarkdown": "Hi @dimitreoliveira,\n\nThis is a more tensorflow-related question rather than TPU specific.\n\nDo you have any clue onhow to use a custom gradient computation within your training loop?\nI mean, if i define a custom loss and a custom gradient function and use it with standard custom training loop (tf.gradientTape), I get an error that says: \"gradients cannot flow correctly\".\n\nDo you have any idea on how to use a custom implementation of gradient computation?",
      "votes": null
    },
    {
      "id": "1122232",
      "postDate": "12/22/2020 09:20:30",
      "content": "<p>Actually my loss function was computing both softmax probabilities and loss values, thus returning 2 tensors.</p>\n<p><code>probabilities, loss_values = loss_fn(...)</code></p>\n<p>Splitting those two and making the loss output a single tensor solves the \"FakeParam OpKernel\" problem I had.</p>\n<pre><code>probabilities = softmax(...)\nloss_values = loss_fn(...)\n</code></pre>\n<p>Thanks</p>",
      "rawMarkdown": "Actually my loss function was computing both softmax probabilities and loss values, thus returning 2 tensors.\n\n`probabilities, loss_values = loss_fn(...)`\n\nSplitting those two and making the loss output a single tensor solves the \"FakeParam OpKernel\" problem I had.\n\n```\nprobabilities = softmax(...)\nloss_values = loss_fn(...)\n```\n\nThanks",
      "votes": null
    },
    {
      "id": "1122288",
      "postDate": "12/22/2020 10:33:47",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1122448",
      "postDate": "12/22/2020 12:52:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , one hint that the gradients are <code>0</code> is that the training gets stucks, this happens because the network is not being updated, but I am not sure how to make the gradients become <code>None</code>.</p>",
      "rawMarkdown": "Hi @lazcoder , one hint that the gradients are `0` is that the training gets stucks, this happens because the network is not being updated, but I am not sure how to make the gradients become `None`.",
      "votes": null
    },
    {
      "id": "1122456",
      "postDate": "12/22/2020 12:58:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> ,</p>\n<p>If you are using a code like mine <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods#Training\" target=\"_blank\">custom training loop</a>, you can do it inside the <code>train_step</code> function.</p>\n<p>In my code this is how I did:</p>\n<pre><code>with tf.GradientTape() as tape:\n    probabilities = model(x, training=True)\n    loss = loss_fn(y, probabilities, label_smoothing=.3)\ngradients = tape.gradient(loss, model.trainable_variables)\noptimizer.apply_gradients(zip(gradients, model.trainable_variables))\n</code></pre>\n<p>There I am generating the gradient by applying a <code>loss_fn</code> that is <code>categorical_crossentropy</code> loss, but you could also do something like gradient clipping after that.</p>",
      "rawMarkdown": "Hi @lazcoder ,\n\nIf you are using a code like mine [custom training loop](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods#Training), you can do it inside the `train_step` function.\n\nIn my code this is how I did:\n```\nwith tf.GradientTape() as tape:\n    probabilities = model(x, training=True)\n    loss = loss_fn(y, probabilities, label_smoothing=.3)\ngradients = tape.gradient(loss, model.trainable_variables)\noptimizer.apply_gradients(zip(gradients, model.trainable_variables))\n```\n\nThere I am generating the gradient by applying a `loss_fn` that is `categorical_crossentropy` loss, but you could also do something like gradient clipping after that.",
      "votes": null
    },
    {
      "id": "1122975",
      "postDate": "12/22/2020 20:32:36",
      "content": "<p>FYI I created <a href=\"https://www.kaggle.com/lazcoder/cassava-leaf-disease-2019-extraimages-tfrecords\" target=\"_blank\">extraimages tfrecords</a> with the same format used by <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> by using extra-images from 2019 competition</p>",
      "rawMarkdown": "FYI I created [extraimages tfrecords](https://www.kaggle.com/lazcoder/cassava-leaf-disease-2019-extraimages-tfrecords) with the same format used by @dimitreoliveira by using extra-images from 2019 competition",
      "votes": null
    },
    {
      "id": "1128523",
      "postDate": "12/27/2020 14:13:33",
      "content": "<p>Thanks for sharing.. Great work</p>",
      "rawMarkdown": "Thanks for sharing.. Great work",
      "votes": null
    },
    {
      "id": "1129530",
      "postDate": "12/28/2020 12:13:56",
      "content": "<p>Hi, Can you please give out links for external data? thank you!</p>",
      "rawMarkdown": "Hi, Can you please give out links for external data? thank you!",
      "votes": null
    },
    {
      "id": "1130001",
      "postDate": "12/28/2020 17:33:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jiaweil\" target=\"_blank\">@jiaweil</a> , you can find the all the datasets at this link: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198744\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198744</a></p>",
      "rawMarkdown": "Hi @jiaweil , you can find the all the datasets at this link: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198744",
      "votes": null
    },
    {
      "id": "1131266",
      "postDate": "12/29/2020 16:16:36",
      "content": "<p>How many epochs are you training for? I was using B3 for 20 epochs with some of the techniques mentioned above but couldn't get more than 0.89.<br>\nIs it epochs or maybe something else???🤔🤔</p>",
      "rawMarkdown": "How many epochs are you training for? I was using B3 for 20 epochs with some of the techniques mentioned above but couldn't get more than 0.89.\nIs it epochs or maybe something else???🤔🤔",
      "votes": null
    },
    {
      "id": "1131332",
      "postDate": "12/29/2020 16:58:57",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I was usually training for 10~15 epochs, but that can also change depending on your parameters like learning rate and batch size, probably TTA can boost your score a little.</p>",
      "rawMarkdown": "mrinath I was usually training for 10~15 epochs, but that can also change depending on your parameters like learning rate and batch size, probably TTA can boost your score a little.",
      "votes": null
    },
    {
      "id": "1131335",
      "postDate": "12/29/2020 17:02:56",
      "content": "<p>I'm appreciated! </p>",
      "rawMarkdown": "I'm appreciated!",
      "votes": null
    },
    {
      "id": "1132349",
      "postDate": "12/30/2020 10:30:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>,</p>\n<p>I was doing some experiments with optimizers from tensorflow_addons. If I use for instance RAdam, training proceeds without problems. </p>\n<p>Actually I'm in trobule with Lookahead optimizer instead, and when I use it training suddenly stops saying \"socket closed\". </p>\n<p>Did you ever faced a similar problem?</p>",
      "rawMarkdown": "Hi @dimitreoliveira,\n\nI was doing some experiments with optimizers from tensorflow_addons. If I use for instance RAdam, training proceeds without problems. \n\nActually I'm in trobule with Lookahead optimizer instead, and when I use it training suddenly stops saying \"socket closed\". \n\nDid you ever faced a similar problem?",
      "votes": null
    },
    {
      "id": "1132901",
      "postDate": "12/30/2020 19:14:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , actually I was also going to try Lookahead + Adam, but I could not yet because I am on vacation, but I remember that I had used it with success on other competitions, maybe you could try the default parameters, or first try to use lookahead with Adam or sgd, please let me know if you get nice results.</p>",
      "rawMarkdown": "Hi @lazcoder , actually I was also going to try Lookahead + Adam, but I could not yet because I am on vacation, but I remember that I had used it with success on other competitions, maybe you could try the default parameters, or first try to use lookahead with Adam or sgd, please let me know if you get nice results.",
      "votes": null
    },
    {
      "id": "1133562",
      "postDate": "12/31/2020 10:12:39",
      "content": "<p>Do you have any idea why keeping batch normalisation layer frozen worked here?<br>\nBatch Normalisation is usually tightly bound to the 'batch' of input.</p>",
      "rawMarkdown": "Do you have any idea why keeping batch normalisation layer frozen worked here?\nBatch Normalisation is usually tightly bound to the 'batch' of input.",
      "votes": null
    },
    {
      "id": "1133845",
      "postDate": "12/31/2020 15:25:56",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1134399",
      "postDate": "01/01/2021 08:57:39",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1134604",
      "postDate": "01/01/2021 12:20:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> , I had this discussion with another member by the beginning of this post, if you scroll down a little you will find, fell free to share your thoughts.</p>",
      "rawMarkdown": "Hi @prvnkmr , I had this discussion with another member by the beginning of this post, if you scroll down a little you will find, fell free to share your thoughts.",
      "votes": null
    },
    {
      "id": "1135064",
      "postDate": "01/01/2021 21:05:43",
      "content": "<p>Hi,</p>\n<p>I also tried the Lookahead optimizer with RAdam but for some reason, when training begins, it seems to just freeze. After a while, I also get the socket closed error as <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> mentioned. I am training on a TPU. Could this be related to the issue?</p>",
      "rawMarkdown": "Hi,\n\nI also tried the Lookahead optimizer with RAdam but for some reason, when training begins, it seems to just freeze. After a while, I also get the socket closed error as @lazcoder mentioned. I am training on a TPU. Could this be related to the issue?",
      "votes": null
    },
    {
      "id": "1135639",
      "postDate": "01/02/2021 11:49:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> , sometimes this kind of error is related to memory issues, are you using the implementation from the tensorflow add-ons page? Maybe you could also try to see if it works with a smaller network and small images so it could free up the memory a little.</p>",
      "rawMarkdown": "Hi @ayu055 , sometimes this kind of error is related to memory issues, are you using the implementation from the tensorflow add-ons page? Maybe you could also try to see if it works with a smaller network and small images so it could free up the memory a little.",
      "votes": null
    },
    {
      "id": "1137116",
      "postDate": "01/03/2021 17:34:24",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> I read your discussion with Harveen. The point about the mean and variance of the data distribution being the same or different is very insightful. <br>\nThanks a lot for this !! :)</p>",
      "rawMarkdown": "dimitreoliveira I read your discussion with Harveen. The point about the mean and variance of the data distribution being the same or different is very insightful. \nThanks a lot for this !! :)",
      "votes": null
    },
    {
      "id": "1142869",
      "postDate": "01/07/2021 16:30:41",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> Thanks for sharing your notebooks and discussions. I have learned a ton about computer vision from your flower competition notebooks and the cassava leaf competition notebooks. </p>\n<p>When you say maximize MXU what range between percent would you say is good? Was there a specific change that really increased this for you?</p>\n<p>Thanks again!</p>\n<p>Brendan</p>",
      "rawMarkdown": "dimitreoliveira Thanks for sharing your notebooks and discussions. I have learned a ton about computer vision from your flower competition notebooks and the cassava leaf competition notebooks. \n\nWhen you say maximize MXU what range between percent would you say is good? Was there a specific change that really increased this for you?\n\nThanks again!\n\nBrendan",
      "votes": null
    },
    {
      "id": "1143326",
      "postDate": "01/07/2021 21:01:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> , I am glad I could help you.</p>\n<p>About the MXU usage it really depends, maybe a MXU usage of 30% is already too good, depending on the model and data, and batch size, some models can achieve higher usage than others, if I am not mistaken models like transformers can do a lot of parallel computing, while some CNN might not achieve very high usage, and if your batch size is too small the TPU won't be busy all the time, and it will result in lower MXU usages.</p>\n<p>At this notebook the higher MXU impact  came from using an adequate batch size, since EfficientNet already has a good performance among CNN models.</p>",
      "rawMarkdown": "Hi @brendanartley , I am glad I could help you.\n\nAbout the MXU usage it really depends, maybe a MXU usage of 30% is already too good, depending on the model and data, and batch size, some models can achieve higher usage than others, if I am not mistaken models like transformers can do a lot of parallel computing, while some CNN might not achieve very high usage, and if your batch size is too small the TPU won't be busy all the time, and it will result in lower MXU usages.\n\nAt this notebook the higher MXU impact  came from using an adequate batch size, since EfficientNet already has a good performance among CNN models.",
      "votes": null
    },
    {
      "id": "1145483",
      "postDate": "01/09/2021 07:12:09",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> Thanks for Sharing this insights , How did you oversampled the class 0 , 1 ,2,4  ? </p>",
      "rawMarkdown": "dimitreoliveira Thanks for Sharing this insights , How did you oversampled the class 0 , 1 ,2,4  ?",
      "votes": null
    },
    {
      "id": "1145496",
      "postDate": "01/09/2021 07:27:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> I used different datasets that had TFRecords for each class, you can check the code at the linked training notebook.</p>",
      "rawMarkdown": "Hi @sayedathar11 I used different datasets that had TFRecords for each class, you can check the code at the linked training notebook.",
      "votes": null
    },
    {
      "id": "1145590",
      "postDate": "01/09/2021 08:42:06",
      "content": "<p>Okay <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>  I thought oversampling was something related to generating new samples for class using Gans .</p>",
      "rawMarkdown": "Okay @dimitreoliveira  I thought oversampling was something related to generating new samples for class using Gans .",
      "votes": null
    },
    {
      "id": "1148394",
      "postDate": "01/11/2021 05:25:25",
      "content": "<p>I have noticed that noisy student effnets perform better.</p>",
      "rawMarkdown": "I have noticed that noisy student effnets perform better.",
      "votes": null
    },
    {
      "id": "1148474",
      "postDate": "01/11/2021 06:50:37",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I haven't done much experiments with regards to the initialization weights, but in most of the tasks, noisy-student usually better.</p>",
      "rawMarkdown": "mrinath I haven't done much experiments with regards to the initialization weights, but in most of the tasks, noisy-student usually better.",
      "votes": null
    },
    {
      "id": "1148507",
      "postDate": "01/11/2021 07:27:05",
      "content": "<p>Hi, I see you haven't submitted for a few weeks. Got something good brewing? 😁 <br>\nI was wondering if your current LB score (0.9) was from a single model (and it's folds) or if it was an ensemble of several models? My best single model is a B5 at LB 0.899, but I have an ensemble of 2 B5's and a B7 getting LB 0.902. All largely based on your earlier work. That's not to say I haven't tried a bazillion things suggested in the discussion or from other notebooks.<br>\nThanks for the work and discussion points you have written.</p>",
      "rawMarkdown": "Hi, I see you haven't submitted for a few weeks. Got something good brewing? 😁 \nI was wondering if your current LB score (0.9) was from a single model (and it's folds) or if it was an ensemble of several models? My best single model is a B5 at LB 0.899, but I have an ensemble of 2 B5's and a B7 getting LB 0.902. All largely based on your earlier work. That's not to say I haven't tried a bazillion things suggested in the discussion or from other notebooks.\nThanks for the work and discussion points you have written.",
      "votes": null
    },
    {
      "id": "1148875",
      "postDate": "01/11/2021 12:56:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mutantspore\" target=\"_blank\">@mutantspore</a> , Actually I was travelling for 15 days, and just got back, I hope I can catch up with everything 😄.</p>\n<p>My score of <code>0.900</code> was from my public notebooks <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference\" target=\"_blank\">this one</a> and it is simply a 5-fold EfficientNetB4, I got a few ideas, but nothing that seems to be too disruptive yet, previously I did not get good results with <code>B7</code>, so this is already one thing to try.</p>\n<p>I am glad my notebooks could help you, if you have any suggestions please let me know, I am trying to get more proficient with Tensorflow, so any corrections on the code are welcomed.</p>",
      "rawMarkdown": "Hi @mutantspore , Actually I was travelling for 15 days, and just got back, I hope I can catch up with everything 😄.\n\nMy score of `0.900` was from my public notebooks [this one](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference) and it is simply a 5-fold EfficientNetB4, I got a few ideas, but nothing that seems to be too disruptive yet, previously I did not get good results with `B7`, so this is already one thing to try.\n\nI am glad my notebooks could help you, if you have any suggestions please let me know, I am trying to get more proficient with Tensorflow, so any corrections on the code are welcomed.",
      "votes": null
    },
    {
      "id": "1151675",
      "postDate": "01/13/2021 13:41:29",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> , if possible can you please mention the difference (both CV and LB) you got in the following two cases</p>\n<ol>\n<li>With TTA</li>\n<li>Without TTA</li>\n</ol>",
      "rawMarkdown": "dimitreoliveira , if possible can you please mention the difference (both CV and LB) you got in the following two cases\n1. With TTA\n2. Without TTA",
      "votes": null
    },
    {
      "id": "1151769",
      "postDate": "01/13/2021 14:42:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/soumya5891\" target=\"_blank\">@soumya5891</a> </p>\n<p>For my experiments, I have not measured the difference in CV with and without TTA, sometimes the randomness can add noise to model comparison. For the LB score, it can be tricky to say how much TTA improved because it depends on the TTA augmentations and the TTA number, but usually, when I get it right, it can improve my score for about <code>0.003</code></p>",
      "rawMarkdown": "Hi @soumya5891 \n\nFor my experiments, I have not measured the difference in CV with and without TTA, sometimes the randomness can add noise to model comparison. For the LB score, it can be tricky to say how much TTA improved because it depends on the TTA augmentations and the TTA number, but usually, when I get it right, it can improve my score for about `0.003`",
      "votes": null
    },
    {
      "id": "1153942",
      "postDate": "01/15/2021 09:17:38",
      "content": "<p>Thanks. Do You use 2019 training dataset, or pseudo labeling on extra also?</p>",
      "rawMarkdown": "Thanks. Do You use 2019 training dataset, or pseudo labeling on extra also?",
      "votes": null
    },
    {
      "id": "1154027",
      "postDate": "01/15/2021 10:49:04",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/theodotus\" target=\"_blank\">@theodotus</a> ,</p>\n<p>I had similar results by using only 2020 data and using 2020+2019 data, I have not tried pseudo labeling yet, but should try soon.</p>",
      "rawMarkdown": "Hi @theodotus ,\n\nI had similar results by using only 2020 data and using 2020+2019 data, I have not tried pseudo labeling yet, but should try soon.",
      "votes": null
    },
    {
      "id": "1171214",
      "postDate": "01/26/2021 17:34:16",
      "content": "<p>Hi Dimitre, Thank you very much for the valuable suggestions. </p>",
      "rawMarkdown": "Hi Dimitre, Thank you very much for the valuable suggestions.",
      "votes": null
    },
    {
      "id": "1200667",
      "postDate": "02/14/2021 21:01:04",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1115081,
      "author_name": "capiru",
      "author_url": "",
      "post_date": "12/16/2020 01:04:44",
      "content": "<p>Hey, firstly thanks for sharing as this is a very nicely written post. My goal for this competition was to learn how to build a custom TPU training loop, but all my efforts lead to a much slower training time, even if i don't do augmentations inside the training loop. How much faster have you been able to achieve comparing to vanilla keras?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1115113,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/16/2020 01:59:01",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> , you're welcome.</p>\n<p>Custom training loops can be tricky for the first time, but after some hacking you get comfortable with it, I recommend you to fork my kernel and try to play a little.</p>\n<p>About the training time, if you look at the <code>version 22</code> of <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training#Model-evaluation\" target=\"_blank\">this kernel</a><br>\nI am using EfficientNetB4 and the training time is something like this:</p>\n<pre><code>Epoch 1/30\n133/133 - 106s - loss: 1.7036 - sparse_categorical_accuracy: 0.1527 - val_loss: 1.5900 - val_sparse_categorical_accuracy: 0.2801 - lr: 1.0000e-08\nEpoch 2/30\n133/133 - 71s - loss: 0.8568 - sparse_categorical_accuracy: 0.6930 - val_loss: 0.5133 - val_sparse_categorical_accuracy: 0.8229 - lr: 8.0007e-05\n</code></pre>\n<p>On the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">kernel</a> with custom training loop I also use EfficientNetB4 and the training time is:</p>\n<pre><code>EPOCH 1/10\ntime: 296.3s loss: 1.2660 accuracy: 0.6615 val_loss: 0.6311 val_accuracy: 0.8638 lr: 0.00032\nEPOCH 2/10\ntime: 38.9s loss: 1.1040 accuracy: 0.8217 val_loss: 0.5619 val_accuracy: 0.8760 lr: 0.0003104\n</code></pre>\n<p>As you can see the first epoch is slower, but after that, the epochs are much faster.<br>\nNote that the first uses TPUv3 and the second uses the TPUv2 Pod, I will try to run the second kernel with a regular Keras fitting, then I will post here the results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115218,
      "author_name": "harveenchadha",
      "author_url": "",
      "post_date": "12/16/2020 05:09:51",
      "content": "<p>Thank you! Finally someone raised the issue for Batch Normalization layers but the point here I feel is, if you are doing finetuning then only freezing of BN matters. If you are doing training from scratch then you must not freeze BN layers. I have a complete thread <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203095\" target=\"_blank\">here</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F29c16e38ae51f4bb709114afca3e258b%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095363441873&amp;alt=media\" alt=\"pic\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1115631,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/16/2020 12:32:24",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> those are some interesting points, here is what I think.</p>\n<p><strong>What <code>Batch Normalization</code> actually do?</strong>: In simple words, it would be what you pointed, \"Learn the mean and variance of the input data\".</p>\n<p>So, if you are training from scratch you should keep them trainable so they can learn the data distribution.</p>\n<p>But for fine-tunning, I can see two cases:</p>\n<ol>\n<li><strong>Fine-tuning for a domain close to the pre-train domain</strong> (e.g. pre-train on <code>Imagenet</code> but fine-tune to <code>Cat vs Dog</code>), those two domains are very close, so it is intuitive to think that the distribution learned by the pre-train step will be useful for the fine-tune task.</li>\n<li><strong>Fine-tuning for a domain very different to the pre-train domain</strong> (e.g. pre-train on <code>Imagenet</code> but fine-tune to <code>Melanoma classification</code>), there are no images similar to Melanomas on the <code>Imagenet</code> data, so it would be intuitive to think that the data distribution (mean and variance) from <code>Imagenet</code> won't help as much as the <code>Cat vs Dog</code> task.</li>\n</ol>\n<p>Here in this competition, the task domain may be closer to the <code>Imagenet</code>, and keeping the <code>Batch Normalization</code> layers frozen may help.</p>\n<p>As I mentioned, at the Melanoma competition I tried the same but it did not work, and here there was an improvement but it was not too big.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115655,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/16/2020 12:54:43",
          "content": "<p>I think whatever you just said makes a lot of sense. To check if the input data's mean and variance is same as imagenet, we can calculate that as well before making a decision. Right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115952,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/16/2020 17:35:12",
          "content": "<p><a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> this may be a good idea, but in practice, it may not be this simple.</p>\n<p>It may be a little tricky to figure out if the distribution from one dataset is close enough to 'Imagenet'.</p>\n<p>I would suggest just doing an exploration over the dataset, and see if the images are similar to <code>Imagenet</code>, then it is probably worth it to just try both options anyway and see if there is any improvement 😄.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115267,
      "author_name": "slowlearnermack",
      "author_url": "",
      "post_date": "12/16/2020 06:37:24",
      "content": "<p>thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1115613,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/16/2020 12:16:55",
          "content": "<p>You're welcome <a href=\"https://www.kaggle.com/slowlearnermack\" target=\"_blank\">@slowlearnermack</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115877,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "12/16/2020 16:48:53",
      "content": "<p>Nice observations! I wonder how you did oversampling here? Using class weights &gt; 1 or just repeating the images in the dataset? I've never done anything like that, so sorry if this is too naive question. Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116005,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/16/2020 18:24:31",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> I actually did both options, but <code>oversampling</code> refers to <code>increasing the data</code>, in the notebook at the beginning of the training loop you can see that I am <code>adding more data</code> to the minority classes.</p>\n<p>The second option is just <code>increasing class weights</code> for the minority classes, I got no improvements by doing it here, but in theory, the results should be similar.</p>\n<p>Both are very easy to implement, especially if you are using a regular Keras <code>.fit</code> training, but you can feel free to fork the kernel and do some experimentations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116026,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "12/16/2020 18:51:43",
          "content": "<p>Okay. Thanks for the information!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1116881,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/17/2020 14:38:03",
      "content": "<p>Thanks for sharing!!! Are you using the crossentropy loss or some custom loss function?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1117238,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/17/2020 20:27:42",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> , I am using basic <code>Categorical crossentropy</code> with <code>label smoothing</code>, it is defined at the training loop by <code>loss_fn = losses.categorical_crossentropy</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120968,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/21/2020 08:50:02",
          "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> did you used any custom loss on TPU?</p>\n<p>Actually I'm trying to implement my own loss function (that works fine if I run training on GPU) but experiencing some weird problem when training on TPU.</p>\n<p>My loss function uses tf.cond and tf.while_loop, but everything should be TPU compatible, right? Anyway I'm receiving some compilation error from TPU without any clear explanation, something like:</p>\n<p><code>FakeParam OpKernel not implemented on XLA_JIT</code></p>\n<p>Did you ever experienced a similar error?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121283,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/21/2020 14:18:19",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> I used a \"custom loss\" to apply class weights to <code>categorical cross-entropy</code> and it worked, by your error, I guess you might be using some function that can't be compiled by the <code>XLA</code> compiler (TPU uses XLA), also note that some functions are not available to TPUs, for example, some paddings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122232,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/22/2020 09:20:30",
          "content": "<p>Actually my loss function was computing both softmax probabilities and loss values, thus returning 2 tensors.</p>\n<p><code>probabilities, loss_values = loss_fn(...)</code></p>\n<p>Splitting those two and making the loss output a single tensor solves the \"FakeParam OpKernel\" problem I had.</p>\n<pre><code>probabilities = softmax(...)\nloss_values = loss_fn(...)\n</code></pre>\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1116915,
      "author_name": "jbolaifa",
      "author_url": "",
      "post_date": "12/17/2020 15:01:16",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> <br>\nPlease I am new to Kaggle competitions, I have been trying to submit but having difficulties. It says I can't submit with the internet turned on, but when I turn off internet it does not generate output.</p>\n<p>Any help will be appreciated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116963,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "12/17/2020 15:40:16",
          "content": "<p>hello<br>\nyou can make two notebooks; one for training and other for inference, while training you can use the internet and save your weights, later you can use these weights in offline mode for predicting on the test. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116979,
          "author_name": "jbolaifa",
          "author_url": "",
          "post_date": "12/17/2020 16:01:22",
          "content": "<p>When I eventually got a submission file submitted. It says submission file not found. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116981,
          "author_name": "jbolaifa",
          "author_url": "",
          "post_date": "12/17/2020 16:02:40",
          "content": "<p>See this it’s frustrating 😭</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117240,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/17/2020 20:29:58",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jbolaifa\" target=\"_blank\">@jbolaifa</a> , as pointed by <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> , for this competition you need a notebook that does not use the internet to make the submission.</p>\n<p>If you just want to try making the submission you can fork this <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference\" target=\"_blank\">notebook</a> and submit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1119910,
      "author_name": "wwangli",
      "author_url": "",
      "post_date": "12/20/2020 12:52:15",
      "content": "<p>Great job, thanks for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1120046,
      "author_name": "tyqiangz",
      "author_url": "",
      "post_date": "12/20/2020 15:04:47",
      "content": "<p>Thanks for sharing, appreciate it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1120625,
      "author_name": "sjtuyxc",
      "author_url": "",
      "post_date": "12/21/2020 01:44:16",
      "content": "<p>Hello, I would like to ask whether the use of the 2019 data set has improved LB or CV?<br>\nI have used it before, but there seems to be no change, do I need to take more processing？</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120639,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/21/2020 02:11:10",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sjtuyxc\" target=\"_blank\">@sjtuyxc</a> , I have experimented with 2019 data, and the training kernel uses it, but as I mentioned in the <code>Experiments</code> section: </p>\n<pre><code>Small improvements using external data (2019 competition): to be honest, I was expecting better results using data from the 2019 competition, I remember that at the APTOS competition data from the previous competition was key, at the Melanoma competition it also helped more.\n</code></pre>\n<p>I decided to leave the <code>2019</code> for training since it has similar quality and may help with generalization, but in terms of score, the improvements were very minimal, so far I have not seen people reporting great improvements using the external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120960,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/21/2020 08:45:25",
          "content": "<p>Hi, I'm interested in this discussion too. I have experimented same small change or no change at all in both CV and LB score.</p>\n<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> thanks for your notebooks and work. Actually I'm using your tfrecords for 2019 data, but I cannot find the notebook to generate them. Did you check for duplicates?</p>\n<p>I was wondering if there is some way of using 2019 data in this competition also taking into account of unlabeled images and maybe go for a semi-supervised approach.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121219,
          "author_name": "sjtuyxc",
          "author_url": "",
          "post_date": "12/21/2020 13:05:27",
          "content": "<p>Thanks~I got it！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121287,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/21/2020 14:22:38",
          "content": "<p>You're welcome <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , I have not made the notebook to create the external data public because it is essentially the same code to create the regular <code>TFRecords</code>, but if you want to check out the code to evaluate or see how it is done I can make it public.</p>\n<p>About the unlabeled data, I was also thinking to use them as <code>semi-supervised</code>, as you said, but I had no time yet, it is probably worth it to try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121701,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/21/2020 20:39:21",
          "content": "<p>Okay, thanks <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>.</p>\n<p>Did you used all the dataset from 2019 competition (i.e. both training and test set)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121766,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/21/2020 22:24:32",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , I have used only the training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122975,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/22/2020 20:32:36",
          "content": "<p>FYI I created <a href=\"https://www.kaggle.com/lazcoder/cassava-leaf-disease-2019-extraimages-tfrecords\" target=\"_blank\">extraimages tfrecords</a> with the same format used by <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> by using extra-images from 2019 competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121380,
      "author_name": "denizyigit",
      "author_url": "",
      "post_date": "12/21/2020 15:53:12",
      "content": "<p>thanks for sharing! Have you tried using a customized cost function? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1121770,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/21/2020 22:26:54",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/denizyigit\" target=\"_blank\">@denizyigit</a> , For this notebook, the only custom loss was <code>categorical cross-entropy</code> with class weights, but I am doing some experiments with <code>supervised contrastive learning</code>, I will release the notebook in a day or two.</p>\n<p>Have you tried with success any custom loss?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121812,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/21/2020 23:36:11",
          "content": "<p>I'm trying to implement bi-tempered logistic loss as from <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">this discussion</a> it seems to give better performances on noisy datasets. In my experiments anyway I don't get so much improvements. </p>\n<p>I adapted <a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">pytorch implementation</a> from, but the <a href=\"https://github.com/google/bi-tempered-loss\" target=\"_blank\">original version</a> is in tensorflow</p>\n<p>My version seems slightly faster, but when changing temperatures the training get stuck and no error is thrown 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121839,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/22/2020 00:23:52",
          "content": "<p>Nice <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , about the behaviour, if the training freezes maybe your implementation is turning the gradients to 0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122210,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/22/2020 09:10:58",
          "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> is there any way to check for 0 gradients? On CPU/GPU usually training stops throwing an error by setting gradients to None and not to 0. Is there any way to force none gradients on TPU and avoid stucking loops?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122214,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/22/2020 09:12:24",
          "content": "<p>Anyway <a href=\"https://www.kaggle.com/denizyigit\" target=\"_blank\">@denizyigit</a> have a look at <a href=\"https://www.kaggle.com/nasirkhalid24/loss-functions-to-help-with-noisy-labelled-data\" target=\"_blank\">this nootebook</a> for some interesting loss functions implemented in tensorflow (it also includes tf version of bi-tempered logistic loss i was talking about)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122448,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/22/2020 12:52:31",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , one hint that the gradients are <code>0</code> is that the training gets stucks, this happens because the network is not being updated, but I am not sure how to make the gradients become <code>None</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122221,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "12/22/2020 09:16:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>,</p>\n<p>This is a more tensorflow-related question rather than TPU specific.</p>\n<p>Do you have any clue onhow to use a custom gradient computation within your training loop?<br>\nI mean, if i define a custom loss and a custom gradient function and use it with standard custom training loop (tf.gradientTape), I get an error that says: \"gradients cannot flow correctly\".</p>\n<p>Do you have any idea on how to use a custom implementation of gradient computation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122456,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/22/2020 12:58:23",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> ,</p>\n<p>If you are using a code like mine <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods#Training\" target=\"_blank\">custom training loop</a>, you can do it inside the <code>train_step</code> function.</p>\n<p>In my code this is how I did:</p>\n<pre><code>with tf.GradientTape() as tape:\n    probabilities = model(x, training=True)\n    loss = loss_fn(y, probabilities, label_smoothing=.3)\ngradients = tape.gradient(loss, model.trainable_variables)\noptimizer.apply_gradients(zip(gradients, model.trainable_variables))\n</code></pre>\n<p>There I am generating the gradient by applying a <code>loss_fn</code> that is <code>categorical_crossentropy</code> loss, but you could also do something like gradient clipping after that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122288,
      "author_name": "fireheart7",
      "author_url": "",
      "post_date": "12/22/2020 10:33:47",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128523,
      "author_name": "ashikm96",
      "author_url": "",
      "post_date": "12/27/2020 14:13:33",
      "content": "<p>Thanks for sharing.. Great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1129530,
      "author_name": "jiaweil",
      "author_url": "",
      "post_date": "12/28/2020 12:13:56",
      "content": "<p>Hi, Can you please give out links for external data? thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1130001,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/28/2020 17:33:50",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jiaweil\" target=\"_blank\">@jiaweil</a> , you can find the all the datasets at this link: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198744\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198744</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131335,
          "author_name": "jiaweil",
          "author_url": "",
          "post_date": "12/29/2020 17:02:56",
          "content": "<p>I'm appreciated! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1131266,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/29/2020 16:16:36",
      "content": "<p>How many epochs are you training for? I was using B3 for 20 epochs with some of the techniques mentioned above but couldn't get more than 0.89.<br>\nIs it epochs or maybe something else???🤔🤔</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131332,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/29/2020 16:58:57",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I was usually training for 10~15 epochs, but that can also change depending on your parameters like learning rate and batch size, probably TTA can boost your score a little.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1132349,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "12/30/2020 10:30:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>,</p>\n<p>I was doing some experiments with optimizers from tensorflow_addons. If I use for instance RAdam, training proceeds without problems. </p>\n<p>Actually I'm in trobule with Lookahead optimizer instead, and when I use it training suddenly stops saying \"socket closed\". </p>\n<p>Did you ever faced a similar problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1132901,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/30/2020 19:14:28",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> , actually I was also going to try Lookahead + Adam, but I could not yet because I am on vacation, but I remember that I had used it with success on other competitions, maybe you could try the default parameters, or first try to use lookahead with Adam or sgd, please let me know if you get nice results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133562,
      "author_name": "prvnkmr",
      "author_url": "",
      "post_date": "12/31/2020 10:12:39",
      "content": "<p>Do you have any idea why keeping batch normalisation layer frozen worked here?<br>\nBatch Normalisation is usually tightly bound to the 'batch' of input.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1134604,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/01/2021 12:20:35",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> , I had this discussion with another member by the beginning of this post, if you scroll down a little you will find, fell free to share your thoughts.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137116,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/03/2021 17:34:24",
          "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> I read your discussion with Harveen. The point about the mean and variance of the data distribution being the same or different is very insightful. <br>\nThanks a lot for this !! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133845,
      "author_name": "arshpratap",
      "author_url": "",
      "post_date": "12/31/2020 15:25:56",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1134399,
      "author_name": "hiromoon166",
      "author_url": "",
      "post_date": "01/01/2021 08:57:39",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1135064,
      "author_name": "ayu055",
      "author_url": "",
      "post_date": "01/01/2021 21:05:43",
      "content": "<p>Hi,</p>\n<p>I also tried the Lookahead optimizer with RAdam but for some reason, when training begins, it seems to just freeze. After a while, I also get the socket closed error as <a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> mentioned. I am training on a TPU. Could this be related to the issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1135639,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/02/2021 11:49:55",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> , sometimes this kind of error is related to memory issues, are you using the implementation from the tensorflow add-ons page? Maybe you could also try to see if it works with a smaller network and small images so it could free up the memory a little.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1142869,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "01/07/2021 16:30:41",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> Thanks for sharing your notebooks and discussions. I have learned a ton about computer vision from your flower competition notebooks and the cassava leaf competition notebooks. </p>\n<p>When you say maximize MXU what range between percent would you say is good? Was there a specific change that really increased this for you?</p>\n<p>Thanks again!</p>\n<p>Brendan</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143326,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/07/2021 21:01:26",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> , I am glad I could help you.</p>\n<p>About the MXU usage it really depends, maybe a MXU usage of 30% is already too good, depending on the model and data, and batch size, some models can achieve higher usage than others, if I am not mistaken models like transformers can do a lot of parallel computing, while some CNN might not achieve very high usage, and if your batch size is too small the TPU won't be busy all the time, and it will result in lower MXU usages.</p>\n<p>At this notebook the higher MXU impact  came from using an adequate batch size, since EfficientNet already has a good performance among CNN models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1145483,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "01/09/2021 07:12:09",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> Thanks for Sharing this insights , How did you oversampled the class 0 , 1 ,2,4  ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1145496,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/09/2021 07:27:42",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> I used different datasets that had TFRecords for each class, you can check the code at the linked training notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145590,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "01/09/2021 08:42:06",
          "content": "<p>Okay <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>  I thought oversampling was something related to generating new samples for class using Gans .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148394,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "01/11/2021 05:25:25",
      "content": "<p>I have noticed that noisy student effnets perform better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148474,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/11/2021 06:50:37",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I haven't done much experiments with regards to the initialization weights, but in most of the tasks, noisy-student usually better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148507,
      "author_name": "mutantspore",
      "author_url": "",
      "post_date": "01/11/2021 07:27:05",
      "content": "<p>Hi, I see you haven't submitted for a few weeks. Got something good brewing? 😁 <br>\nI was wondering if your current LB score (0.9) was from a single model (and it's folds) or if it was an ensemble of several models? My best single model is a B5 at LB 0.899, but I have an ensemble of 2 B5's and a B7 getting LB 0.902. All largely based on your earlier work. That's not to say I haven't tried a bazillion things suggested in the discussion or from other notebooks.<br>\nThanks for the work and discussion points you have written.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148875,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/11/2021 12:56:24",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mutantspore\" target=\"_blank\">@mutantspore</a> , Actually I was travelling for 15 days, and just got back, I hope I can catch up with everything 😄.</p>\n<p>My score of <code>0.900</code> was from my public notebooks <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference\" target=\"_blank\">this one</a> and it is simply a 5-fold EfficientNetB4, I got a few ideas, but nothing that seems to be too disruptive yet, previously I did not get good results with <code>B7</code>, so this is already one thing to try.</p>\n<p>I am glad my notebooks could help you, if you have any suggestions please let me know, I am trying to get more proficient with Tensorflow, so any corrections on the code are welcomed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1151675,
      "author_name": "soumya5891",
      "author_url": "",
      "post_date": "01/13/2021 13:41:29",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a> , if possible can you please mention the difference (both CV and LB) you got in the following two cases</p>\n<ol>\n<li>With TTA</li>\n<li>Without TTA</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1151769,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/13/2021 14:42:33",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/soumya5891\" target=\"_blank\">@soumya5891</a> </p>\n<p>For my experiments, I have not measured the difference in CV with and without TTA, sometimes the randomness can add noise to model comparison. For the LB score, it can be tricky to say how much TTA improved because it depends on the TTA augmentations and the TTA number, but usually, when I get it right, it can improve my score for about <code>0.003</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1153942,
      "author_name": "theodotus",
      "author_url": "",
      "post_date": "01/15/2021 09:17:38",
      "content": "<p>Thanks. Do You use 2019 training dataset, or pseudo labeling on extra also?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1154027,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "01/15/2021 10:49:04",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/theodotus\" target=\"_blank\">@theodotus</a> ,</p>\n<p>I had similar results by using only 2020 data and using 2020+2019 data, I have not tried pseudo labeling yet, but should try soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1171214,
      "author_name": "gautamv",
      "author_url": "",
      "post_date": "01/26/2021 17:34:16",
      "content": "<p>Hi Dimitre, Thank you very much for the valuable suggestions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1200667,
      "author_name": "index707",
      "author_url": "",
      "post_date": "02/14/2021 21:01:04",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1113996": "Recently I have been doing many experiments to improve my first notebook [Cassava Leaf Disease - TPU Tensorflow - Training\n](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training), and wanted to share my results and maybe have some discussions about what some of you think about.\n\nAs part of the `TPU star` program, I got access to `TPU-v2 Pods` for a few weeks and wanted to make sure I was optimizing the runtime, so I also improved the Tensorflow pipeline.\n\nHere is the link for the [training notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods) and here is the [inference notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference).\n\n5-Fold CV was ranging between `0.890` and `0.898`\n\n___\n\n#### Improvements\n\n- **Custom training loop**: Using a custom training loop greatly improves the training time and resource usage, if you are using Tensorflow I highly recommend to use it, your training time will be a lot less, it also makes it easier to do some nice tricks during training.\n- **Maximize MXU and minimize Idle time**: I have made a few adjustments to the Tensorflow pipeline to improve performance.\n\n___\n\n#### Experiments\n\n**Small improvements using external data (2019 competition)**: to be honest, I was expecting better results using data from the 2019 competition, I remember that at the `APTOS` competition data from the previous competition was key, at the `Melanoma` competition it also helped more.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F0568c1bf1afe589c7ac5a3e058637053%2FScreenshot%20from%202020-12-15%2019-03-28.png?generation=1608069848078572&alt=media)\n___\n**Small improvements from using CCE label smoothing**: This was expected since the competition has some noise, as reported by other users, the thing was that here I got slighter better results by using `label smoothing of 0.3`, for me this value is a little high.\n___\n**Small improvements from using CutOut**: Using `CutOut` improved generalization, I am using a custom `CutOut` function and have not tweaked its parameters a lot.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2Fd15c79d66df79f98360a231be5f6668c%2FScreenshot%20from%202020-12-15%2018-59-53.png?generation=1608069787618497&alt=media)\n___\n\n**Small improvements from oversampling classes 0, 1, 2, and 4**:  Oversampling the dataset with the minority classes improved the training, interestingly this worked a lot better than using class weights.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F280fbb4d2f8dac94a356112f823bdf62%2FScreenshot%20from%202020-12-15%2019-00-04.png?generation=1608069808112772&alt=media)\n___\n**Small improvements from keeping batch normalization layers frozen**:  This was curious for me, I tried to do the same during the `Melanoma` competition,  but there, it did not work, maybe here the plant's images are closer to the `Imagenet` data?\n\nKeeping the `Batch Normalization` layers frozen during fine-tuning is recommended on the [Image classification via fine-tuning with EfficientNet](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/) at the Keras repository.\n___\n**No relevant improvements by using class weights**: I have tried generating class weights in different ways, but got no improvements from it, did not matter using it within a custom training loop or the regular training.\n___\n**No relevant improvements by using MixUp**: I was not expecting it to work, but it could also be a problem with my implementation.\n___\n**No relevant improvements by using different backbones**: I have tried the other `EfficientNet` backbones and some others as well, but for me, I got better results with `EfficientNet` `B3` to `B6`.\n___\n**Worse performance by using different image resolution even the default EfficientNet input size**: At my experiments, the best results came from image resolution `512x512`.\n___\n**Was not able to make progressive unfreezing work**: The concept is cool, and I tried to unfreeze a block of `EfficientNet` at a time but did not get improvements.\n___\n**Changing the learning rate batch-wise seems more efficient than epoch wise, especially for the warm-up phase**: The thing with changing the learning rate epoch-wise is that the model will train with a single learning rate for the complete epoch, changing it batch-wise gives a \"continuous\" change of the learning rate, this is interesting for the warm-up phase.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F9b2c734d02550c9e43c07e4471db12d8%2FScreenshot%20from%202020-12-15%2019-00-13.png?generation=1608069761044709&alt=media)\n___\n\nI will keep doing some more experiments during the week and will update this thread with the results.",
    "1115081": "Hey, firstly thanks for sharing as this is a very nicely written post. My goal for this competition was to learn how to build a custom TPU training loop, but all my efforts lead to a much slower training time, even if i don't do augmentations inside the training loop. How much faster have you been able to achieve comparing to vanilla keras?",
    "1115113": "Hey @capiru , you're welcome.\n\nCustom training loops can be tricky for the first time, but after some hacking you get comfortable with it, I recommend you to fork my kernel and try to play a little.\n\nAbout the training time, if you look at the `version 22` of [this kernel](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training#Model-evaluation)\nI am using EfficientNetB4 and the training time is something like this:\n```\nEpoch 1/30\n133/133 - 106s - loss: 1.7036 - sparse_categorical_accuracy: 0.1527 - val_loss: 1.5900 - val_sparse_categorical_accuracy: 0.2801 - lr: 1.0000e-08\nEpoch 2/30\n133/133 - 71s - loss: 0.8568 - sparse_categorical_accuracy: 0.6930 - val_loss: 0.5133 - val_sparse_categorical_accuracy: 0.8229 - lr: 8.0007e-05\n```\n\nOn the [kernel](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods) with custom training loop I also use EfficientNetB4 and the training time is:\n```\nEPOCH 1/10\ntime: 296.3s loss: 1.2660 accuracy: 0.6615 val_loss: 0.6311 val_accuracy: 0.8638 lr: 0.00032\nEPOCH 2/10\ntime: 38.9s loss: 1.1040 accuracy: 0.8217 val_loss: 0.5619 val_accuracy: 0.8760 lr: 0.0003104\n```\n\nAs you can see the first epoch is slower, but after that, the epochs are much faster.\nNote that the first uses TPUv3 and the second uses the TPUv2 Pod, I will try to run the second kernel with a regular Keras fitting, then I will post here the results.",
    "1115218": "Thank you! Finally someone raised the issue for Batch Normalization layers but the point here I feel is, if you are doing finetuning then only freezing of BN matters. If you are doing training from scratch then you must not freeze BN layers. I have a complete thread [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203095)\n\n![pic](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F29c16e38ae51f4bb709114afca3e258b%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095363441873&alt=media)",
    "1115267": "thanks for sharing!",
    "1115613": "You're welcome @slowlearnermack",
    "1115631": "Hey @harveenchadha those are some interesting points, here is what I think.\n\n**What `Batch Normalization` actually do?**: In simple words, it would be what you pointed, \"Learn the mean and variance of the input data\".\n\nSo, if you are training from scratch you should keep them trainable so they can learn the data distribution.\n\nBut for fine-tunning, I can see two cases:\n1. **Fine-tuning for a domain close to the pre-train domain** (e.g. pre-train on `Imagenet` but fine-tune to `Cat vs Dog`), those two domains are very close, so it is intuitive to think that the distribution learned by the pre-train step will be useful for the fine-tune task.\n2. **Fine-tuning for a domain very different to the pre-train domain** (e.g. pre-train on `Imagenet` but fine-tune to `Melanoma classification`), there are no images similar to Melanomas on the `Imagenet` data, so it would be intuitive to think that the data distribution (mean and variance) from `Imagenet` won't help as much as the `Cat vs Dog` task.\n\nHere in this competition, the task domain may be closer to the `Imagenet`, and keeping the `Batch Normalization` layers frozen may help.\n\nAs I mentioned, at the Melanoma competition I tried the same but it did not work, and here there was an improvement but it was not too big.",
    "1115655": "I think whatever you just said makes a lot of sense. To check if the input data's mean and variance is same as imagenet, we can calculate that as well before making a decision. Right?",
    "1115877": "Nice observations! I wonder how you did oversampling here? Using class weights > 1 or just repeating the images in the dataset? I've never done anything like that, so sorry if this is too naive question. Thanks",
    "1115952": "harveenchadha this may be a good idea, but in practice, it may not be this simple.\n\nIt may be a little tricky to figure out if the distribution from one dataset is close enough to 'Imagenet'.\n\nI would suggest just doing an exploration over the dataset, and see if the images are similar to `Imagenet`, then it is probably worth it to just try both options anyway and see if there is any improvement 😄.",
    "1116005": "Hi @kaushal2896 I actually did both options, but `oversampling` refers to `increasing the data`, in the notebook at the beginning of the training loop you can see that I am `adding more data` to the minority classes.\n\nThe second option is just `increasing class weights` for the minority classes, I got no improvements by doing it here, but in theory, the results should be similar.\n\nBoth are very easy to implement, especially if you are using a regular Keras `.fit` training, but you can feel free to fork the kernel and do some experimentations.",
    "1116026": "Okay. Thanks for the information!",
    "1116881": "Thanks for sharing!!! Are you using the crossentropy loss or some custom loss function?",
    "1116915": "Hello @dimitreoliveira \nPlease I am new to Kaggle competitions, I have been trying to submit but having difficulties. It says I can't submit with the internet turned on, but when I turn off internet it does not generate output.\n\nAny help will be appreciated.",
    "1116963": "hello\nyou can make two notebooks; one for training and other for inference, while training you can use the internet and save your weights, later you can use these weights in offline mode for predicting on the test.",
    "1116979": "When I eventually got a submission file submitted. It says submission file not found.",
    "1116981": "See this it’s frustrating 😭",
    "1117238": "Hi @mrinath , I am using basic `Categorical crossentropy` with `label smoothing`, it is defined at the training loop by `loss_fn = losses.categorical_crossentropy`.",
    "1117240": "Hi @jbolaifa , as pointed by @mrinath , for this competition you need a notebook that does not use the internet to make the submission.\n\nIf you just want to try making the submission you can fork this [notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference) and submit.",
    "1119910": "Great job, thanks for sharing!",
    "1120046": "Thanks for sharing, appreciate it!",
    "1120625": "Hello, I would like to ask whether the use of the 2019 data set has improved LB or CV?\nI have used it before, but there seems to be no change, do I need to take more processing？",
    "1120639": "Hi @sjtuyxc , I have experimented with 2019 data, and the training kernel uses it, but as I mentioned in the `Experiments` section: \n\n```\nSmall improvements using external data (2019 competition): to be honest, I was expecting better results using data from the 2019 competition, I remember that at the APTOS competition data from the previous competition was key, at the Melanoma competition it also helped more.\n```\n\nI decided to leave the `2019` for training since it has similar quality and may help with generalization, but in terms of score, the improvements were very minimal, so far I have not seen people reporting great improvements using the external data.",
    "1120960": "Hi, I'm interested in this discussion too. I have experimented same small change or no change at all in both CV and LB score.\n\n@dimitreoliveira thanks for your notebooks and work. Actually I'm using your tfrecords for 2019 data, but I cannot find the notebook to generate them. Did you check for duplicates?\n\nI was wondering if there is some way of using 2019 data in this competition also taking into account of unlabeled images and maybe go for a semi-supervised approach.",
    "1120968": "dimitreoliveira did you used any custom loss on TPU?\n\nActually I'm trying to implement my own loss function (that works fine if I run training on GPU) but experiencing some weird problem when training on TPU.\n\nMy loss function uses tf.cond and tf.while_loop, but everything should be TPU compatible, right? Anyway I'm receiving some compilation error from TPU without any clear explanation, something like:\n\n`FakeParam OpKernel not implemented on XLA_JIT`\n\nDid you ever experienced a similar error?",
    "1121219": "Thanks~I got it！",
    "1121283": "Hey @lazcoder I used a \"custom loss\" to apply class weights to `categorical cross-entropy` and it worked, by your error, I guess you might be using some function that can't be compiled by the `XLA` compiler (TPU uses XLA), also note that some functions are not available to TPUs, for example, some paddings.",
    "1121287": "You're welcome @lazcoder , I have not made the notebook to create the external data public because it is essentially the same code to create the regular `TFRecords`, but if you want to check out the code to evaluate or see how it is done I can make it public.\n\nAbout the unlabeled data, I was also thinking to use them as `semi-supervised`, as you said, but I had no time yet, it is probably worth it to try.",
    "1121380": "thanks for sharing! Have you tried using a customized cost function?",
    "1121701": "Okay, thanks @dimitreoliveira.\n\nDid you used all the dataset from 2019 competition (i.e. both training and test set)?",
    "1121766": "Hey @lazcoder , I have used only the training data.",
    "1121770": "Hi @denizyigit , For this notebook, the only custom loss was `categorical cross-entropy` with class weights, but I am doing some experiments with `supervised contrastive learning`, I will release the notebook in a day or two.\n\nHave you tried with success any custom loss?",
    "1121812": "I'm trying to implement bi-tempered logistic loss as from [this discussion](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017) it seems to give better performances on noisy datasets. In my experiments anyway I don't get so much improvements. \n\nI adapted [pytorch implementation](https://github.com/mlpanda/bi-tempered-loss-pytorch) from, but the [original version](https://github.com/google/bi-tempered-loss) is in tensorflow\n\nMy version seems slightly faster, but when changing temperatures the training get stuck and no error is thrown 👍",
    "1121839": "Nice @lazcoder , about the behaviour, if the training freezes maybe your implementation is turning the gradients to 0.",
    "1122210": "dimitreoliveira is there any way to check for 0 gradients? On CPU/GPU usually training stops throwing an error by setting gradients to None and not to 0. Is there any way to force none gradients on TPU and avoid stucking loops?",
    "1122214": "Anyway @denizyigit have a look at [this nootebook](https://www.kaggle.com/nasirkhalid24/loss-functions-to-help-with-noisy-labelled-data) for some interesting loss functions implemented in tensorflow (it also includes tf version of bi-tempered logistic loss i was talking about)",
    "1122221": "Hi @dimitreoliveira,\n\nThis is a more tensorflow-related question rather than TPU specific.\n\nDo you have any clue onhow to use a custom gradient computation within your training loop?\nI mean, if i define a custom loss and a custom gradient function and use it with standard custom training loop (tf.gradientTape), I get an error that says: \"gradients cannot flow correctly\".\n\nDo you have any idea on how to use a custom implementation of gradient computation?",
    "1122232": "Actually my loss function was computing both softmax probabilities and loss values, thus returning 2 tensors.\n\n`probabilities, loss_values = loss_fn(...)`\n\nSplitting those two and making the loss output a single tensor solves the \"FakeParam OpKernel\" problem I had.\n\n```\nprobabilities = softmax(...)\nloss_values = loss_fn(...)\n```\n\nThanks",
    "1122288": "Thank you for sharing!",
    "1122448": "Hi @lazcoder , one hint that the gradients are `0` is that the training gets stucks, this happens because the network is not being updated, but I am not sure how to make the gradients become `None`.",
    "1122456": "Hi @lazcoder ,\n\nIf you are using a code like mine [custom training loop](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods#Training), you can do it inside the `train_step` function.\n\nIn my code this is how I did:\n```\nwith tf.GradientTape() as tape:\n    probabilities = model(x, training=True)\n    loss = loss_fn(y, probabilities, label_smoothing=.3)\ngradients = tape.gradient(loss, model.trainable_variables)\noptimizer.apply_gradients(zip(gradients, model.trainable_variables))\n```\n\nThere I am generating the gradient by applying a `loss_fn` that is `categorical_crossentropy` loss, but you could also do something like gradient clipping after that.",
    "1122975": "FYI I created [extraimages tfrecords](https://www.kaggle.com/lazcoder/cassava-leaf-disease-2019-extraimages-tfrecords) with the same format used by @dimitreoliveira by using extra-images from 2019 competition",
    "1128523": "Thanks for sharing.. Great work",
    "1129530": "Hi, Can you please give out links for external data? thank you!",
    "1130001": "Hi @jiaweil , you can find the all the datasets at this link: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198744",
    "1131266": "How many epochs are you training for? I was using B3 for 20 epochs with some of the techniques mentioned above but couldn't get more than 0.89.\nIs it epochs or maybe something else???🤔🤔",
    "1131332": "mrinath I was usually training for 10~15 epochs, but that can also change depending on your parameters like learning rate and batch size, probably TTA can boost your score a little.",
    "1131335": "I'm appreciated!",
    "1132349": "Hi @dimitreoliveira,\n\nI was doing some experiments with optimizers from tensorflow_addons. If I use for instance RAdam, training proceeds without problems. \n\nActually I'm in trobule with Lookahead optimizer instead, and when I use it training suddenly stops saying \"socket closed\". \n\nDid you ever faced a similar problem?",
    "1132901": "Hi @lazcoder , actually I was also going to try Lookahead + Adam, but I could not yet because I am on vacation, but I remember that I had used it with success on other competitions, maybe you could try the default parameters, or first try to use lookahead with Adam or sgd, please let me know if you get nice results.",
    "1133562": "Do you have any idea why keeping batch normalisation layer frozen worked here?\nBatch Normalisation is usually tightly bound to the 'batch' of input.",
    "1133845": "Thanks for sharing",
    "1134399": "Thank you for sharing!",
    "1134604": "Hi @prvnkmr , I had this discussion with another member by the beginning of this post, if you scroll down a little you will find, fell free to share your thoughts.",
    "1135064": "Hi,\n\nI also tried the Lookahead optimizer with RAdam but for some reason, when training begins, it seems to just freeze. After a while, I also get the socket closed error as @lazcoder mentioned. I am training on a TPU. Could this be related to the issue?",
    "1135639": "Hi @ayu055 , sometimes this kind of error is related to memory issues, are you using the implementation from the tensorflow add-ons page? Maybe you could also try to see if it works with a smaller network and small images so it could free up the memory a little.",
    "1137116": "dimitreoliveira I read your discussion with Harveen. The point about the mean and variance of the data distribution being the same or different is very insightful. \nThanks a lot for this !! :)",
    "1142869": "dimitreoliveira Thanks for sharing your notebooks and discussions. I have learned a ton about computer vision from your flower competition notebooks and the cassava leaf competition notebooks. \n\nWhen you say maximize MXU what range between percent would you say is good? Was there a specific change that really increased this for you?\n\nThanks again!\n\nBrendan",
    "1143326": "Hi @brendanartley , I am glad I could help you.\n\nAbout the MXU usage it really depends, maybe a MXU usage of 30% is already too good, depending on the model and data, and batch size, some models can achieve higher usage than others, if I am not mistaken models like transformers can do a lot of parallel computing, while some CNN might not achieve very high usage, and if your batch size is too small the TPU won't be busy all the time, and it will result in lower MXU usages.\n\nAt this notebook the higher MXU impact  came from using an adequate batch size, since EfficientNet already has a good performance among CNN models.",
    "1145483": "dimitreoliveira Thanks for Sharing this insights , How did you oversampled the class 0 , 1 ,2,4  ?",
    "1145496": "Hi @sayedathar11 I used different datasets that had TFRecords for each class, you can check the code at the linked training notebook.",
    "1145590": "Okay @dimitreoliveira  I thought oversampling was something related to generating new samples for class using Gans .",
    "1148394": "I have noticed that noisy student effnets perform better.",
    "1148474": "mrinath I haven't done much experiments with regards to the initialization weights, but in most of the tasks, noisy-student usually better.",
    "1148507": "Hi, I see you haven't submitted for a few weeks. Got something good brewing? 😁 \nI was wondering if your current LB score (0.9) was from a single model (and it's folds) or if it was an ensemble of several models? My best single model is a B5 at LB 0.899, but I have an ensemble of 2 B5's and a B7 getting LB 0.902. All largely based on your earlier work. That's not to say I haven't tried a bazillion things suggested in the discussion or from other notebooks.\nThanks for the work and discussion points you have written.",
    "1148875": "Hi @mutantspore , Actually I was travelling for 15 days, and just got back, I hope I can catch up with everything 😄.\n\nMy score of `0.900` was from my public notebooks [this one](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference) and it is simply a 5-fold EfficientNetB4, I got a few ideas, but nothing that seems to be too disruptive yet, previously I did not get good results with `B7`, so this is already one thing to try.\n\nI am glad my notebooks could help you, if you have any suggestions please let me know, I am trying to get more proficient with Tensorflow, so any corrections on the code are welcomed.",
    "1151675": "dimitreoliveira , if possible can you please mention the difference (both CV and LB) you got in the following two cases\n1. With TTA\n2. Without TTA",
    "1151769": "Hi @soumya5891 \n\nFor my experiments, I have not measured the difference in CV with and without TTA, sometimes the randomness can add noise to model comparison. For the LB score, it can be tricky to say how much TTA improved because it depends on the TTA augmentations and the TTA number, but usually, when I get it right, it can improve my score for about `0.003`",
    "1153942": "Thanks. Do You use 2019 training dataset, or pseudo labeling on extra also?",
    "1154027": "Hi @theodotus ,\n\nI had similar results by using only 2020 data and using 2020+2019 data, I have not tried pseudo labeling yet, but should try soon.",
    "1171214": "Hi Dimitre, Thank you very much for the valuable suggestions.",
    "1200667": "Thanks for sharing"
  },
  "source": "meta"
}