{
  "id": 214559,
  "title": "Does TTA really work?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214559",
  "author_name": "",
  "post_date": "2021-01-27T05:14:01.466057Z",
  "votes": 31,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I have taken several days to train my models. 01/23 is my first day to probe the LB score.</p>\n<p>After some experiments, Here are my current results.</p>\n<p>Training 2 models (seresnext50, effb4ns) with,</p>\n<p>image size = 512,<br>\nmerged dataset (drop the duplicates based on <a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">this notebook</a> by <a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a>),<br>\nbi-tempered loss,<br>\nA little bit heavy random aug,<br>\nno valid time aug,<br>\nstratified 5-fold</p>\n<p>The results are,</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Seresnext50</th>\n<th>Effb4ns</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV without TTA</td>\n<td>0.895+</td>\n<td>0.891+</td>\n</tr>\n<tr>\n<td>LB without TTA</td>\n<td>0.903</td>\n<td>0.900</td>\n</tr>\n<tr>\n<td>LB with 5x light TTA</td>\n<td>0.900</td>\n<td>0.902</td>\n</tr>\n</tbody>\n</table>\n<p>light TTA means only doing flips, rotates, and crops, while in training process doing coarseout, cutout, HueSaturationValue, blur and so on.</p>\n<p>The baseline seresnext50 and effb4ns scored LB 0.899 and 0.898 respectively with CV scores 0.891+ and 0.890+.</p>\n<p>I have not done many experiments (will do). But, from the current results, it seems TTA randomly affects the LB socre. Without TTA, we can see that the CV and LB are probably aligned. Or, probably I am using the wrong way to implement TTA.</p>\n<p>So, does TTA really work for this competition? Or, should we trust CV with valid time aug? Is it useful to analyze the OOF with and without valid time aug?</p>",
  "messages": [
    {
      "id": "1171773",
      "postDate": "01/27/2021 05:14:01",
      "content": "<p>I have taken several days to train my models. 01/23 is my first day to probe the LB score.</p>\n<p>After some experiments, Here are my current results.</p>\n<p>Training 2 models (seresnext50, effb4ns) with,</p>\n<p>image size = 512,<br>\nmerged dataset (drop the duplicates based on <a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">this notebook</a> by <a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a>),<br>\nbi-tempered loss,<br>\nA little bit heavy random aug,<br>\nno valid time aug,<br>\nstratified 5-fold</p>\n<p>The results are,</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Seresnext50</th>\n<th>Effb4ns</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV without TTA</td>\n<td>0.895+</td>\n<td>0.891+</td>\n</tr>\n<tr>\n<td>LB without TTA</td>\n<td>0.903</td>\n<td>0.900</td>\n</tr>\n<tr>\n<td>LB with 5x light TTA</td>\n<td>0.900</td>\n<td>0.902</td>\n</tr>\n</tbody>\n</table>\n<p>light TTA means only doing flips, rotates, and crops, while in training process doing coarseout, cutout, HueSaturationValue, blur and so on.</p>\n<p>The baseline seresnext50 and effb4ns scored LB 0.899 and 0.898 respectively with CV scores 0.891+ and 0.890+.</p>\n<p>I have not done many experiments (will do). But, from the current results, it seems TTA randomly affects the LB socre. Without TTA, we can see that the CV and LB are probably aligned. Or, probably I am using the wrong way to implement TTA.</p>\n<p>So, does TTA really work for this competition? Or, should we trust CV with valid time aug? Is it useful to analyze the OOF with and without valid time aug?</p>",
      "rawMarkdown": "I have taken several days to train my models. 01/23 is my first day to probe the LB score.\n\nAfter some experiments, Here are my current results.\n\nTraining 2 models (seresnext50, effb4ns) with,\n\nimage size = 512,\nmerged dataset (drop the duplicates based on [this notebook](https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan) by @graf10a),\nbi-tempered loss,\nA little bit heavy random aug,\nno valid time aug,\nstratified 5-fold\n\nThe results are,\n\n|| Seresnext50 |Effb4ns  |\n| --- | --- |\n|CV without TTA| 0.895+ |0.891+ |\n|LB without TTA| 0.903 | 0.900 |\n|LB with 5x light TTA|0.900  |0.902  |\n\nlight TTA means only doing flips, rotates, and crops, while in training process doing coarseout, cutout, HueSaturationValue, blur and so on.\n\nThe baseline seresnext50 and effb4ns scored LB 0.899 and 0.898 respectively with CV scores 0.891+ and 0.890+.\n\nI have not done many experiments (will do). But, from the current results, it seems TTA randomly affects the LB socre. Without TTA, we can see that the CV and LB are probably aligned. Or, probably I am using the wrong way to implement TTA.\n\nSo, does TTA really work for this competition? Or, should we trust CV with valid time aug? Is it useful to analyze the OOF with and without valid time aug?",
      "votes": null
    },
    {
      "id": "1171803",
      "postDate": "01/27/2021 05:36:17",
      "content": "<p>IMO </p>\n<p>TTA done correctly always results is a better result.  We don't know how you did your TTA or how many loops you made.  My limited testing indicates that 10 is the minimum number to establish a stable result.</p>\n<p>Statistics done without understanding variation is always going to lead to poor decisions.</p>\n<p>The test set is 15,000 images.  If my math is correct the LB score is reported for 4650 images.</p>\n<p>0.903 LB score is 4198 images that are correct<br>\n0.900 LB score is  4185 images that are correct</p>\n<p>Do you believe 13 images is a statistically valid difference for this noisy data set?</p>\n<p>Can you use the eye ball test to decide  a 0.003 difference on 4650 images?</p>",
      "rawMarkdown": "IMO \n\nTTA done correctly always results is a better result.  We don't know how you did your TTA or how many loops you made.  My limited testing indicates that 10 is the minimum number to establish a stable result.\n\nStatistics done without understanding variation is always going to lead to poor decisions.\n\nThe test set is 15,000 images.  If my math is correct the LB score is reported for 4650 images.\n\n0.903 LB score is 4198 images that are correct\n0.900 LB score is  4185 images that are correct\n\nDo you believe 13 images is a statistically valid difference for this noisy data set?\n\nCan you use the eye ball test to decide  a 0.003 difference on 4650 images?",
      "votes": null
    },
    {
      "id": "1171894",
      "postDate": "01/27/2021 06:55:45",
      "content": "<p>In most cases, TTA is effective</p>",
      "rawMarkdown": "In most cases, TTA is effective",
      "votes": null
    },
    {
      "id": "1172055",
      "postDate": "01/27/2021 08:40:06",
      "content": "<p>In my experience, TTA is effective for LB score. For TTA, I used only light aug, too. But, as <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> said below, I don't know how TTA is really effective for our private LB score. </p>",
      "rawMarkdown": "In my experience, TTA is effective for LB score. For TTA, I used only light aug, too. But, as @pcjimmmy said below, I don't know how TTA is really effective for our private LB score.",
      "votes": null
    },
    {
      "id": "1172850",
      "postDate": "01/27/2021 15:14:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> thanks for your tips. I will try 10+ TTA on both CV and LB.</p>\n<p>Yes, you might be right. You remind me that there are noisy labels in hidden test set. We already find an example in graf10a's notebook, so I truly believe it. </p>\n<p>The decreased 0.003 may not be the problem of TTA but the problem of noisy labels. Because, after TTA, the model can predict an image correctly but it was labeled improperly. So, it would be scored 0.</p>\n<p>But, the question is how noisy the test set is. If there are lots of noisy labels but we have a robust model that generalizes well, the score may be a bit lower than a wrong model.</p>",
      "rawMarkdown": "Hi @pcjimmmy thanks for your tips. I will try 10+ TTA on both CV and LB.\n\nYes, you might be right. You remind me that there are noisy labels in hidden test set. We already find an example in graf10a's notebook, so I truly believe it. \n\nThe decreased 0.003 may not be the problem of TTA but the problem of noisy labels. Because, after TTA, the model can predict an image correctly but it was labeled improperly. So, it would be scored 0.\n\nBut, the question is how noisy the test set is. If there are lots of noisy labels but we have a robust model that generalizes well, the score may be a bit lower than a wrong model.",
      "votes": null
    },
    {
      "id": "1172961",
      "postDate": "01/27/2021 16:18:38",
      "content": "<p>I am doing all my submissions now with 3 loops - once I find the models that I love than I will increment the number until submission timeout.  </p>",
      "rawMarkdown": "I am doing all my submissions now with 3 loops - once I find the models that I love than I will increment the number until submission timeout.",
      "votes": null
    },
    {
      "id": "1173012",
      "postDate": "01/27/2021 16:47:44",
      "content": "<p><a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> Which augmentations do you use for TTA? I had random rotations and cropping but it made my leaderboard scores inconsistent. </p>",
      "rawMarkdown": "woshifym Which augmentations do you use for TTA? I had random rotations and cropping but it made my leaderboard scores inconsistent.",
      "votes": null
    },
    {
      "id": "1173230",
      "postDate": "01/27/2021 18:34:45",
      "content": "<p>In previous computer vision competitions, TTA was mostly applied in 1st place solutions</p>\n<p>My personal experience underlines that, TTA did increase public LB aswell as private LB.</p>",
      "rawMarkdown": "In previous computer vision competitions, TTA was mostly applied in 1st place solutions\n\nMy personal experience underlines that, TTA did increase public LB aswell as private LB.",
      "votes": null
    },
    {
      "id": "1173367",
      "postDate": "01/27/2021 20:28:48",
      "content": "<p>As I mentioned, I just used simple random flips, rotates, and crops.</p>",
      "rawMarkdown": "As I mentioned, I just used simple random flips, rotates, and crops.",
      "votes": null
    },
    {
      "id": "1173369",
      "postDate": "01/27/2021 20:29:16",
      "content": "<p>Thanks for tips, I am just confused with this competition.</p>",
      "rawMarkdown": "Thanks for tips, I am just confused with this competition.",
      "votes": null
    },
    {
      "id": "1173383",
      "postDate": "01/27/2021 20:46:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> what did you mean loops? Is it about how many times TTA used?</p>",
      "rawMarkdown": "Hi @pcjimmmy what did you mean loops? Is it about how many times TTA used?",
      "votes": null
    },
    {
      "id": "1173390",
      "postDate": "01/27/2021 20:56:08",
      "content": "<p><a href=\"https://www.kaggle.com/yimin22\" target=\"_blank\">@yimin22</a> For crops, are you doing random crops? If so, does your LB score fluctuate a lot?</p>",
      "rawMarkdown": "yimin22 For crops, are you doing random crops? If so, does your LB score fluctuate a lot?",
      "votes": null
    },
    {
      "id": "1173476",
      "postDate": "01/27/2021 23:26:19",
      "content": "<p>Hi yimin,   great achievement you have made.   </p>\n<p>I saw people using seresnext50. but I checked keras applications,  it does not have seresnext50, only had resnet50.  I wonder how would you create model to use seresnext50 in keras?   Really beginner question, don't laugh </p>",
      "rawMarkdown": "Hi yimin,   great achievement you have made.   \n\nI saw people using seresnext50. but I checked keras applications,  it does not have seresnext50, only had resnet50.  I wonder how would you create model to use seresnext50 in keras?   Really beginner question, don't laugh",
      "votes": null
    },
    {
      "id": "1173549",
      "postDate": "01/28/2021 00:27:38",
      "content": "<p>Loops = times used.  </p>",
      "rawMarkdown": "Loops = times used.",
      "votes": null
    },
    {
      "id": "1173635",
      "postDate": "01/28/2021 02:59:57",
      "content": "<p><a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> yes. LB decreases but not fluctuate.</p>",
      "rawMarkdown": "ayu055 yes. LB decreases but not fluctuate.",
      "votes": null
    },
    {
      "id": "1173639",
      "postDate": "01/28/2021 03:00:44",
      "content": "<p>Thanks. I used pytorch to implement the seresnext50.</p>",
      "rawMarkdown": "Thanks. I used pytorch to implement the seresnext50.",
      "votes": null
    },
    {
      "id": "1174082",
      "postDate": "01/28/2021 08:54:27",
      "content": "<p>Do you apply a subset of the augmentation used in training or do you design TTA independently of training? </p>",
      "rawMarkdown": "Do you apply a subset of the augmentation used in training or do you design TTA independently of training?",
      "votes": null
    },
    {
      "id": "1174413",
      "postDate": "01/28/2021 13:07:03",
      "content": "<p>In my attempt, I chose TTA from 2 to 8, and the LB score was slightly improved.</p>",
      "rawMarkdown": "In my attempt, I chose TTA from 2 to 8, and the LB score was slightly improved.",
      "votes": null
    },
    {
      "id": "1174423",
      "postDate": "01/28/2021 13:18:15",
      "content": "<p>Thats right, normally higher TTA increases LB but therefore takes longer to submit due to higher amount of inferences.</p>",
      "rawMarkdown": "Thats right, normally higher TTA increases LB but therefore takes longer to submit due to higher amount of inferences.",
      "votes": null
    },
    {
      "id": "1175182",
      "postDate": "01/29/2021 01:40:54",
      "content": "<p>Not work from my side. Only tried Resnet34 for experiment with different tta and you can check <a href=\"https://www.kaggle.com/leonshangguan/inference-test\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Not work from my side. Only tried Resnet34 for experiment with different tta and you can check [here](https://www.kaggle.com/leonshangguan/inference-test).",
      "votes": null
    },
    {
      "id": "1175578",
      "postDate": "01/29/2021 08:28:34",
      "content": "<p>TTA boosts my score<br>\nMy EfficientNEt5, 512x512 size without TTA -&gt; 0.898<br>\n3x TTA -&gt; 0.901</p>",
      "rawMarkdown": "TTA boosts my score\nMy EfficientNEt5, 512x512 size without TTA -> 0.898\n3x TTA -> 0.901",
      "votes": null
    },
    {
      "id": "1184454",
      "postDate": "02/03/2021 14:33:05",
      "content": "<p><a href=\"https://www.kaggle.com/luqing2\" target=\"_blank\">@luqing2</a> You can use seresnext50 in keras by installing these GitHub repositories: <a href=\"https://github.com/qubvel/classification_models\" target=\"_blank\">Classification models </a> and <a href=\"https://github.com/keras-team/keras-applications\" target=\"_blank\">Keras applications </a>.<br>\nFor training you can keep the internet on and install the libraries directly:</p>\n<pre><code>!pip install image-classifiers\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(512,512,3))\n\n    base = seresnext50(input_shape=(512,512,3),weights = 'imagenet', include_top = False)\n\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(5,activation='softmax')(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n</code></pre>\n<p>In inference, with the internet off, you need to create kaggle datasets of those 2 repositories and attach them to your notebook:</p>\n<pre><code>import sys\nsys.path.append('../input/keras-applications/')\nsys.path.append('../input/keras-seresnext-50-101/')\nimport keras_applications as ka\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n</code></pre>",
      "rawMarkdown": "luqing2 You can use seresnext50 in keras by installing these GitHub repositories: [Classification models ](https://github.com/qubvel/classification_models) and [Keras applications ](https://github.com/keras-team/keras-applications).\nFor training you can keep the internet on and install the libraries directly:\n```\n!pip install image-classifiers\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(512,512,3))\n\n    base = seresnext50(input_shape=(512,512,3),weights = 'imagenet', include_top = False)\n    \n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(5,activation='softmax')(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n```\n\nIn inference, with the internet off, you need to create kaggle datasets of those 2 repositories and attach them to your notebook:\n\n```\nimport sys\nsys.path.append('../input/keras-applications/')\nsys.path.append('../input/keras-seresnext-50-101/')\nimport keras_applications as ka\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n```",
      "votes": null
    },
    {
      "id": "1184461",
      "postDate": "02/03/2021 14:36:00",
      "content": "<p><a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> In addition to coarseout, cutout, HueSaturationValue… Did you use cutmix, mixup… to get LB=0.903 with seresnext50?</p>",
      "rawMarkdown": "woshifym In addition to coarseout, cutout, HueSaturationValue... Did you use cutmix, mixup... to get LB=0.903 with seresnext50?",
      "votes": null
    },
    {
      "id": "1185720",
      "postDate": "02/04/2021 10:19:14",
      "content": "<p>Do you use the timm's seresnext50 or pretrainedmodel's seresnext50? Thanks a lot! My effb4 is same as yours, but seresnext50 is much lower.</p>",
      "rawMarkdown": "Do you use the timm's seresnext50 or pretrainedmodel's seresnext50? Thanks a lot! My effb4 is same as yours, but seresnext50 is much lower.",
      "votes": null
    },
    {
      "id": "1186459",
      "postDate": "02/04/2021 20:38:47",
      "content": "<p>I used pretrained seresnext50</p>",
      "rawMarkdown": "I used pretrained seresnext50",
      "votes": null
    },
    {
      "id": "1186461",
      "postDate": "02/04/2021 20:41:22",
      "content": "<p><a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> no, cutmix decreases both CV and LB score of my model.</p>",
      "rawMarkdown": "amiiiney no, cutmix decreases both CV and LB score of my model.",
      "votes": null
    },
    {
      "id": "1186468",
      "postDate": "02/04/2021 20:42:39",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Did you use the TTA same as the aug used in training?</p>",
      "rawMarkdown": "deepkim Did you use the TTA same as the aug used in training?",
      "votes": null
    },
    {
      "id": "1186474",
      "postDate": "02/04/2021 20:43:45",
      "content": "<p><a href=\"https://www.kaggle.com/tiboas\" target=\"_blank\">@tiboas</a> it is subset of the aug used in training</p>",
      "rawMarkdown": "tiboas it is subset of the aug used in training",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1171803,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/27/2021 05:36:17",
      "content": "<p>IMO </p>\n<p>TTA done correctly always results is a better result.  We don't know how you did your TTA or how many loops you made.  My limited testing indicates that 10 is the minimum number to establish a stable result.</p>\n<p>Statistics done without understanding variation is always going to lead to poor decisions.</p>\n<p>The test set is 15,000 images.  If my math is correct the LB score is reported for 4650 images.</p>\n<p>0.903 LB score is 4198 images that are correct<br>\n0.900 LB score is  4185 images that are correct</p>\n<p>Do you believe 13 images is a statistically valid difference for this noisy data set?</p>\n<p>Can you use the eye ball test to decide  a 0.003 difference on 4650 images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1172850,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "01/27/2021 15:14:18",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> thanks for your tips. I will try 10+ TTA on both CV and LB.</p>\n<p>Yes, you might be right. You remind me that there are noisy labels in hidden test set. We already find an example in graf10a's notebook, so I truly believe it. </p>\n<p>The decreased 0.003 may not be the problem of TTA but the problem of noisy labels. Because, after TTA, the model can predict an image correctly but it was labeled improperly. So, it would be scored 0.</p>\n<p>But, the question is how noisy the test set is. If there are lots of noisy labels but we have a robust model that generalizes well, the score may be a bit lower than a wrong model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1172961,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/27/2021 16:18:38",
          "content": "<p>I am doing all my submissions now with 3 loops - once I find the models that I love than I will increment the number until submission timeout.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173012,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/27/2021 16:47:44",
          "content": "<p><a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> Which augmentations do you use for TTA? I had random rotations and cropping but it made my leaderboard scores inconsistent. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173367,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "01/27/2021 20:28:48",
          "content": "<p>As I mentioned, I just used simple random flips, rotates, and crops.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173383,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "01/27/2021 20:46:20",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> what did you mean loops? Is it about how many times TTA used?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173390,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/27/2021 20:56:08",
          "content": "<p><a href=\"https://www.kaggle.com/yimin22\" target=\"_blank\">@yimin22</a> For crops, are you doing random crops? If so, does your LB score fluctuate a lot?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173549,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/28/2021 00:27:38",
          "content": "<p>Loops = times used.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173635,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "01/28/2021 02:59:57",
          "content": "<p><a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> yes. LB decreases but not fluctuate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1171894,
      "author_name": "darknesszx",
      "author_url": "",
      "post_date": "01/27/2021 06:55:45",
      "content": "<p>In most cases, TTA is effective</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1172055,
      "author_name": "vkehfdl1",
      "author_url": "",
      "post_date": "01/27/2021 08:40:06",
      "content": "<p>In my experience, TTA is effective for LB score. For TTA, I used only light aug, too. But, as <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> said below, I don't know how TTA is really effective for our private LB score. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1173230,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "01/27/2021 18:34:45",
      "content": "<p>In previous computer vision competitions, TTA was mostly applied in 1st place solutions</p>\n<p>My personal experience underlines that, TTA did increase public LB aswell as private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1173369,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "01/27/2021 20:29:16",
          "content": "<p>Thanks for tips, I am just confused with this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1173476,
      "author_name": "luqing2",
      "author_url": "",
      "post_date": "01/27/2021 23:26:19",
      "content": "<p>Hi yimin,   great achievement you have made.   </p>\n<p>I saw people using seresnext50. but I checked keras applications,  it does not have seresnext50, only had resnet50.  I wonder how would you create model to use seresnext50 in keras?   Really beginner question, don't laugh </p>",
      "votes": null,
      "replies": [
        {
          "id": 1173639,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "01/28/2021 03:00:44",
          "content": "<p>Thanks. I used pytorch to implement the seresnext50.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1184454,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "02/03/2021 14:33:05",
          "content": "<p><a href=\"https://www.kaggle.com/luqing2\" target=\"_blank\">@luqing2</a> You can use seresnext50 in keras by installing these GitHub repositories: <a href=\"https://github.com/qubvel/classification_models\" target=\"_blank\">Classification models </a> and <a href=\"https://github.com/keras-team/keras-applications\" target=\"_blank\">Keras applications </a>.<br>\nFor training you can keep the internet on and install the libraries directly:</p>\n<pre><code>!pip install image-classifiers\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(512,512,3))\n\n    base = seresnext50(input_shape=(512,512,3),weights = 'imagenet', include_top = False)\n\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(5,activation='softmax')(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n</code></pre>\n<p>In inference, with the internet off, you need to create kaggle datasets of those 2 repositories and attach them to your notebook:</p>\n<pre><code>import sys\nsys.path.append('../input/keras-applications/')\nsys.path.append('../input/keras-seresnext-50-101/')\nimport keras_applications as ka\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1174082,
      "author_name": "tiboas",
      "author_url": "",
      "post_date": "01/28/2021 08:54:27",
      "content": "<p>Do you apply a subset of the augmentation used in training or do you design TTA independently of training? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1186474,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "02/04/2021 20:43:45",
          "content": "<p><a href=\"https://www.kaggle.com/tiboas\" target=\"_blank\">@tiboas</a> it is subset of the aug used in training</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1174413,
      "author_name": "yindachen",
      "author_url": "",
      "post_date": "01/28/2021 13:07:03",
      "content": "<p>In my attempt, I chose TTA from 2 to 8, and the LB score was slightly improved.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1174423,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/28/2021 13:18:15",
          "content": "<p>Thats right, normally higher TTA increases LB but therefore takes longer to submit due to higher amount of inferences.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1175182,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "01/29/2021 01:40:54",
      "content": "<p>Not work from my side. Only tried Resnet34 for experiment with different tta and you can check <a href=\"https://www.kaggle.com/leonshangguan/inference-test\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1175578,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "01/29/2021 08:28:34",
      "content": "<p>TTA boosts my score<br>\nMy EfficientNEt5, 512x512 size without TTA -&gt; 0.898<br>\n3x TTA -&gt; 0.901</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186468,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "02/04/2021 20:42:39",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Did you use the TTA same as the aug used in training?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1184461,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "02/03/2021 14:36:00",
      "content": "<p><a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> In addition to coarseout, cutout, HueSaturationValue… Did you use cutmix, mixup… to get LB=0.903 with seresnext50?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186461,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "02/04/2021 20:41:22",
          "content": "<p><a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> no, cutmix decreases both CV and LB score of my model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1185720,
      "author_name": "clw5180",
      "author_url": "",
      "post_date": "02/04/2021 10:19:14",
      "content": "<p>Do you use the timm's seresnext50 or pretrainedmodel's seresnext50? Thanks a lot! My effb4 is same as yours, but seresnext50 is much lower.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186459,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "02/04/2021 20:38:47",
          "content": "<p>I used pretrained seresnext50</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1171773": "I have taken several days to train my models. 01/23 is my first day to probe the LB score.\n\nAfter some experiments, Here are my current results.\n\nTraining 2 models (seresnext50, effb4ns) with,\n\nimage size = 512,\nmerged dataset (drop the duplicates based on [this notebook](https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan) by @graf10a),\nbi-tempered loss,\nA little bit heavy random aug,\nno valid time aug,\nstratified 5-fold\n\nThe results are,\n\n|| Seresnext50 |Effb4ns  |\n| --- | --- |\n|CV without TTA| 0.895+ |0.891+ |\n|LB without TTA| 0.903 | 0.900 |\n|LB with 5x light TTA|0.900  |0.902  |\n\nlight TTA means only doing flips, rotates, and crops, while in training process doing coarseout, cutout, HueSaturationValue, blur and so on.\n\nThe baseline seresnext50 and effb4ns scored LB 0.899 and 0.898 respectively with CV scores 0.891+ and 0.890+.\n\nI have not done many experiments (will do). But, from the current results, it seems TTA randomly affects the LB socre. Without TTA, we can see that the CV and LB are probably aligned. Or, probably I am using the wrong way to implement TTA.\n\nSo, does TTA really work for this competition? Or, should we trust CV with valid time aug? Is it useful to analyze the OOF with and without valid time aug?",
    "1171803": "IMO \n\nTTA done correctly always results is a better result.  We don't know how you did your TTA or how many loops you made.  My limited testing indicates that 10 is the minimum number to establish a stable result.\n\nStatistics done without understanding variation is always going to lead to poor decisions.\n\nThe test set is 15,000 images.  If my math is correct the LB score is reported for 4650 images.\n\n0.903 LB score is 4198 images that are correct\n0.900 LB score is  4185 images that are correct\n\nDo you believe 13 images is a statistically valid difference for this noisy data set?\n\nCan you use the eye ball test to decide  a 0.003 difference on 4650 images?",
    "1171894": "In most cases, TTA is effective",
    "1172055": "In my experience, TTA is effective for LB score. For TTA, I used only light aug, too. But, as @pcjimmmy said below, I don't know how TTA is really effective for our private LB score.",
    "1172850": "Hi @pcjimmmy thanks for your tips. I will try 10+ TTA on both CV and LB.\n\nYes, you might be right. You remind me that there are noisy labels in hidden test set. We already find an example in graf10a's notebook, so I truly believe it. \n\nThe decreased 0.003 may not be the problem of TTA but the problem of noisy labels. Because, after TTA, the model can predict an image correctly but it was labeled improperly. So, it would be scored 0.\n\nBut, the question is how noisy the test set is. If there are lots of noisy labels but we have a robust model that generalizes well, the score may be a bit lower than a wrong model.",
    "1172961": "I am doing all my submissions now with 3 loops - once I find the models that I love than I will increment the number until submission timeout.",
    "1173012": "woshifym Which augmentations do you use for TTA? I had random rotations and cropping but it made my leaderboard scores inconsistent.",
    "1173230": "In previous computer vision competitions, TTA was mostly applied in 1st place solutions\n\nMy personal experience underlines that, TTA did increase public LB aswell as private LB.",
    "1173367": "As I mentioned, I just used simple random flips, rotates, and crops.",
    "1173369": "Thanks for tips, I am just confused with this competition.",
    "1173383": "Hi @pcjimmmy what did you mean loops? Is it about how many times TTA used?",
    "1173390": "yimin22 For crops, are you doing random crops? If so, does your LB score fluctuate a lot?",
    "1173476": "Hi yimin,   great achievement you have made.   \n\nI saw people using seresnext50. but I checked keras applications,  it does not have seresnext50, only had resnet50.  I wonder how would you create model to use seresnext50 in keras?   Really beginner question, don't laugh",
    "1173549": "Loops = times used.",
    "1173635": "ayu055 yes. LB decreases but not fluctuate.",
    "1173639": "Thanks. I used pytorch to implement the seresnext50.",
    "1174082": "Do you apply a subset of the augmentation used in training or do you design TTA independently of training?",
    "1174413": "In my attempt, I chose TTA from 2 to 8, and the LB score was slightly improved.",
    "1174423": "Thats right, normally higher TTA increases LB but therefore takes longer to submit due to higher amount of inferences.",
    "1175182": "Not work from my side. Only tried Resnet34 for experiment with different tta and you can check [here](https://www.kaggle.com/leonshangguan/inference-test).",
    "1175578": "TTA boosts my score\nMy EfficientNEt5, 512x512 size without TTA -> 0.898\n3x TTA -> 0.901",
    "1184454": "luqing2 You can use seresnext50 in keras by installing these GitHub repositories: [Classification models ](https://github.com/qubvel/classification_models) and [Keras applications ](https://github.com/keras-team/keras-applications).\nFor training you can keep the internet on and install the libraries directly:\n```\n!pip install image-classifiers\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(512,512,3))\n\n    base = seresnext50(input_shape=(512,512,3),weights = 'imagenet', include_top = False)\n    \n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(5,activation='softmax')(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n```\n\nIn inference, with the internet off, you need to create kaggle datasets of those 2 repositories and attach them to your notebook:\n\n```\nimport sys\nsys.path.append('../input/keras-applications/')\nsys.path.append('../input/keras-seresnext-50-101/')\nimport keras_applications as ka\nfrom classification_models.tfkeras import Classifiers\nseresnext50 , _ = Classifiers.get('seresnext50')\n```",
    "1184461": "woshifym In addition to coarseout, cutout, HueSaturationValue... Did you use cutmix, mixup... to get LB=0.903 with seresnext50?",
    "1185720": "Do you use the timm's seresnext50 or pretrainedmodel's seresnext50? Thanks a lot! My effb4 is same as yours, but seresnext50 is much lower.",
    "1186459": "I used pretrained seresnext50",
    "1186461": "amiiiney no, cutmix decreases both CV and LB score of my model.",
    "1186468": "deepkim Did you use the TTA same as the aug used in training?",
    "1186474": "tiboas it is subset of the aug used in training"
  },
  "source": "meta"
}