{
  "id": 210016,
  "title": "Other Baseline Scores (SE-ResNeXt, ResNeSt, RegNetY)",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/210016",
  "author_name": "",
  "post_date": "2021-01-09T11:12:17.195250700Z",
  "votes": 75,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Thanks to <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> , we know the baseline scores by EfficientNet variants in this topic:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950</a></p>\n<p>Here I share you the scores by some other models.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size</th>\n<th>OOF Score</th>\n<th>Public LB(5-fold avg)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SE-ResNeXt50_32x4d</td>\n<td>512x512</td>\n<td>0.936</td>\n<td>0.954</td>\n</tr>\n<tr>\n<td>SE-ResNeXt101_32x4d</td>\n<td>512x512</td>\n<td>0.940</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>ResNeSt50-fast</td>\n<td>512x512</td>\n<td>0.939</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>ResNeSt101</td>\n<td>512x512</td>\n<td>0.944</td>\n<td><strong>0.959</strong></td>\n</tr>\n<tr>\n<td>ResNeSt200</td>\n<td>512x512</td>\n<td>0.946</td>\n<td>0.957</td>\n</tr>\n<tr>\n<td>RegNetY_032</td>\n<td>512x512</td>\n<td>0.941</td>\n<td>0.954</td>\n</tr>\n<tr>\n<td>RegNetY_064</td>\n<td>512x512</td>\n<td>0.943</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>RegNetY_080</td>\n<td>512x512</td>\n<td>0.945</td>\n<td><strong>0.959</strong></td>\n</tr>\n<tr>\n<td>RegNetY_120</td>\n<td>512x512</td>\n<td>0.945</td>\n<td>0.958</td>\n</tr>\n</tbody>\n</table>\n<p>I use <a href=\"https://www.kaggle.com/ttahara/resnest-package\" target=\"_blank\">ResNeSt Package</a> and <a href=\"https://www.kaggle.com/yasufuminakama/pytorch-image-models\" target=\"_blank\">PyTorch Image Models(timm)</a> for training.</p>\n<p>I didn't use special architectures, training techniques, and so on. There is still a lot of room for improvement.</p>\n<p>Among these models, especially big models(ResNeSt200, RegNetY_120) seems to overfit.  <br>\nThis can be prevented by heavy data augmentation like <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> 's benchmark:  <br>\n<a href=\"https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\" target=\"_blank\">https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965</a></p>\n<p>It may be possible to improve these models by special techniques like <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's 3 stage training:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577</a></p>\n<p>I would be glad that this topic helps you :)</p>",
  "messages": [
    {
      "id": "1145826",
      "postDate": "01/09/2021 11:12:17",
      "content": "<p>Thanks to <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> , we know the baseline scores by EfficientNet variants in this topic:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950</a></p>\n<p>Here I share you the scores by some other models.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>image size</th>\n<th>OOF Score</th>\n<th>Public LB(5-fold avg)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SE-ResNeXt50_32x4d</td>\n<td>512x512</td>\n<td>0.936</td>\n<td>0.954</td>\n</tr>\n<tr>\n<td>SE-ResNeXt101_32x4d</td>\n<td>512x512</td>\n<td>0.940</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>ResNeSt50-fast</td>\n<td>512x512</td>\n<td>0.939</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>ResNeSt101</td>\n<td>512x512</td>\n<td>0.944</td>\n<td><strong>0.959</strong></td>\n</tr>\n<tr>\n<td>ResNeSt200</td>\n<td>512x512</td>\n<td>0.946</td>\n<td>0.957</td>\n</tr>\n<tr>\n<td>RegNetY_032</td>\n<td>512x512</td>\n<td>0.941</td>\n<td>0.954</td>\n</tr>\n<tr>\n<td>RegNetY_064</td>\n<td>512x512</td>\n<td>0.943</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>RegNetY_080</td>\n<td>512x512</td>\n<td>0.945</td>\n<td><strong>0.959</strong></td>\n</tr>\n<tr>\n<td>RegNetY_120</td>\n<td>512x512</td>\n<td>0.945</td>\n<td>0.958</td>\n</tr>\n</tbody>\n</table>\n<p>I use <a href=\"https://www.kaggle.com/ttahara/resnest-package\" target=\"_blank\">ResNeSt Package</a> and <a href=\"https://www.kaggle.com/yasufuminakama/pytorch-image-models\" target=\"_blank\">PyTorch Image Models(timm)</a> for training.</p>\n<p>I didn't use special architectures, training techniques, and so on. There is still a lot of room for improvement.</p>\n<p>Among these models, especially big models(ResNeSt200, RegNetY_120) seems to overfit.  <br>\nThis can be prevented by heavy data augmentation like <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> 's benchmark:  <br>\n<a href=\"https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\" target=\"_blank\">https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965</a></p>\n<p>It may be possible to improve these models by special techniques like <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's 3 stage training:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577</a></p>\n<p>I would be glad that this topic helps you :)</p>",
      "rawMarkdown": "Thanks to @xhlulu , we know the baseline scores by EfficientNet variants in this topic:\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\n\nHere I share you the scores by some other models.\n\n| model                             | image size | OOF Score | Public LB(5-fold avg) |\n|:-----------------------:|:-----------:|:------------|:---------:|\n| SE-ResNeXt50_32x4d  | 512x512 | 0.936 | 0.954 |\n| SE-ResNeXt101_32x4d | 512x512 | 0.940 | 0.956 |\n| ResNeSt50-fast            | 512x512 | 0.939 | 0.956 |\n| ResNeSt101                   | 512x512 | 0.944 | **0.959** |\n| ResNeSt200                   | 512x512 | 0.946 | 0.957 |\n| RegNetY_032  | 512x512 | 0.941 | 0.954 |\n| RegNetY_064 | 512x512 | 0.943 | 0.956 |\n| RegNetY_080 | 512x512 | 0.945 | **0.959** |\n| RegNetY_120 | 512x512 | 0.945 | 0.958 |\n\nI use [ResNeSt Package](https://www.kaggle.com/ttahara/resnest-package) and [PyTorch Image Models(timm)](https://www.kaggle.com/yasufuminakama/pytorch-image-models) for training.\n\n\nI didn't use special architectures, training techniques, and so on. There is still a lot of room for improvement.\n\nAmong these models, especially big models(ResNeSt200, RegNetY_120) seems to overfit.  \nThis can be prevented by heavy data augmentation like @underwearfitting 's benchmark:  \nhttps://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\n\nIt may be possible to improve these models by special techniques like @yasufuminakama 's 3 stage training:\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\n\n I would be glad that this topic helps you :)",
      "votes": null
    },
    {
      "id": "1145873",
      "postDate": "01/09/2021 11:50:47",
      "content": "<p>Good job! Can you share how many lr do you use?And the lr scheduler do you use?Thanks!</p>",
      "rawMarkdown": "Good job! Can you share how many lr do you use?And the lr scheduler do you use?Thanks!",
      "votes": null
    },
    {
      "id": "1146694",
      "postDate": "01/10/2021 00:35:17",
      "content": "<p>I use 1e-3 for small models and 3e-4 for others.</p>",
      "rawMarkdown": "I use 1e-3 for small models and 3e-4 for others.",
      "votes": null
    },
    {
      "id": "1146729",
      "postDate": "01/10/2021 01:47:27",
      "content": "<p>new release <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb50f453e4cd2be77326399e0828ce60c%2FSelection_026.png?generation=1610243230213152&amp;alt=media\" alt=\"\"> </p>",
      "rawMarkdown": "new release \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb50f453e4cd2be77326399e0828ce60c%2FSelection_026.png?generation=1610243230213152&alt=media)",
      "votes": null
    },
    {
      "id": "1148091",
      "postDate": "01/10/2021 22:08:07",
      "content": "<p>I've tried it, didn't work well for me. </p>",
      "rawMarkdown": "I've tried it, didn't work well for me.",
      "votes": null
    },
    {
      "id": "1148093",
      "postDate": "01/10/2021 22:10:32",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for this info. I was reading ResNeSt paper-work today and wanted to try. I think to make ResNeSt-200D work, it may need a bit larger dimension. </p>",
      "rawMarkdown": "Thanks, @ttahara for this info. I was reading ResNeSt paper-work today and wanted to try. I think to make ResNeSt-200D work, it may need a bit larger dimension.",
      "votes": null
    },
    {
      "id": "1148567",
      "postDate": "01/11/2021 08:16:48",
      "content": "<p>I also tried SE-ResNet-152D by using <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's 3 stages training (same hyperparameters as original notebooks).<br>\nIn my case, CV is 0.9528 for ResNet200D, 0.9413 for SE-ResNet-152D. Both are the results of single fold.</p>",
      "rawMarkdown": "I also tried SE-ResNet-152D by using @yasufuminakama 's 3 stages training (same hyperparameters as original notebooks).\nIn my case, CV is 0.9528 for ResNet200D, 0.9413 for SE-ResNet-152D. Both are the results of single fold.",
      "votes": null
    },
    {
      "id": "1149729",
      "postDate": "01/12/2021 05:03:36",
      "content": "<p>Hi,did you test the efficientnet_b7 using <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's 3 stages training? <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> </p>",
      "rawMarkdown": "Hi,did you test the efficientnet_b7 using @yasufuminakama 's 3 stages training? @yosukeyama",
      "votes": null
    },
    {
      "id": "1149760",
      "postDate": "01/12/2021 05:56:32",
      "content": "<p>Yes, I tried b5 and b7.<br>\nCV is 0.950 by b5, 0.944 by b7.<br>\nI think it is possible to make the models better by parameter tuning, especially for b7.</p>",
      "rawMarkdown": "Yes, I tried b5 and b7.\nCV is 0.950 by b5, 0.944 by b7.\nI think it is possible to make the models better by parameter tuning, especially for b7.",
      "votes": null
    },
    {
      "id": "1150253",
      "postDate": "01/12/2021 13:25:31",
      "content": "<p>Hmm, that requires more GPUs 😂</p>",
      "rawMarkdown": "Hmm, that requires more GPUs 😂",
      "votes": null
    },
    {
      "id": "1150310",
      "postDate": "01/12/2021 14:03:43",
      "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> Thanks for your reply!Can your share more training details like lr_scheduler or learning rate.Thanks!</p>",
      "rawMarkdown": "yosukeyama Thanks for your reply!Can your share more training details like lr_scheduler or learning rate.Thanks!",
      "votes": null
    },
    {
      "id": "1150383",
      "postDate": "01/12/2021 14:46:55",
      "content": "<p>Almost the same parameters as the original notebooks.<br>\nI might change some augmentations. Anyway, in same settings, B5 was as good as ResNet200D (slightly ResNet200D was better).<br>\nCertainly, I think both the model can be improved much more, because I have not tried enough parameter tuning for the models yet.</p>",
      "rawMarkdown": "Almost the same parameters as the original notebooks.\nI might change some augmentations. Anyway, in same settings, B5 was as good as ResNet200D (slightly ResNet200D was better).\nCertainly, I think both the model can be improved much more, because I have not tried enough parameter tuning for the models yet.",
      "votes": null
    },
    {
      "id": "1150611",
      "postDate": "01/12/2021 17:46:50",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> are you guys getting consistent results i.e. CV/LB with your models? Both CV and LB varies with my initial models right now i.e. the same run without changes produce different result. I am using only flipped image augmentation.</p>",
      "rawMarkdown": "hengck23, @ttahara, @yosukeyama are you guys getting consistent results i.e. CV/LB with your models? Both CV and LB varies with my initial models right now i.e. the same run without changes produce different result. I am using only flipped image augmentation.",
      "votes": null
    },
    {
      "id": "1151032",
      "postDate": "01/13/2021 04:48:09",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fd920927ed8a6cb81356e467d9b67a53f%2F987.jpg?generation=1610513245295349&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fd920927ed8a6cb81356e467d9b67a53f%2F987.jpg?generation=1610513245295349&alt=media)",
      "votes": null
    },
    {
      "id": "1151693",
      "postDate": "01/13/2021 13:53:46",
      "content": "<p>there is a limit to the accuracy that can be reached by low resolution (e.g. 512x512).<br>\neven with the better models, the gain may not be significant.</p>\n<p>i think we may have reached the limited.</p>\n<p>maybe it worths more to go to the next stage, think of how to use higher resolution image. e.g. crops at original resolution, combining low and high resolution, tiling the images (like WSI in tissue image) etc</p>",
      "rawMarkdown": "there is a limit to the accuracy that can be reached by low resolution (e.g. 512x512).\neven with the better models, the gain may not be significant.\n\ni think we may have reached the limited.\n\nmaybe it worths more to go to the next stage, think of how to use higher resolution image. e.g. crops at original resolution, combining low and high resolution, tiling the images (like WSI in tissue image) etc",
      "votes": null
    },
    {
      "id": "1151924",
      "postDate": "01/13/2021 16:42:45",
      "content": "<p>Its time for Ensembles</p>",
      "rawMarkdown": "Its time for Ensembles",
      "votes": null
    },
    {
      "id": "1152484",
      "postDate": "01/14/2021 07:51:49",
      "content": "<p>In my case, the improvement of Public LB by simple average is only 0.001.</p>",
      "rawMarkdown": "In my case, the improvement of Public LB by simple average is only 0.001.",
      "votes": null
    },
    {
      "id": "1170335",
      "postDate": "01/26/2021 06:36:00",
      "content": "<p>I'm curious what hardware you are using that you are able to train 5 folds of all of these models. Seems like quite a heavy lift. </p>",
      "rawMarkdown": "I'm curious what hardware you are using that you are able to train 5 folds of all of these models. Seems like quite a heavy lift.",
      "votes": null
    },
    {
      "id": "1170337",
      "postDate": "01/26/2021 06:37:20",
      "content": "<p>It seems to me like the model architecture does not have a huge impact on performance but it seems like for some reason people have seen fairly large impact based on initial learning rate and lr scheduling. I wonder why that is the case</p>",
      "rawMarkdown": "It seems to me like the model architecture does not have a huge impact on performance but it seems like for some reason people have seen fairly large impact based on initial learning rate and lr scheduling. I wonder why that is the case",
      "votes": null
    },
    {
      "id": "1189569",
      "postDate": "02/07/2021 05:49:01",
      "content": "<p>I used one Titan RTX.</p>",
      "rawMarkdown": "I used one Titan RTX.",
      "votes": null
    },
    {
      "id": "1189576",
      "postDate": "02/07/2021 05:53:24",
      "content": "<p>[Update]<br>\nI've achieved CV 0.954 and LB 0.965 by ResNeSt200 on 640x640 with more data augmentations.</p>",
      "rawMarkdown": "[Update]\nI've achieved CV 0.954 and LB 0.965 by ResNeSt200 on 640x640 with more data augmentations.",
      "votes": null
    },
    {
      "id": "1191156",
      "postDate": "02/08/2021 09:16:42",
      "content": "<p>Yes, I got 95.5 with B7.</p>",
      "rawMarkdown": "Yes, I got 95.5 with B7.",
      "votes": null
    },
    {
      "id": "1219797",
      "postDate": "02/27/2021 08:56:02",
      "content": "<p><a href=\"https://www.kaggle.com/bcwang\" target=\"_blank\">@bcwang</a> I have trained a single fold of b6 for 10 epochs, AdamW optimizer, and BCEwithLogitsLoss.<br>\nROC AUC score reached 0.904 in the 4th epoch then gradually decreased to 0.890 till the 10th epoch. Can you give any tips to solve this issue?</p>",
      "rawMarkdown": "bcwang I have trained a single fold of b6 for 10 epochs, AdamW optimizer, and BCEwithLogitsLoss.\nROC AUC score reached 0.904 in the 4th epoch then gradually decreased to 0.890 till the 10th epoch. Can you give any tips to solve this issue?",
      "votes": null
    },
    {
      "id": "1220087",
      "postDate": "02/27/2021 14:34:46",
      "content": "<p>This is the curse of DL called OverFitting. Some things you can do:<br>\n[1] Try use *train *and *test *<em>time data augmentations</em>, Augmentations boost your model in terms of robustness and stability and thus prevents also overfitting<br>\n[2] Try use scheldurer in which it will graually change learning rate depending on the val auc, saving also the best auc.<br>\n[3] Try diferent initial random seeds. Sometimes the initial state of a model may doomed it to be stacked in an initial local minimum preventing your model find better solution.<br>\n[4] Expirement with the initial image resolutions eg from 500 to 850. The image resolutios plays a vital part in every image classification model. Sometimes high resolution overfit the model since it is more vurnelable on learning noise from images, while in lower resolution the model may find only the important and vital patterns for buiding a robust decision function.</p>",
      "rawMarkdown": "This is the curse of DL called OverFitting. Some things you can do:\n[1] Try use *train *and *test **time data augmentations*, Augmentations boost your model in terms of robustness and stability and thus prevents also overfitting\n[2] Try use scheldurer in which it will graually change learning rate depending on the val auc, saving also the best auc.\n[3] Try diferent initial random seeds. Sometimes the initial state of a model may doomed it to be stacked in an initial local minimum preventing your model find better solution.\n[4] Expirement with the initial image resolutions eg from 500 to 850. The image resolutios plays a vital part in every image classification model. Sometimes high resolution overfit the model since it is more vurnelable on learning noise from images, while in lower resolution the model may find only the important and vital patterns for buiding a robust decision function.",
      "votes": null
    },
    {
      "id": "1221230",
      "postDate": "02/28/2021 19:13:56",
      "content": "<p>How you load resnest to the time : timm.create_model( _ , pretrained=False)  in  ' _' what you wrote ?  </p>",
      "rawMarkdown": "How you load resnest to the time : timm.create_model( _ , pretrained=False)  in  ' _' what you wrote ?",
      "votes": null
    },
    {
      "id": "1221386",
      "postDate": "02/28/2021 23:50:20",
      "content": "<p><code>'resnest101e'</code> or <code>'resnest200e'</code>. You can find all possible ResNeSt variants in timm <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/resnest.py\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "`'resnest101e'` or `'resnest200e'`. You can find all possible ResNeSt variants in timm [here](https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/resnest.py).",
      "votes": null
    },
    {
      "id": "1221991",
      "postDate": "03/01/2021 13:09:33",
      "content": "<p>'resnest101e' thanks it works now.  Can you suggests me tpu pytorch best baseline for analysis. Thanks in an advance.     </p>",
      "rawMarkdown": "'resnest101e' thanks it works now.  Can you suggests me tpu pytorch best baseline for analysis. Thanks in an advance.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1145873,
      "author_name": "bcwang",
      "author_url": "",
      "post_date": "01/09/2021 11:50:47",
      "content": "<p>Good job! Can you share how many lr do you use?And the lr scheduler do you use?Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146694,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "01/10/2021 00:35:17",
          "content": "<p>I use 1e-3 for small models and 3e-4 for others.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146729,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/10/2021 01:47:27",
      "content": "<p>new release <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb50f453e4cd2be77326399e0828ce60c%2FSelection_026.png?generation=1610243230213152&amp;alt=media\" alt=\"\"> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1148091,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "01/10/2021 22:08:07",
          "content": "<p>I've tried it, didn't work well for me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148567,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "01/11/2021 08:16:48",
          "content": "<p>I also tried SE-ResNet-152D by using <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's 3 stages training (same hyperparameters as original notebooks).<br>\nIn my case, CV is 0.9528 for ResNet200D, 0.9413 for SE-ResNet-152D. Both are the results of single fold.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149729,
          "author_name": "bcwang",
          "author_url": "",
          "post_date": "01/12/2021 05:03:36",
          "content": "<p>Hi,did you test the efficientnet_b7 using <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's 3 stages training? <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149760,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "01/12/2021 05:56:32",
          "content": "<p>Yes, I tried b5 and b7.<br>\nCV is 0.950 by b5, 0.944 by b7.<br>\nI think it is possible to make the models better by parameter tuning, especially for b7.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150310,
          "author_name": "bcwang",
          "author_url": "",
          "post_date": "01/12/2021 14:03:43",
          "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> Thanks for your reply!Can your share more training details like lr_scheduler or learning rate.Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150383,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "01/12/2021 14:46:55",
          "content": "<p>Almost the same parameters as the original notebooks.<br>\nI might change some augmentations. Anyway, in same settings, B5 was as good as ResNet200D (slightly ResNet200D was better).<br>\nCertainly, I think both the model can be improved much more, because I have not tried enough parameter tuning for the models yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150611,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "01/12/2021 17:46:50",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> are you guys getting consistent results i.e. CV/LB with your models? Both CV and LB varies with my initial models right now i.e. the same run without changes produce different result. I am using only flipped image augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1191156,
          "author_name": "kudzayiking",
          "author_url": "",
          "post_date": "02/08/2021 09:16:42",
          "content": "<p>Yes, I got 95.5 with B7.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219797,
          "author_name": "sanchitvj",
          "author_url": "",
          "post_date": "02/27/2021 08:56:02",
          "content": "<p><a href=\"https://www.kaggle.com/bcwang\" target=\"_blank\">@bcwang</a> I have trained a single fold of b6 for 10 epochs, AdamW optimizer, and BCEwithLogitsLoss.<br>\nROC AUC score reached 0.904 in the 4th epoch then gradually decreased to 0.890 till the 10th epoch. Can you give any tips to solve this issue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220087,
          "author_name": "",
          "author_url": "",
          "post_date": "02/27/2021 14:34:46",
          "content": "<p>This is the curse of DL called OverFitting. Some things you can do:<br>\n[1] Try use *train *and *test *<em>time data augmentations</em>, Augmentations boost your model in terms of robustness and stability and thus prevents also overfitting<br>\n[2] Try use scheldurer in which it will graually change learning rate depending on the val auc, saving also the best auc.<br>\n[3] Try diferent initial random seeds. Sometimes the initial state of a model may doomed it to be stacked in an initial local minimum preventing your model find better solution.<br>\n[4] Expirement with the initial image resolutions eg from 500 to 850. The image resolutios plays a vital part in every image classification model. Sometimes high resolution overfit the model since it is more vurnelable on learning noise from images, while in lower resolution the model may find only the important and vital patterns for buiding a robust decision function.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148093,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "01/10/2021 22:10:32",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for this info. I was reading ResNeSt paper-work today and wanted to try. I think to make ResNeSt-200D work, it may need a bit larger dimension. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1150253,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "01/12/2021 13:25:31",
          "content": "<p>Hmm, that requires more GPUs 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151032,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "01/13/2021 04:48:09",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fd920927ed8a6cb81356e467d9b67a53f%2F987.jpg?generation=1610513245295349&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1151693,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/13/2021 13:53:46",
      "content": "<p>there is a limit to the accuracy that can be reached by low resolution (e.g. 512x512).<br>\neven with the better models, the gain may not be significant.</p>\n<p>i think we may have reached the limited.</p>\n<p>maybe it worths more to go to the next stage, think of how to use higher resolution image. e.g. crops at original resolution, combining low and high resolution, tiling the images (like WSI in tissue image) etc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1151924,
      "author_name": "",
      "author_url": "",
      "post_date": "01/13/2021 16:42:45",
      "content": "<p>Its time for Ensembles</p>",
      "votes": null,
      "replies": [
        {
          "id": 1152484,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "01/14/2021 07:51:49",
          "content": "<p>In my case, the improvement of Public LB by simple average is only 0.001.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1170335,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "01/26/2021 06:36:00",
      "content": "<p>I'm curious what hardware you are using that you are able to train 5 folds of all of these models. Seems like quite a heavy lift. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1189569,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "02/07/2021 05:49:01",
          "content": "<p>I used one Titan RTX.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1170337,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "01/26/2021 06:37:20",
      "content": "<p>It seems to me like the model architecture does not have a huge impact on performance but it seems like for some reason people have seen fairly large impact based on initial learning rate and lr scheduling. I wonder why that is the case</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1189576,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "02/07/2021 05:53:24",
      "content": "<p>[Update]<br>\nI've achieved CV 0.954 and LB 0.965 by ResNeSt200 on 640x640 with more data augmentations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1221230,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "02/28/2021 19:13:56",
      "content": "<p>How you load resnest to the time : timm.create_model( _ , pretrained=False)  in  ' _' what you wrote ?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1221386,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "02/28/2021 23:50:20",
          "content": "<p><code>'resnest101e'</code> or <code>'resnest200e'</code>. You can find all possible ResNeSt variants in timm <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/resnest.py\" target=\"_blank\">here</a>. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221991,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "03/01/2021 13:09:33",
          "content": "<p>'resnest101e' thanks it works now.  Can you suggests me tpu pytorch best baseline for analysis. Thanks in an advance.     </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1145826": "Thanks to @xhlulu , we know the baseline scores by EfficientNet variants in this topic:\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\n\nHere I share you the scores by some other models.\n\n| model                             | image size | OOF Score | Public LB(5-fold avg) |\n|:-----------------------:|:-----------:|:------------|:---------:|\n| SE-ResNeXt50_32x4d  | 512x512 | 0.936 | 0.954 |\n| SE-ResNeXt101_32x4d | 512x512 | 0.940 | 0.956 |\n| ResNeSt50-fast            | 512x512 | 0.939 | 0.956 |\n| ResNeSt101                   | 512x512 | 0.944 | **0.959** |\n| ResNeSt200                   | 512x512 | 0.946 | 0.957 |\n| RegNetY_032  | 512x512 | 0.941 | 0.954 |\n| RegNetY_064 | 512x512 | 0.943 | 0.956 |\n| RegNetY_080 | 512x512 | 0.945 | **0.959** |\n| RegNetY_120 | 512x512 | 0.945 | 0.958 |\n\nI use [ResNeSt Package](https://www.kaggle.com/ttahara/resnest-package) and [PyTorch Image Models(timm)](https://www.kaggle.com/yasufuminakama/pytorch-image-models) for training.\n\n\nI didn't use special architectures, training techniques, and so on. There is still a lot of room for improvement.\n\nAmong these models, especially big models(ResNeSt200, RegNetY_120) seems to overfit.  \nThis can be prevented by heavy data augmentation like @underwearfitting 's benchmark:  \nhttps://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\n\nIt may be possible to improve these models by special techniques like @yasufuminakama 's 3 stage training:\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\n\n I would be glad that this topic helps you :)",
    "1145873": "Good job! Can you share how many lr do you use?And the lr scheduler do you use?Thanks!",
    "1146694": "I use 1e-3 for small models and 3e-4 for others.",
    "1146729": "new release \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb50f453e4cd2be77326399e0828ce60c%2FSelection_026.png?generation=1610243230213152&alt=media)",
    "1148091": "I've tried it, didn't work well for me.",
    "1148093": "Thanks, @ttahara for this info. I was reading ResNeSt paper-work today and wanted to try. I think to make ResNeSt-200D work, it may need a bit larger dimension.",
    "1148567": "I also tried SE-ResNet-152D by using @yasufuminakama 's 3 stages training (same hyperparameters as original notebooks).\nIn my case, CV is 0.9528 for ResNet200D, 0.9413 for SE-ResNet-152D. Both are the results of single fold.",
    "1149729": "Hi,did you test the efficientnet_b7 using @yasufuminakama 's 3 stages training? @yosukeyama",
    "1149760": "Yes, I tried b5 and b7.\nCV is 0.950 by b5, 0.944 by b7.\nI think it is possible to make the models better by parameter tuning, especially for b7.",
    "1150253": "Hmm, that requires more GPUs 😂",
    "1150310": "yosukeyama Thanks for your reply!Can your share more training details like lr_scheduler or learning rate.Thanks!",
    "1150383": "Almost the same parameters as the original notebooks.\nI might change some augmentations. Anyway, in same settings, B5 was as good as ResNet200D (slightly ResNet200D was better).\nCertainly, I think both the model can be improved much more, because I have not tried enough parameter tuning for the models yet.",
    "1150611": "hengck23, @ttahara, @yosukeyama are you guys getting consistent results i.e. CV/LB with your models? Both CV and LB varies with my initial models right now i.e. the same run without changes produce different result. I am using only flipped image augmentation.",
    "1151032": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fd920927ed8a6cb81356e467d9b67a53f%2F987.jpg?generation=1610513245295349&alt=media)",
    "1151693": "there is a limit to the accuracy that can be reached by low resolution (e.g. 512x512).\neven with the better models, the gain may not be significant.\n\ni think we may have reached the limited.\n\nmaybe it worths more to go to the next stage, think of how to use higher resolution image. e.g. crops at original resolution, combining low and high resolution, tiling the images (like WSI in tissue image) etc",
    "1151924": "Its time for Ensembles",
    "1152484": "In my case, the improvement of Public LB by simple average is only 0.001.",
    "1170335": "I'm curious what hardware you are using that you are able to train 5 folds of all of these models. Seems like quite a heavy lift.",
    "1170337": "It seems to me like the model architecture does not have a huge impact on performance but it seems like for some reason people have seen fairly large impact based on initial learning rate and lr scheduling. I wonder why that is the case",
    "1189569": "I used one Titan RTX.",
    "1189576": "[Update]\nI've achieved CV 0.954 and LB 0.965 by ResNeSt200 on 640x640 with more data augmentations.",
    "1191156": "Yes, I got 95.5 with B7.",
    "1219797": "bcwang I have trained a single fold of b6 for 10 epochs, AdamW optimizer, and BCEwithLogitsLoss.\nROC AUC score reached 0.904 in the 4th epoch then gradually decreased to 0.890 till the 10th epoch. Can you give any tips to solve this issue?",
    "1220087": "This is the curse of DL called OverFitting. Some things you can do:\n[1] Try use *train *and *test **time data augmentations*, Augmentations boost your model in terms of robustness and stability and thus prevents also overfitting\n[2] Try use scheldurer in which it will graually change learning rate depending on the val auc, saving also the best auc.\n[3] Try diferent initial random seeds. Sometimes the initial state of a model may doomed it to be stacked in an initial local minimum preventing your model find better solution.\n[4] Expirement with the initial image resolutions eg from 500 to 850. The image resolutios plays a vital part in every image classification model. Sometimes high resolution overfit the model since it is more vurnelable on learning noise from images, while in lower resolution the model may find only the important and vital patterns for buiding a robust decision function.",
    "1221230": "How you load resnest to the time : timm.create_model( _ , pretrained=False)  in  ' _' what you wrote ?",
    "1221386": "`'resnest101e'` or `'resnest200e'`. You can find all possible ResNeSt variants in timm [here](https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/resnest.py).",
    "1221991": "'resnest101e' thanks it works now.  Can you suggests me tpu pytorch best baseline for analysis. Thanks in an advance."
  },
  "source": "meta"
}