{
  "id": 229416,
  "title": "[Update score] Multi Label Approach Baseline [LB:0.616]",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/229416",
  "author_name": "",
  "post_date": "2021-03-30T05:14:59.909767800Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>There are only a few notebooks that have been solved as multi-label problems, so I made one.<br>\nI used the following <a href=\"https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter\" target=\"_blank\">notebook</a> as a reference.</p>\n<p>The accuracy of the multi-label solution is about the same as that of the simple solution, and I think the accuracy can be improved by post-processing.</p>\n<p>[update]<br>\nAfter updating the label, a larger model may be better.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Notebook Version (train:infer)</th>\n<th>LB</th>\n<th>memo</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Resnet50</td>\n<td>V4 : V4</td>\n<td>0.778</td>\n<td></td>\n</tr>\n<tr>\n<td>SE-ResNeXt50</td>\n<td>V6 : V5</td>\n<td>0.790</td>\n<td></td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V8 : V6</td>\n<td>0.782</td>\n<td>remove duplicates</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V11 : V7</td>\n<td>0.777</td>\n<td>More epoch, change lr_scheduler</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V14 : V8</td>\n<td>0.776</td>\n<td>torchmetrics F1</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V15 : V9</td>\n<td>0.757</td>\n<td>Focal Loss</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V16 : V11</td>\n<td>0.771</td>\n<td>iterative-stratification</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V17 : V12</td>\n<td>0.758</td>\n<td>60 epoch</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://www.kaggle.com/pegasos/plant2021-multi-label-model-training\" target=\"_blank\">Training Notebook</a><br>\n<a href=\"https://www.kaggle.com/pegasos/plant2021-multi-label-model-inference\" target=\"_blank\">Inference Notebook</a></p>",
  "messages": [
    {
      "id": "1256627",
      "postDate": "03/30/2021 05:14:59",
      "content": "<p>There are only a few notebooks that have been solved as multi-label problems, so I made one.<br>\nI used the following <a href=\"https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter\" target=\"_blank\">notebook</a> as a reference.</p>\n<p>The accuracy of the multi-label solution is about the same as that of the simple solution, and I think the accuracy can be improved by post-processing.</p>\n<p>[update]<br>\nAfter updating the label, a larger model may be better.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Notebook Version (train:infer)</th>\n<th>LB</th>\n<th>memo</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Resnet50</td>\n<td>V4 : V4</td>\n<td>0.778</td>\n<td></td>\n</tr>\n<tr>\n<td>SE-ResNeXt50</td>\n<td>V6 : V5</td>\n<td>0.790</td>\n<td></td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V8 : V6</td>\n<td>0.782</td>\n<td>remove duplicates</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V11 : V7</td>\n<td>0.777</td>\n<td>More epoch, change lr_scheduler</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V14 : V8</td>\n<td>0.776</td>\n<td>torchmetrics F1</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V15 : V9</td>\n<td>0.757</td>\n<td>Focal Loss</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V16 : V11</td>\n<td>0.771</td>\n<td>iterative-stratification</td>\n</tr>\n<tr>\n<td>Resnet50</td>\n<td>V17 : V12</td>\n<td>0.758</td>\n<td>60 epoch</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://www.kaggle.com/pegasos/plant2021-multi-label-model-training\" target=\"_blank\">Training Notebook</a><br>\n<a href=\"https://www.kaggle.com/pegasos/plant2021-multi-label-model-inference\" target=\"_blank\">Inference Notebook</a></p>",
      "rawMarkdown": "There are only a few notebooks that have been solved as multi-label problems, so I made one.\nI used the following [notebook](https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter) as a reference.\n\nThe accuracy of the multi-label solution is about the same as that of the simple solution, and I think the accuracy can be improved by post-processing.\n\n[update]\nAfter updating the label, a larger model may be better.\n\n| Model | Notebook Version (train:infer)| LB | memo |\n| --- | --- | --- | --- |\n| Resnet50 | V4 : V4 | 0.778 | |\n| SE-ResNeXt50 | V6 : V5 | 0.790 | |\n| Resnet50 | V8 : V6 | 0.782 | remove duplicates |\n| Resnet50 | V11 : V7 | 0.777 | More epoch, change lr_scheduler |\n| Resnet50 | V14 : V8 | 0.776 | torchmetrics F1 |\n| Resnet50 | V15 : V9 | 0.757 | Focal Loss |\n| Resnet50 | V16 : V11 | 0.771 | iterative-stratification |\n| Resnet50 | V17 : V12 | 0.758 | 60 epoch |\n\n[Training Notebook](https://www.kaggle.com/pegasos/plant2021-multi-label-model-training)\n[Inference Notebook](https://www.kaggle.com/pegasos/plant2021-multi-label-model-inference)",
      "votes": null
    },
    {
      "id": "1256921",
      "postDate": "03/30/2021 11:18:12",
      "content": "<p>Thanks for sharing. I was adapting your previous notebooks to multi-labels :-) <br>\nLooks like in <a href=\"https://pytorch-lightning.readthedocs.io/en/stable/extensions/metrics.html\" target=\"_blank\">the F1 metric definition</a> the default is <code>multilabel=False</code>, so you have to deliberately set it to true <br>\n<code>self.metric = pl.metrics.F1(num_classes=CFG.num_classes, multilabel=True)</code></p>",
      "rawMarkdown": "Thanks for sharing. I was adapting your previous notebooks to multi-labels :-) \nLooks like in [the F1 metric definition](https://pytorch-lightning.readthedocs.io/en/stable/extensions/metrics.html) the default is `multilabel=False`, so you have to deliberately set it to true \n`self.metric = pl.metrics.F1(num_classes=CFG.num_classes, multilabel=True)`",
      "votes": null
    },
    {
      "id": "1256969",
      "postDate": "03/30/2021 12:21:37",
      "content": "<p>Thank you for pointing out the mistake. I'm glad it helps you.</p>",
      "rawMarkdown": "Thank you for pointing out the mistake. I'm glad it helps you.",
      "votes": null
    },
    {
      "id": "1257662",
      "postDate": "03/31/2021 02:50:34",
      "content": "<p>Interesting, thanks for sharing this.</p>",
      "rawMarkdown": "Interesting, thanks for sharing this.",
      "votes": null
    },
    {
      "id": "1271921",
      "postDate": "04/13/2021 04:07:06",
      "content": "<p>I have a question but not related to this.</p>\n<p>I'm training the same model with same configs using pytorch lightning with kaggle GPU. But, it takes 54 min for a single epoch. I also use AMP. But When I looked into your script, I found that it took only about 8min for a single epoch.</p>\n<p>Do you have any idea about this? </p>",
      "rawMarkdown": "I have a question but not related to this.\n\nI'm training the same model with same configs using pytorch lightning with kaggle GPU. But, it takes 54 min for a single epoch. I also use AMP. But When I looked into your script, I found that it took only about 8min for a single epoch.\n\nDo you have any idea about this?",
      "votes": null
    },
    {
      "id": "1272229",
      "postDate": "04/13/2021 10:25:29",
      "content": "<p>In my notebook, I use pre-resized images instead of the original ones.<br>\nThis reduces the time it takes to load and so on.<br>\n<a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227032\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227032</a></p>",
      "rawMarkdown": "In my notebook, I use pre-resized images instead of the original ones.\nThis reduces the time it takes to load and so on.\nhttps://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227032",
      "votes": null
    },
    {
      "id": "1273000",
      "postDate": "04/14/2021 02:03:57",
      "content": "<p>oh okay thank you, Does augmentations like cutout and affine transformations take time with albumentations.</p>",
      "rawMarkdown": "oh okay thank you, Does augmentations like cutout and affine transformations take time with albumentations.",
      "votes": null
    },
    {
      "id": "1273107",
      "postDate": "04/14/2021 05:05:26",
      "content": "<p>I know it is small compared to the time it takes to load etc., but it is a time consuming part of the process.<br>\nIf you want to speed up the process, you can use the GPU to do the augmentation.<br>\n<a href=\"https://kornia.readthedocs.io/en/latest/introduction.html\" target=\"_blank\">https://kornia.readthedocs.io/en/latest/introduction.html</a></p>",
      "rawMarkdown": "I know it is small compared to the time it takes to load etc., but it is a time consuming part of the process.\nIf you want to speed up the process, you can use the GPU to do the augmentation.\nhttps://kornia.readthedocs.io/en/latest/introduction.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1256921,
      "author_name": "buinyi",
      "author_url": "",
      "post_date": "03/30/2021 11:18:12",
      "content": "<p>Thanks for sharing. I was adapting your previous notebooks to multi-labels :-) <br>\nLooks like in <a href=\"https://pytorch-lightning.readthedocs.io/en/stable/extensions/metrics.html\" target=\"_blank\">the F1 metric definition</a> the default is <code>multilabel=False</code>, so you have to deliberately set it to true <br>\n<code>self.metric = pl.metrics.F1(num_classes=CFG.num_classes, multilabel=True)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1256969,
          "author_name": "pegasos",
          "author_url": "",
          "post_date": "03/30/2021 12:21:37",
          "content": "<p>Thank you for pointing out the mistake. I'm glad it helps you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1257662,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "03/31/2021 02:50:34",
      "content": "<p>Interesting, thanks for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1271921,
      "author_name": "kingofarmy",
      "author_url": "",
      "post_date": "04/13/2021 04:07:06",
      "content": "<p>I have a question but not related to this.</p>\n<p>I'm training the same model with same configs using pytorch lightning with kaggle GPU. But, it takes 54 min for a single epoch. I also use AMP. But When I looked into your script, I found that it took only about 8min for a single epoch.</p>\n<p>Do you have any idea about this? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1272229,
          "author_name": "pegasos",
          "author_url": "",
          "post_date": "04/13/2021 10:25:29",
          "content": "<p>In my notebook, I use pre-resized images instead of the original ones.<br>\nThis reduces the time it takes to load and so on.<br>\n<a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227032\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227032</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1273000,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "04/14/2021 02:03:57",
          "content": "<p>oh okay thank you, Does augmentations like cutout and affine transformations take time with albumentations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1273107,
          "author_name": "pegasos",
          "author_url": "",
          "post_date": "04/14/2021 05:05:26",
          "content": "<p>I know it is small compared to the time it takes to load etc., but it is a time consuming part of the process.<br>\nIf you want to speed up the process, you can use the GPU to do the augmentation.<br>\n<a href=\"https://kornia.readthedocs.io/en/latest/introduction.html\" target=\"_blank\">https://kornia.readthedocs.io/en/latest/introduction.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1256627": "There are only a few notebooks that have been solved as multi-label problems, so I made one.\nI used the following [notebook](https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter) as a reference.\n\nThe accuracy of the multi-label solution is about the same as that of the simple solution, and I think the accuracy can be improved by post-processing.\n\n[update]\nAfter updating the label, a larger model may be better.\n\n| Model | Notebook Version (train:infer)| LB | memo |\n| --- | --- | --- | --- |\n| Resnet50 | V4 : V4 | 0.778 | |\n| SE-ResNeXt50 | V6 : V5 | 0.790 | |\n| Resnet50 | V8 : V6 | 0.782 | remove duplicates |\n| Resnet50 | V11 : V7 | 0.777 | More epoch, change lr_scheduler |\n| Resnet50 | V14 : V8 | 0.776 | torchmetrics F1 |\n| Resnet50 | V15 : V9 | 0.757 | Focal Loss |\n| Resnet50 | V16 : V11 | 0.771 | iterative-stratification |\n| Resnet50 | V17 : V12 | 0.758 | 60 epoch |\n\n[Training Notebook](https://www.kaggle.com/pegasos/plant2021-multi-label-model-training)\n[Inference Notebook](https://www.kaggle.com/pegasos/plant2021-multi-label-model-inference)",
    "1256921": "Thanks for sharing. I was adapting your previous notebooks to multi-labels :-) \nLooks like in [the F1 metric definition](https://pytorch-lightning.readthedocs.io/en/stable/extensions/metrics.html) the default is `multilabel=False`, so you have to deliberately set it to true \n`self.metric = pl.metrics.F1(num_classes=CFG.num_classes, multilabel=True)`",
    "1256969": "Thank you for pointing out the mistake. I'm glad it helps you.",
    "1257662": "Interesting, thanks for sharing this.",
    "1271921": "I have a question but not related to this.\n\nI'm training the same model with same configs using pytorch lightning with kaggle GPU. But, it takes 54 min for a single epoch. I also use AMP. But When I looked into your script, I found that it took only about 8min for a single epoch.\n\nDo you have any idea about this?",
    "1272229": "In my notebook, I use pre-resized images instead of the original ones.\nThis reduces the time it takes to load and so on.\nhttps://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227032",
    "1273000": "oh okay thank you, Does augmentations like cutout and affine transformations take time with albumentations.",
    "1273107": "I know it is small compared to the time it takes to load etc., but it is a time consuming part of the process.\nIf you want to speed up the process, you can use the GPU to do the augmentation.\nhttps://kornia.readthedocs.io/en/latest/introduction.html"
  },
  "source": "meta"
}