{
  "id": 198241,
  "title": "[11/26 Updated] Pytorch Baseline Starter (Val: 0.932, Public LB: 0.900)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198241",
  "author_name": "",
  "post_date": "2020-11-20T11:10:22.033984Z",
  "votes": 81,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Dear Kagglers,</p>\n<p>I've shared two kernels showing how to do training and inference in this competition. I'm going to also summarize some skills here to help new coming kagglers.</p>\n<h4><strong>Train</strong>: <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug</a></h4>\n<ol>\n<li>Use Pytorch built-in Automatic Mixed Precision (AMP)</li>\n<li>Batch accumulation</li>\n<li>Transfer learning with efficientnet b0</li>\n<li>Heavy train-time augmentation: h\\vflip, coarsedropout, … to increase generalization</li>\n</ol>\n<p>With 1 and 2, the model would train faster and you could use a larger model with effectively larger batch size.</p>\n<h4><strong>Inference</strong>: <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta</a></h4>\n<ol>\n<li>Checkpoint result averaging</li>\n<li>Test time augmentation</li>\n</ol>\n<h4><strong>[11/21 Update]</strong></h4>\n<p><strong>Train</strong>: </p>\n<ol>\n<li>Use adam and learning rate scheduler setting from <a href=\"https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training</a></li>\n<li>Use noisy-student pretrained efficientnet b3</li>\n<li>More augmentations</li>\n</ol>\n<p><strong>Valid</strong>: </p>\n<ol>\n<li>Use last 4 epochs' checkpoints + TTA 3 times</li>\n<li>More augmentation for TTA</li>\n<li>Logloss\\Accuracy Improvements: 0.44460/0.84650 to 0.26788/0.91192</li>\n</ol>\n<p><strong>Public LB</strong></p>\n<ol>\n<li>Improve from 0.853 to 0.895</li>\n</ol>\n<h4><strong>[11/24 Update]</strong></h4>\n<p><strong>Train</strong>: </p>\n<ol>\n<li>Scale up to test: Use image size=512, noisy-student pretrained efficientnet b4</li>\n</ol>\n<p><strong>Valid</strong>: </p>\n<ol>\n<li>Logloss\\Accuracy Improvements: 0.26788/0.91192 to 0.23218/0.92804</li>\n</ol>\n<p><strong>Public LB</strong></p>\n<ol>\n<li>Improve from 0.895 to 0.899</li>\n</ol>\n<h4><strong>[11/26 Update]</strong></h4>\n<p><strong>Valid</strong>: </p>\n<ol>\n<li>Use the last 4 epochs with 3 TTA</li>\n<li>Logloss\\Accuracy Improvements: 0.23218/0.92804 to 0.21773/0.93224</li>\n</ol>\n<p><strong>Public LB</strong></p>\n<ol>\n<li>Improve from 0.899 to 0.900</li>\n</ol>\n<p>Feel free to advise &amp; ask questions :)</p>",
  "messages": [
    {
      "id": "1084762",
      "postDate": "11/20/2020 11:10:22",
      "content": "<p>Dear Kagglers,</p>\n<p>I've shared two kernels showing how to do training and inference in this competition. I'm going to also summarize some skills here to help new coming kagglers.</p>\n<h4><strong>Train</strong>: <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug</a></h4>\n<ol>\n<li>Use Pytorch built-in Automatic Mixed Precision (AMP)</li>\n<li>Batch accumulation</li>\n<li>Transfer learning with efficientnet b0</li>\n<li>Heavy train-time augmentation: h\\vflip, coarsedropout, … to increase generalization</li>\n</ol>\n<p>With 1 and 2, the model would train faster and you could use a larger model with effectively larger batch size.</p>\n<h4><strong>Inference</strong>: <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta</a></h4>\n<ol>\n<li>Checkpoint result averaging</li>\n<li>Test time augmentation</li>\n</ol>\n<h4><strong>[11/21 Update]</strong></h4>\n<p><strong>Train</strong>: </p>\n<ol>\n<li>Use adam and learning rate scheduler setting from <a href=\"https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training</a></li>\n<li>Use noisy-student pretrained efficientnet b3</li>\n<li>More augmentations</li>\n</ol>\n<p><strong>Valid</strong>: </p>\n<ol>\n<li>Use last 4 epochs' checkpoints + TTA 3 times</li>\n<li>More augmentation for TTA</li>\n<li>Logloss\\Accuracy Improvements: 0.44460/0.84650 to 0.26788/0.91192</li>\n</ol>\n<p><strong>Public LB</strong></p>\n<ol>\n<li>Improve from 0.853 to 0.895</li>\n</ol>\n<h4><strong>[11/24 Update]</strong></h4>\n<p><strong>Train</strong>: </p>\n<ol>\n<li>Scale up to test: Use image size=512, noisy-student pretrained efficientnet b4</li>\n</ol>\n<p><strong>Valid</strong>: </p>\n<ol>\n<li>Logloss\\Accuracy Improvements: 0.26788/0.91192 to 0.23218/0.92804</li>\n</ol>\n<p><strong>Public LB</strong></p>\n<ol>\n<li>Improve from 0.895 to 0.899</li>\n</ol>\n<h4><strong>[11/26 Update]</strong></h4>\n<p><strong>Valid</strong>: </p>\n<ol>\n<li>Use the last 4 epochs with 3 TTA</li>\n<li>Logloss\\Accuracy Improvements: 0.23218/0.92804 to 0.21773/0.93224</li>\n</ol>\n<p><strong>Public LB</strong></p>\n<ol>\n<li>Improve from 0.899 to 0.900</li>\n</ol>\n<p>Feel free to advise &amp; ask questions :)</p>",
      "rawMarkdown": "Dear Kagglers,\n\nI've shared two kernels showing how to do training and inference in this competition. I'm going to also summarize some skills here to help new coming kagglers.\n\n#### **Train**: https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug \n1. Use Pytorch built-in Automatic Mixed Precision (AMP)\n2. Batch accumulation\n3. Transfer learning with efficientnet b0\n4. Heavy train-time augmentation: h\\vflip, coarsedropout, ... to increase generalization\n\nWith 1 and 2, the model would train faster and you could use a larger model with effectively larger batch size.\n\n#### **Inference**: https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta\n1. Checkpoint result averaging\n2. Test time augmentation\n\n#### **[11/21 Update]**\n**Train**: \n1. Use adam and learning rate scheduler setting from https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training\n2. Use noisy-student pretrained efficientnet b3\n3. More augmentations\n\n**Valid**: \n1. Use last 4 epochs' checkpoints + TTA 3 times\n2. More augmentation for TTA\n3. Logloss\\Accuracy Improvements: 0.44460/0.84650 to 0.26788/0.91192\n\n**Public LB**\n1. Improve from 0.853 to 0.895\n\n#### **[11/24 Update]**\n**Train**: \n1. Scale up to test: Use image size=512, noisy-student pretrained efficientnet b4\n\n**Valid**: \n1. Logloss\\Accuracy Improvements: 0.26788/0.91192 to 0.23218/0.92804\n\n**Public LB**\n1. Improve from 0.895 to 0.899\n\n#### **[11/26 Update]**\n**Valid**: \n1. Use the last 4 epochs with 3 TTA\n2. Logloss\\Accuracy Improvements: 0.23218/0.92804 to 0.21773/0.93224\n\n**Public LB**\n1. Improve from 0.899 to 0.900\n\nFeel free to advise & ask questions :)",
      "votes": null
    },
    {
      "id": "1110763",
      "postDate": "12/13/2020 03:52:49",
      "content": "<p>What does \"Use the last 4 epochs with 3 TTA\" mean ? </p>",
      "rawMarkdown": "What does \"Use the last 4 epochs with 3 TTA\" mean ?",
      "votes": null
    },
    {
      "id": "1110867",
      "postDate": "12/13/2020 06:41:41",
      "content": "<p>In his code, he trained 10 epochs. <br>\n\"Use the last 4 epochs\" means he chose last 4 models (i.e. 6, 7, 8, 9th models) for inference.<br>\nAnd he applied \"TTA (Test Time Augmentation)\"</p>",
      "rawMarkdown": "In his code, he trained 10 epochs. \n\"Use the last 4 epochs\" means he chose last 4 models (i.e. 6, 7, 8, 9th models) for inference.\nAnd he applied \"TTA (Test Time Augmentation)\"",
      "votes": null
    },
    {
      "id": "1112306",
      "postDate": "12/14/2020 13:12:17",
      "content": "<p>Please share any notebook that use log loss with pytorch. Thanks in advanced.   </p>",
      "rawMarkdown": "Please share any notebook that use log loss with pytorch. Thanks in advanced.",
      "votes": null
    },
    {
      "id": "1146739",
      "postDate": "01/10/2021 02:16:29",
      "content": "<p>Thank you for your share.I feel confused that i got val/87.5,but the LB got only 82.7,how could it be such a difference</p>",
      "rawMarkdown": "Thank you for your share.I feel confused that i got val/87.5,but the LB got only 82.7,how could it be such a difference",
      "votes": null
    },
    {
      "id": "1156101",
      "postDate": "01/16/2021 23:33:34",
      "content": "<p>Thanks for your methods and notebook sharing. Do you apply the model from any paper ? Do you have the paper reference? Thanks,</p>",
      "rawMarkdown": "Thanks for your methods and notebook sharing. Do you apply the model from any paper ? Do you have the paper reference? Thanks,",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1110763,
      "author_name": "nishikattt",
      "author_url": "",
      "post_date": "12/13/2020 03:52:49",
      "content": "<p>What does \"Use the last 4 epochs with 3 TTA\" mean ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1110867,
          "author_name": "jinkyh",
          "author_url": "",
          "post_date": "12/13/2020 06:41:41",
          "content": "<p>In his code, he trained 10 epochs. <br>\n\"Use the last 4 epochs\" means he chose last 4 models (i.e. 6, 7, 8, 9th models) for inference.<br>\nAnd he applied \"TTA (Test Time Augmentation)\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112306,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "12/14/2020 13:12:17",
      "content": "<p>Please share any notebook that use log loss with pytorch. Thanks in advanced.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146739,
      "author_name": "cnruiwang",
      "author_url": "",
      "post_date": "01/10/2021 02:16:29",
      "content": "<p>Thank you for your share.I feel confused that i got val/87.5,but the LB got only 82.7,how could it be such a difference</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156101,
      "author_name": "denisechendd",
      "author_url": "",
      "post_date": "01/16/2021 23:33:34",
      "content": "<p>Thanks for your methods and notebook sharing. Do you apply the model from any paper ? Do you have the paper reference? Thanks,</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1084762": "Dear Kagglers,\n\nI've shared two kernels showing how to do training and inference in this competition. I'm going to also summarize some skills here to help new coming kagglers.\n\n#### **Train**: https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug \n1. Use Pytorch built-in Automatic Mixed Precision (AMP)\n2. Batch accumulation\n3. Transfer learning with efficientnet b0\n4. Heavy train-time augmentation: h\\vflip, coarsedropout, ... to increase generalization\n\nWith 1 and 2, the model would train faster and you could use a larger model with effectively larger batch size.\n\n#### **Inference**: https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta\n1. Checkpoint result averaging\n2. Test time augmentation\n\n#### **[11/21 Update]**\n**Train**: \n1. Use adam and learning rate scheduler setting from https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training\n2. Use noisy-student pretrained efficientnet b3\n3. More augmentations\n\n**Valid**: \n1. Use last 4 epochs' checkpoints + TTA 3 times\n2. More augmentation for TTA\n3. Logloss\\Accuracy Improvements: 0.44460/0.84650 to 0.26788/0.91192\n\n**Public LB**\n1. Improve from 0.853 to 0.895\n\n#### **[11/24 Update]**\n**Train**: \n1. Scale up to test: Use image size=512, noisy-student pretrained efficientnet b4\n\n**Valid**: \n1. Logloss\\Accuracy Improvements: 0.26788/0.91192 to 0.23218/0.92804\n\n**Public LB**\n1. Improve from 0.895 to 0.899\n\n#### **[11/26 Update]**\n**Valid**: \n1. Use the last 4 epochs with 3 TTA\n2. Logloss\\Accuracy Improvements: 0.23218/0.92804 to 0.21773/0.93224\n\n**Public LB**\n1. Improve from 0.899 to 0.900\n\nFeel free to advise & ask questions :)",
    "1110763": "What does \"Use the last 4 epochs with 3 TTA\" mean ?",
    "1110867": "In his code, he trained 10 epochs. \n\"Use the last 4 epochs\" means he chose last 4 models (i.e. 6, 7, 8, 9th models) for inference.\nAnd he applied \"TTA (Test Time Augmentation)\"",
    "1112306": "Please share any notebook that use log loss with pytorch. Thanks in advanced.",
    "1146739": "Thank you for your share.I feel confused that i got val/87.5,but the LB got only 82.7,how could it be such a difference",
    "1156101": "Thanks for your methods and notebook sharing. Do you apply the model from any paper ? Do you have the paper reference? Thanks,"
  },
  "source": "meta"
}