{
  "id": 212623,
  "title": "What I am doing wrong?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212623",
  "author_name": "",
  "post_date": "2021-01-19T15:25:31.549107900Z",
  "votes": 10,
  "comment_count": 15,
  "views": 0,
  "content": "<p>At this <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111\" target=\"_blank\">post</a>, I read that multiple people get 0.900+ CV accuracy with their model, and far better LB than mine. I understand that the dataset is noisy, but others could get a much better score than I, with the effnet or resnext architecture.</p>\n<p>Here is the training cooperation table. I made the validation process from <code>k = 5</code> fold 0. I generated a .csv file, where I wrote what is fold 0, so the validation set is the same. I tested two architectures <code>resnet50_32x4d</code> and <code>effnet-b4</code>, and several methods, but somehow couldn't pass the 0.898 validation accuracy. I tried:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">Bi-Tempered Logistic Loss</a> loss</li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203271\" target=\"_blank\">Focal Cosine Loss</a></li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208887#1139841\" target=\"_blank\">RandAugment</a></li>\n<li>Up-sample the <code>`cbb', 'cbsd' ,'cgm', 'healthy'</code> classes to get the same distribution, or by 2x to reduce the class imbalance</li>\n<li>Combine the 2020 dataset with  2019 (merged column in the image)</li>\n<li>Snapmix</li>\n<li>Learning scheduler (Effnet-b4-scheduler column in the image)<br>\nNeither one approach improved my LB or CV.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fe230cf84b270c5c0e19e3bede722e06d%2FSelection_085.jpg?generation=1611069573949012&amp;alt=media\" alt=\"\"></p>\n<p>The <code>basic</code> approach is. In the last case (column), I used a scheduler.</p>\n<pre><code>Image size      512\nOptimizer       Ranger (Lookahed + RAdam)\nDataset         2020\nFold            0/5\nAugmentation    random scale, random rotate, random contrast, random hsv\n                random flip transpose, do_random_cutout\nsmooth          0.05\n</code></pre>\n<p>I tried with light augmentation based on this <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206489\" target=\"_blank\">topic</a>, so I reduced the augmentation (other configure parts same with  <code>basic</code> approach).</p>\n<pre><code>Augmentation    random flip transpose,\n</code></pre>\n<p>The  random flip transpose:</p>\n<pre><code>def do_random_flip_transpose(image):\n    if np.random.rand() &lt; 0.5:\n        image = cv2.flip(image, 0)\n    if np.random.rand() &lt; 0.5:\n        image = cv2.flip(image, 1)\n    if np.random.rand() &lt; 0.5:\n        image = image.transpose(1, 0, 2)\n    return np.ascontiguousarray(image)\n</code></pre>\n<h3>Questions:</h3>\n<ol>\n<li>What am I doing wrong?</li>\n<li>For the same validation fold, for CV 0.898, I got LB between [0.883, 0.896]. Should I entirely pay attention to the LB score?</li>\n<li>What else can I try?</li>\n</ol>\n<p>I started this competition 16 days ago and spent nearly \"full-time\" daily without any considerable achievement. I ran out of ideas, so I need external advice.</p>\n<p>Thanks for the help!</p>",
  "messages": [
    {
      "id": "1159968",
      "postDate": "01/19/2021 15:25:31",
      "content": "<p>At this <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111\" target=\"_blank\">post</a>, I read that multiple people get 0.900+ CV accuracy with their model, and far better LB than mine. I understand that the dataset is noisy, but others could get a much better score than I, with the effnet or resnext architecture.</p>\n<p>Here is the training cooperation table. I made the validation process from <code>k = 5</code> fold 0. I generated a .csv file, where I wrote what is fold 0, so the validation set is the same. I tested two architectures <code>resnet50_32x4d</code> and <code>effnet-b4</code>, and several methods, but somehow couldn't pass the 0.898 validation accuracy. I tried:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">Bi-Tempered Logistic Loss</a> loss</li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203271\" target=\"_blank\">Focal Cosine Loss</a></li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208887#1139841\" target=\"_blank\">RandAugment</a></li>\n<li>Up-sample the <code>`cbb', 'cbsd' ,'cgm', 'healthy'</code> classes to get the same distribution, or by 2x to reduce the class imbalance</li>\n<li>Combine the 2020 dataset with  2019 (merged column in the image)</li>\n<li>Snapmix</li>\n<li>Learning scheduler (Effnet-b4-scheduler column in the image)<br>\nNeither one approach improved my LB or CV.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fe230cf84b270c5c0e19e3bede722e06d%2FSelection_085.jpg?generation=1611069573949012&amp;alt=media\" alt=\"\"></p>\n<p>The <code>basic</code> approach is. In the last case (column), I used a scheduler.</p>\n<pre><code>Image size      512\nOptimizer       Ranger (Lookahed + RAdam)\nDataset         2020\nFold            0/5\nAugmentation    random scale, random rotate, random contrast, random hsv\n                random flip transpose, do_random_cutout\nsmooth          0.05\n</code></pre>\n<p>I tried with light augmentation based on this <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206489\" target=\"_blank\">topic</a>, so I reduced the augmentation (other configure parts same with  <code>basic</code> approach).</p>\n<pre><code>Augmentation    random flip transpose,\n</code></pre>\n<p>The  random flip transpose:</p>\n<pre><code>def do_random_flip_transpose(image):\n    if np.random.rand() &lt; 0.5:\n        image = cv2.flip(image, 0)\n    if np.random.rand() &lt; 0.5:\n        image = cv2.flip(image, 1)\n    if np.random.rand() &lt; 0.5:\n        image = image.transpose(1, 0, 2)\n    return np.ascontiguousarray(image)\n</code></pre>\n<h3>Questions:</h3>\n<ol>\n<li>What am I doing wrong?</li>\n<li>For the same validation fold, for CV 0.898, I got LB between [0.883, 0.896]. Should I entirely pay attention to the LB score?</li>\n<li>What else can I try?</li>\n</ol>\n<p>I started this competition 16 days ago and spent nearly \"full-time\" daily without any considerable achievement. I ran out of ideas, so I need external advice.</p>\n<p>Thanks for the help!</p>",
      "rawMarkdown": "At this [post](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111), I read that multiple people get 0.900+ CV accuracy with their model, and far better LB than mine. I understand that the dataset is noisy, but others could get a much better score than I, with the effnet or resnext architecture.\n\nHere is the training cooperation table. I made the validation process from `k = 5` fold 0. I generated a .csv file, where I wrote what is fold 0, so the validation set is the same. I tested two architectures `resnet50_32x4d` and `effnet-b4`, and several methods, but somehow couldn't pass the 0.898 validation accuracy. I tried:\n1. [Bi-Tempered Logistic Loss](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017) loss\n2. [Focal Cosine Loss](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203271)\n3. [RandAugment](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208887#1139841)\n4. Up-sample the ``cbb', 'cbsd' ,'cgm', 'healthy'` classes to get the same distribution, or by 2x to reduce the class imbalance\n5. Combine the 2020 dataset with  2019 (merged column in the image)\n6. Snapmix\n7. Learning scheduler (Effnet-b4-scheduler column in the image)\nNeither one approach improved my LB or CV.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fe230cf84b270c5c0e19e3bede722e06d%2FSelection_085.jpg?generation=1611069573949012&alt=media)\n\nThe `basic` approach is. In the last case (column), I used a scheduler.\n\n```bash\nImage size      512\nOptimizer       Ranger (Lookahed + RAdam)\nDataset         2020\nFold            0/5\nAugmentation    random scale, random rotate, random contrast, random hsv\n                random flip transpose, do_random_cutout\nsmooth          0.05\n```\n\nI tried with light augmentation based on this [topic](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206489), so I reduced the augmentation (other configure parts same with  `basic` approach).\n\n```bash\nAugmentation\trandom flip transpose,\n```\n\nThe  random flip transpose:\n\n```python\ndef do_random_flip_transpose(image):\n    if np.random.rand() < 0.5:\n        image = cv2.flip(image, 0)\n    if np.random.rand() < 0.5:\n        image = cv2.flip(image, 1)\n    if np.random.rand() < 0.5:\n        image = image.transpose(1, 0, 2)\n    return np.ascontiguousarray(image)\n```\n\n### Questions:\n\n1. What am I doing wrong?\n2. For the same validation fold, for CV 0.898, I got LB between [0.883, 0.896]. Should I entirely pay attention to the LB score?\n3. What else can I try?\n\nI started this competition 16 days ago and spent nearly \"full-time\" daily without any considerable achievement. I ran out of ideas, so I need external advice.\n\nThanks for the help!",
      "votes": null
    },
    {
      "id": "1160162",
      "postDate": "01/19/2021 17:47:54",
      "content": "<ol>\n<li>I'm also stuck in the same dilemma. I too have spent so much time trying out different things from losses to augmentations to ensembles. Some of them worked, some of them didn't but I learned a lot from other discussions. </li>\n<li>I'm focusing on both Cv(mainly) and Lb because of the noisy labels.</li>\n<li>You could possibly try cutmix, mixup and combining 2019 dataset.</li>\n</ol>",
      "rawMarkdown": "1. I'm also stuck in the same dilemma. I too have spent so much time trying out different things from losses to augmentations to ensembles. Some of them worked, some of them didn't but I learned a lot from other discussions. \n2. I'm focusing on both Cv(mainly) and Lb because of the noisy labels.\n3. You could possibly try cutmix, mixup and combining 2019 dataset.",
      "votes": null
    },
    {
      "id": "1160205",
      "postDate": "01/19/2021 18:25:33",
      "content": "<p>Light Augs didn't help for me and made my model overfitting, thus poor CV and LB.  I use heavy augmentations. </p>\n<p>My CV/LB relation is pretty much stable .  The gap is between [- 0.002; +0.002]. But my CV is higher than 0.898 though.</p>\n<p>Adding 2019 dataset may help too.</p>\n<p>That being said,  beware that while doing stuffs that basically everyone else  knows/is doing, may help to climb on the LB, you'll also need some clever tricks and out of the box thinking to remain at the top of Private LB, like almost all kaggle competitions. </p>",
      "rawMarkdown": "Light Augs didn't help for me and made my model overfitting, thus poor CV and LB.  I use heavy augmentations. \n\nMy CV/LB relation is pretty much stable .  The gap is between [- 0.002; +0.002]. But my CV is higher than 0.898 though.\n\nAdding 2019 dataset may help too.\n\nThat being said,  beware that while doing stuffs that basically everyone else  knows/is doing, may help to climb on the LB, you'll also need some clever tricks and out of the box thinking to remain at the top of Private LB, like almost all kaggle competitions.",
      "votes": null
    },
    {
      "id": "1160368",
      "postDate": "01/19/2021 21:24:46",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/thakurudit\" target=\"_blank\">@thakurudit</a> for the response. I forgot to write that I tried snapmix, and to combine 2020 with the 2019 dataset, so I update the post</p>",
      "rawMarkdown": "thanks @thakurudit for the response. I forgot to write that I tried snapmix, and to combine 2020 with the 2019 dataset, so I update the post",
      "votes": null
    },
    {
      "id": "1160385",
      "postDate": "01/19/2021 21:39:32",
      "content": "<p>Thanks for your response! I forgot to write, but I tried to merge the 2020 and 2019 dataset as well. <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> do you use the simple k-fold method or something else for the cross-validation?<br>\nDo you have any thought why the CV and LB score is changing for the same validation set?</p>",
      "rawMarkdown": "Thanks for your response! I forgot to write, but I tried to merge the 2020 and 2019 dataset as well. @serigne do you use the simple k-fold method or something else for the cross-validation?\nDo you have any thought why the CV and LB score is changing for the same validation set?",
      "votes": null
    },
    {
      "id": "1160392",
      "postDate": "01/19/2021 21:47:26",
      "content": "<p>Start with B0ns(default) and adam,as a baseline, and then start testing around with common setups before using larger models. E.g. I got ValAcc 0.896/Lb 0.893 with singel fold B0ns. Start with heavy aug and reduce while testing instead of vice versa.</p>",
      "rawMarkdown": "Start with B0ns(default) and adam,as a baseline, and then start testing around with common setups before using larger models. E.g. I got ValAcc 0.896/Lb 0.893 with singel fold B0ns. Start with heavy aug and reduce while testing instead of vice versa.",
      "votes": null
    },
    {
      "id": "1160608",
      "postDate": "01/20/2021 03:28:26",
      "content": "<p>Did you use valid time augmentation for CV score?</p>",
      "rawMarkdown": "Did you use valid time augmentation for CV score?",
      "votes": null
    },
    {
      "id": "1160618",
      "postDate": "01/20/2021 03:41:18",
      "content": "<p>One question, <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Do you seed everything in your notebook or do you use a different seed to get different validation splits each time? If so, do you mind sharing your seed number?</p>",
      "rawMarkdown": "One question, @serigne Do you seed everything in your notebook or do you use a different seed to get different validation splits each time? If so, do you mind sharing your seed number?",
      "votes": null
    },
    {
      "id": "1160733",
      "postDate": "01/20/2021 05:10:32",
      "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> </p>\n<blockquote>\n  <p>Light Augs didn't help for me and made my model overfitting</p>\n</blockquote>\n<p>same thing for me, I don't how how people are achieving high Lb but unstable Cv relations with light augs.</p>",
      "rawMarkdown": "serigne \n> Light Augs didn't help for me and made my model overfitting\n\nsame thing for me, I don't how how people are achieving high Lb but unstable Cv relations with light augs.",
      "votes": null
    },
    {
      "id": "1161393",
      "postDate": "01/20/2021 14:44:22",
      "content": "<p><a href=\"https://www.kaggle.com/junyingsg\" target=\"_blank\">@junyingsg</a> <br>\nYeah I seed everything to make my experiments reproducible and see where improvements come from. </p>\n<p>I don't tweak the seed number and use just 42. </p>",
      "rawMarkdown": "junyingsg \nYeah I seed everything to make my experiments reproducible and see where improvements come from. \n\nI don't tweak the seed number and use just 42.",
      "votes": null
    },
    {
      "id": "1161480",
      "postDate": "01/20/2021 15:34:32",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "1164365",
      "postDate": "01/22/2021 10:23:29",
      "content": "<p>Another thing you can do is use 768x768 image sizes. I did that and it helped me jump 100 places on the lb with correlated cv.</p>",
      "rawMarkdown": "Another thing you can do is use 768x768 image sizes. I did that and it helped me jump 100 places on the lb with correlated cv.",
      "votes": null
    },
    {
      "id": "1170000",
      "postDate": "01/25/2021 21:59:37",
      "content": "<p>Try with other augmentations, other libs.<br>\nI know that people got +0.001/0.002+ with 2019, not sure why not significant but still some boost</p>",
      "rawMarkdown": "Try with other augmentations, other libs.\nI know that people got +0.001/0.002+ with 2019, not sure why not significant but still some boost",
      "votes": null
    },
    {
      "id": "1180597",
      "postDate": "02/01/2021 10:58:49",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> for your suggestion, I will try!</p>",
      "rawMarkdown": "Thanks @woshifym for your suggestion, I will try!",
      "votes": null
    },
    {
      "id": "1180632",
      "postDate": "02/01/2021 11:38:52",
      "content": "<p>Whats the point in doing this except pushing your cv?</p>",
      "rawMarkdown": "Whats the point in doing this except pushing your cv?",
      "votes": null
    },
    {
      "id": "1180690",
      "postDate": "02/01/2021 12:18:56",
      "content": "<p>The CV increase is the only benefit</p>",
      "rawMarkdown": "The CV increase is the only benefit",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1160162,
      "author_name": "thakurudit",
      "author_url": "",
      "post_date": "01/19/2021 17:47:54",
      "content": "<ol>\n<li>I'm also stuck in the same dilemma. I too have spent so much time trying out different things from losses to augmentations to ensembles. Some of them worked, some of them didn't but I learned a lot from other discussions. </li>\n<li>I'm focusing on both Cv(mainly) and Lb because of the noisy labels.</li>\n<li>You could possibly try cutmix, mixup and combining 2019 dataset.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1160368,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "01/19/2021 21:24:46",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/thakurudit\" target=\"_blank\">@thakurudit</a> for the response. I forgot to write that I tried snapmix, and to combine 2020 with the 2019 dataset, so I update the post</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1164365,
          "author_name": "thakurudit",
          "author_url": "",
          "post_date": "01/22/2021 10:23:29",
          "content": "<p>Another thing you can do is use 768x768 image sizes. I did that and it helped me jump 100 places on the lb with correlated cv.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1160205,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/19/2021 18:25:33",
      "content": "<p>Light Augs didn't help for me and made my model overfitting, thus poor CV and LB.  I use heavy augmentations. </p>\n<p>My CV/LB relation is pretty much stable .  The gap is between [- 0.002; +0.002]. But my CV is higher than 0.898 though.</p>\n<p>Adding 2019 dataset may help too.</p>\n<p>That being said,  beware that while doing stuffs that basically everyone else  knows/is doing, may help to climb on the LB, you'll also need some clever tricks and out of the box thinking to remain at the top of Private LB, like almost all kaggle competitions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1160385,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "01/19/2021 21:39:32",
          "content": "<p>Thanks for your response! I forgot to write, but I tried to merge the 2020 and 2019 dataset as well. <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> do you use the simple k-fold method or something else for the cross-validation?<br>\nDo you have any thought why the CV and LB score is changing for the same validation set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1160618,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/20/2021 03:41:18",
          "content": "<p>One question, <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Do you seed everything in your notebook or do you use a different seed to get different validation splits each time? If so, do you mind sharing your seed number?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1160733,
          "author_name": "thakurudit",
          "author_url": "",
          "post_date": "01/20/2021 05:10:32",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> </p>\n<blockquote>\n  <p>Light Augs didn't help for me and made my model overfitting</p>\n</blockquote>\n<p>same thing for me, I don't how how people are achieving high Lb but unstable Cv relations with light augs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1161393,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/20/2021 14:44:22",
          "content": "<p><a href=\"https://www.kaggle.com/junyingsg\" target=\"_blank\">@junyingsg</a> <br>\nYeah I seed everything to make my experiments reproducible and see where improvements come from. </p>\n<p>I don't tweak the seed number and use just 42. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1161480,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/20/2021 15:34:32",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1160392,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "01/19/2021 21:47:26",
      "content": "<p>Start with B0ns(default) and adam,as a baseline, and then start testing around with common setups before using larger models. E.g. I got ValAcc 0.896/Lb 0.893 with singel fold B0ns. Start with heavy aug and reduce while testing instead of vice versa.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1160608,
      "author_name": "woshifym",
      "author_url": "",
      "post_date": "01/20/2021 03:28:26",
      "content": "<p>Did you use valid time augmentation for CV score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1180597,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "02/01/2021 10:58:49",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> for your suggestion, I will try!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1180632,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "02/01/2021 11:38:52",
          "content": "<p>Whats the point in doing this except pushing your cv?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1180690,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "02/01/2021 12:18:56",
          "content": "<p>The CV increase is the only benefit</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1170000,
      "author_name": "muhakabartay",
      "author_url": "",
      "post_date": "01/25/2021 21:59:37",
      "content": "<p>Try with other augmentations, other libs.<br>\nI know that people got +0.001/0.002+ with 2019, not sure why not significant but still some boost</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1159968": "At this [post](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111), I read that multiple people get 0.900+ CV accuracy with their model, and far better LB than mine. I understand that the dataset is noisy, but others could get a much better score than I, with the effnet or resnext architecture.\n\nHere is the training cooperation table. I made the validation process from `k = 5` fold 0. I generated a .csv file, where I wrote what is fold 0, so the validation set is the same. I tested two architectures `resnet50_32x4d` and `effnet-b4`, and several methods, but somehow couldn't pass the 0.898 validation accuracy. I tried:\n1. [Bi-Tempered Logistic Loss](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017) loss\n2. [Focal Cosine Loss](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203271)\n3. [RandAugment](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208887#1139841)\n4. Up-sample the ``cbb', 'cbsd' ,'cgm', 'healthy'` classes to get the same distribution, or by 2x to reduce the class imbalance\n5. Combine the 2020 dataset with  2019 (merged column in the image)\n6. Snapmix\n7. Learning scheduler (Effnet-b4-scheduler column in the image)\nNeither one approach improved my LB or CV.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4367831%2Fe230cf84b270c5c0e19e3bede722e06d%2FSelection_085.jpg?generation=1611069573949012&alt=media)\n\nThe `basic` approach is. In the last case (column), I used a scheduler.\n\n```bash\nImage size      512\nOptimizer       Ranger (Lookahed + RAdam)\nDataset         2020\nFold            0/5\nAugmentation    random scale, random rotate, random contrast, random hsv\n                random flip transpose, do_random_cutout\nsmooth          0.05\n```\n\nI tried with light augmentation based on this [topic](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206489), so I reduced the augmentation (other configure parts same with  `basic` approach).\n\n```bash\nAugmentation\trandom flip transpose,\n```\n\nThe  random flip transpose:\n\n```python\ndef do_random_flip_transpose(image):\n    if np.random.rand() < 0.5:\n        image = cv2.flip(image, 0)\n    if np.random.rand() < 0.5:\n        image = cv2.flip(image, 1)\n    if np.random.rand() < 0.5:\n        image = image.transpose(1, 0, 2)\n    return np.ascontiguousarray(image)\n```\n\n### Questions:\n\n1. What am I doing wrong?\n2. For the same validation fold, for CV 0.898, I got LB between [0.883, 0.896]. Should I entirely pay attention to the LB score?\n3. What else can I try?\n\nI started this competition 16 days ago and spent nearly \"full-time\" daily without any considerable achievement. I ran out of ideas, so I need external advice.\n\nThanks for the help!",
    "1160162": "1. I'm also stuck in the same dilemma. I too have spent so much time trying out different things from losses to augmentations to ensembles. Some of them worked, some of them didn't but I learned a lot from other discussions. \n2. I'm focusing on both Cv(mainly) and Lb because of the noisy labels.\n3. You could possibly try cutmix, mixup and combining 2019 dataset.",
    "1160205": "Light Augs didn't help for me and made my model overfitting, thus poor CV and LB.  I use heavy augmentations. \n\nMy CV/LB relation is pretty much stable .  The gap is between [- 0.002; +0.002]. But my CV is higher than 0.898 though.\n\nAdding 2019 dataset may help too.\n\nThat being said,  beware that while doing stuffs that basically everyone else  knows/is doing, may help to climb on the LB, you'll also need some clever tricks and out of the box thinking to remain at the top of Private LB, like almost all kaggle competitions.",
    "1160368": "thanks @thakurudit for the response. I forgot to write that I tried snapmix, and to combine 2020 with the 2019 dataset, so I update the post",
    "1160385": "Thanks for your response! I forgot to write, but I tried to merge the 2020 and 2019 dataset as well. @serigne do you use the simple k-fold method or something else for the cross-validation?\nDo you have any thought why the CV and LB score is changing for the same validation set?",
    "1160392": "Start with B0ns(default) and adam,as a baseline, and then start testing around with common setups before using larger models. E.g. I got ValAcc 0.896/Lb 0.893 with singel fold B0ns. Start with heavy aug and reduce while testing instead of vice versa.",
    "1160608": "Did you use valid time augmentation for CV score?",
    "1160618": "One question, @serigne Do you seed everything in your notebook or do you use a different seed to get different validation splits each time? If so, do you mind sharing your seed number?",
    "1160733": "serigne \n> Light Augs didn't help for me and made my model overfitting\n\nsame thing for me, I don't how how people are achieving high Lb but unstable Cv relations with light augs.",
    "1161393": "junyingsg \nYeah I seed everything to make my experiments reproducible and see where improvements come from. \n\nI don't tweak the seed number and use just 42.",
    "1161480": "Thanks a lot!",
    "1164365": "Another thing you can do is use 768x768 image sizes. I did that and it helped me jump 100 places on the lb with correlated cv.",
    "1170000": "Try with other augmentations, other libs.\nI know that people got +0.001/0.002+ with 2019, not sure why not significant but still some boost",
    "1180597": "Thanks @woshifym for your suggestion, I will try!",
    "1180632": "Whats the point in doing this except pushing your cv?",
    "1180690": "The CV increase is the only benefit"
  },
  "source": "meta"
}