{
  "id": 134844,
  "title": "Anyone managed to make your CV go above 0.99 with only cutout or no aug at all?",
  "url": "/competitions/bengaliai-cv19/discussion/134844",
  "author_name": "",
  "post_date": "2020-03-10T17:58:37.696938900Z",
  "votes": 4,
  "comment_count": 20,
  "views": 0,
  "content": "<p>My top CVs are around 0.989 and I can't seem to make it go above 0.99. I'm using only cutout as my major augmentation and some other basic augmentations like rotation, zoom, shift, etc. I haven't tried mixup, cutmix, and other augs mentioned on this forum. I don't know whether my problem is due to lack of model tuning or lack of augmentation. So I'm here to ask you guys about your experiences. Even a little bit of hint would be an immense help. Thanks in advance!</p>",
  "messages": [
    {
      "id": "768379",
      "postDate": "03/10/2020 17:58:37",
      "content": "<p>My top CVs are around 0.989 and I can't seem to make it go above 0.99. I'm using only cutout as my major augmentation and some other basic augmentations like rotation, zoom, shift, etc. I haven't tried mixup, cutmix, and other augs mentioned on this forum. I don't know whether my problem is due to lack of model tuning or lack of augmentation. So I'm here to ask you guys about your experiences. Even a little bit of hint would be an immense help. Thanks in advance!</p>",
      "rawMarkdown": "My top CVs are around 0.989 and I can't seem to make it go above 0.99. I'm using only cutout as my major augmentation and some other basic augmentations like rotation, zoom, shift, etc. I haven't tried mixup, cutmix, and other augs mentioned on this forum. I don't know whether my problem is due to lack of model tuning or lack of augmentation. So I'm here to ask you guys about your experiences. Even a little bit of hint would be an immense help. Thanks in advance!",
      "votes": null
    },
    {
      "id": "768397",
      "postDate": "03/10/2020 18:35:17",
      "content": "<p>Yes, all of my models have been trained with cutout only (no other augmentation) and all of them achieved around 99.4-99.6 CV so far. </p>",
      "rawMarkdown": "Yes, all of my models have been trained with cutout only (no other augmentation) and all of them achieved around 99.4-99.6 CV so far.",
      "votes": null
    },
    {
      "id": "768399",
      "postDate": "03/10/2020 18:42:03",
      "content": "<p>Thanks heaps!</p>",
      "rawMarkdown": "Thanks heaps!",
      "votes": null
    },
    {
      "id": "771789",
      "postDate": "03/14/2020 16:26:37",
      "content": "<p>that's amazing... I have been trying with various model and cutmix/mixup/gridmask etc... barely get 0.97 on local single model... not sure what's wrong. Seems people are saying that it is easy to get 0.98.</p>",
      "rawMarkdown": "that's amazing... I have been trying with various model and cutmix/mixup/gridmask etc... barely get 0.97 on local single model... not sure what's wrong. Seems people are saying that it is easy to get 0.98.",
      "votes": null
    },
    {
      "id": "772127",
      "postDate": "03/15/2020 04:03:41",
      "content": "<p>My model is about 99.0 without any augmentation :)</p>",
      "rawMarkdown": "My model is about 99.0 without any augmentation :)",
      "votes": null
    },
    {
      "id": "772140",
      "postDate": "03/15/2020 04:55:00",
      "content": "<p><a href=\"/ryunosukeishizaki\">@ryunosukeishizaki</a> wouldn't this cause the model to overfit without augmentation?</p>",
      "rawMarkdown": "ryunosukeishizaki wouldn't this cause the model to overfit without augmentation?",
      "votes": null
    },
    {
      "id": "772144",
      "postDate": "03/15/2020 05:00:48",
      "content": "<p>I'm not sure,\nI just started this competition 2days ago then I didn't test, but\nnow I'm having CV 0.994 with augmentation, not submitting yet Lol</p>",
      "rawMarkdown": "I'm not sure,\nI just started this competition 2days ago then I didn't test, but\nnow I'm having CV 0.994 with augmentation, not submitting yet Lol",
      "votes": null
    },
    {
      "id": "772260",
      "postDate": "03/15/2020 08:47:32",
      "content": "<p>I believe you are using 3 labels for the training. Using the last column in train.csv as a supplyment is going to help a lot.</p>",
      "rawMarkdown": "I believe you are using 3 labels for the training. Using the last column in train.csv as a supplyment is going to help a lot.",
      "votes": null
    },
    {
      "id": "772524",
      "postDate": "03/15/2020 15:34:11",
      "content": "<p>So what magic are you using in your modelling? Hoping to hear about that after the competition ends</p>",
      "rawMarkdown": "So what magic are you using in your modelling? Hoping to hear about that after the competition ends",
      "votes": null
    },
    {
      "id": "772532",
      "postDate": "03/15/2020 15:40:00",
      "content": "<p>Indeed... so u train with 4 label instead?  I almost forgot there is a last column. </p>",
      "rawMarkdown": "Indeed... so u train with 4 label instead?  I almost forgot there is a last column.",
      "votes": null
    },
    {
      "id": "772577",
      "postDate": "03/15/2020 16:34:58",
      "content": "<p>yep, it will easily get 0.98 with 4 labels</p>",
      "rawMarkdown": "yep, it will easily get 0.98 with 4 labels",
      "votes": null
    },
    {
      "id": "772648",
      "postDate": "03/15/2020 18:19:27",
      "content": "<p>omg... I am getting this so late.........no way I can retrain my 5-fold to get this right.. but thanks! I will start training.</p>",
      "rawMarkdown": "omg... I am getting this so late.........no way I can retrain my 5-fold to get this right.. but thanks! I will start training.",
      "votes": null
    },
    {
      "id": "772722",
      "postDate": "03/15/2020 20:20:13",
      "content": "<p>the last col is a string though right? curious how you guys are dealing with that. thinking about label encoding or taking the sum of the ord of the chars in the string but not sure if thats a good solution</p>",
      "rawMarkdown": "the last col is a string though right? curious how you guys are dealing with that. thinking about label encoding or taking the sum of the ord of the chars in the string but not sure if thats a good solution",
      "votes": null
    },
    {
      "id": "775187",
      "postDate": "03/16/2020 11:47:44",
      "content": "<p>2 days CV0.994... it is very impressive..</p>",
      "rawMarkdown": "2 days CV0.994... it is very impressive..",
      "votes": null
    },
    {
      "id": "775193",
      "postDate": "03/16/2020 11:51:19",
      "content": "<p>As you can see, top teams are sharing solutions so openly :)</p>",
      "rawMarkdown": "As you can see, top teams are sharing solutions so openly :)",
      "votes": null
    },
    {
      "id": "775227",
      "postDate": "03/16/2020 12:36:43",
      "content": "<p>\"easy to get 0.98.\" </p>\n\n<p>it is already mentioned in the forum and shown in the public kernel that just predicting the 4th column and use it decode into \"root,vowel,const = decode(grapheme)\" will get slightly better results (+0.001 to 0.002).</p>\n\n<p>almost all tricks already revealed in the discussion forum. that should get you at least 0.985</p>",
      "rawMarkdown": "\"easy to get 0.98.\" \n\nit is already mentioned in the forum and shown in the public kernel that just predicting the 4th column and use it decode into \"root,vowel,const = decode(grapheme)\" will get slightly better results (+0.001 to 0.002).\n\nalmost all tricks already revealed in the discussion forum. that should get you at least 0.985",
      "votes": null
    },
    {
      "id": "775349",
      "postDate": "03/16/2020 15:43:59",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  did you get large gap between cv/lb? we are facing large gap between cv/lb, like for cv 0.9895 lb 0.9811\nshould we still trust our cv for private lb? we had 80-20 split for efficientnet-b3 training</p>",
      "rawMarkdown": "hengck23  did you get large gap between cv/lb? we are facing large gap between cv/lb, like for cv 0.9895 lb 0.9811\nshould we still trust our cv for private lb? we had 80-20 split for efficientnet-b3 training",
      "votes": null
    },
    {
      "id": "775412",
      "postDate": "03/16/2020 16:44:57",
      "content": "<p>I guess I have to wait until the end of the competition, I try using a lot of things mentioned in discussion... so far I couldn't get beyond 0.98, even with the suggestion here saying that I should use the grapheme column for training too.</p>",
      "rawMarkdown": "I guess I have to wait until the end of the competition, I try using a lot of things mentioned in discussion... so far I couldn't get beyond 0.98, even with the suggestion here saying that I should use the grapheme column for training too.",
      "votes": null
    },
    {
      "id": "775513",
      "postDate": "03/16/2020 18:56:41",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> this is a very reasonable gap actually, considering the unseen grapheme cases in the test dataset. In the 'Best single model' discussion thread, you will find a similar gap even for the ~99 scoring models.</p>",
      "rawMarkdown": "mobassir this is a very reasonable gap actually, considering the unseen grapheme cases in the test dataset. In the 'Best single model' discussion thread, you will find a similar gap even for the ~99 scoring models.",
      "votes": null
    },
    {
      "id": "775524",
      "postDate": "03/16/2020 19:10:23",
      "content": "<p><a href=\"/udaykamal\">@udaykamal</a>  it's very hard to choose model in such situation,we have witnessed situation like this : \nmodel 1 : cv 0.9895 and lb 0.9811\nmodel 2 : exact same setting as model 1(just augmentaiton change,like from cutmix to cutout only) and result is CV 0.9893 and lb 0.9827 :(\nhard for me to decide </p>",
      "rawMarkdown": "udaykamal  it's very hard to choose model in such situation,we have witnessed situation like this : \nmodel 1 : cv 0.9895 and lb 0.9811\nmodel 2 : exact same setting as model 1(just augmentaiton change,like from cutmix to cutout only) and result is CV 0.9893 and lb 0.9827 :(\nhard for me to decide",
      "votes": null
    },
    {
      "id": "776409",
      "postDate": "03/17/2020 11:20:37",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks, I really missed this part in the forum discussion. I guess I am somehow lucky for not seeing this as it seems the overfitting problem is very serious and the private LB shocked a lot. I did try add grapheme  just for testing, it does not seem to help much for my local CV</p>",
      "rawMarkdown": "hengck23 Thanks, I really missed this part in the forum discussion. I guess I am somehow lucky for not seeing this as it seems the overfitting problem is very serious and the private LB shocked a lot. I did try add grapheme  just for testing, it does not seem to help much for my local CV",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768397,
      "author_name": "udaykamal",
      "author_url": "",
      "post_date": "03/10/2020 18:35:17",
      "content": "<p>Yes, all of my models have been trained with cutout only (no other augmentation) and all of them achieved around 99.4-99.6 CV so far. </p>",
      "votes": null,
      "replies": [
        {
          "id": 768399,
          "author_name": "max6296",
          "author_url": "",
          "post_date": "03/10/2020 18:42:03",
          "content": "<p>Thanks heaps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 772524,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "03/15/2020 15:34:11",
          "content": "<p>So what magic are you using in your modelling? Hoping to hear about that after the competition ends</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 771789,
      "author_name": "noklamchan",
      "author_url": "",
      "post_date": "03/14/2020 16:26:37",
      "content": "<p>that's amazing... I have been trying with various model and cutmix/mixup/gridmask etc... barely get 0.97 on local single model... not sure what's wrong. Seems people are saying that it is easy to get 0.98.</p>",
      "votes": null,
      "replies": [
        {
          "id": 772260,
          "author_name": "markson14",
          "author_url": "",
          "post_date": "03/15/2020 08:47:32",
          "content": "<p>I believe you are using 3 labels for the training. Using the last column in train.csv as a supplyment is going to help a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 772532,
          "author_name": "noklamchan",
          "author_url": "",
          "post_date": "03/15/2020 15:40:00",
          "content": "<p>Indeed... so u train with 4 label instead?  I almost forgot there is a last column. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 772577,
          "author_name": "markson14",
          "author_url": "",
          "post_date": "03/15/2020 16:34:58",
          "content": "<p>yep, it will easily get 0.98 with 4 labels</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 772648,
          "author_name": "noklamchan",
          "author_url": "",
          "post_date": "03/15/2020 18:19:27",
          "content": "<p>omg... I am getting this so late.........no way I can retrain my 5-fold to get this right.. but thanks! I will start training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 772722,
          "author_name": "ajkaggle1",
          "author_url": "",
          "post_date": "03/15/2020 20:20:13",
          "content": "<p>the last col is a string though right? curious how you guys are dealing with that. thinking about label encoding or taking the sum of the ord of the chars in the string but not sure if thats a good solution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775227,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/16/2020 12:36:43",
          "content": "<p>\"easy to get 0.98.\" </p>\n\n<p>it is already mentioned in the forum and shown in the public kernel that just predicting the 4th column and use it decode into \"root,vowel,const = decode(grapheme)\" will get slightly better results (+0.001 to 0.002).</p>\n\n<p>almost all tricks already revealed in the discussion forum. that should get you at least 0.985</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775349,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "03/16/2020 15:43:59",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  did you get large gap between cv/lb? we are facing large gap between cv/lb, like for cv 0.9895 lb 0.9811\nshould we still trust our cv for private lb? we had 80-20 split for efficientnet-b3 training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775513,
          "author_name": "udaykamal",
          "author_url": "",
          "post_date": "03/16/2020 18:56:41",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> this is a very reasonable gap actually, considering the unseen grapheme cases in the test dataset. In the 'Best single model' discussion thread, you will find a similar gap even for the ~99 scoring models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775524,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "03/16/2020 19:10:23",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a>  it's very hard to choose model in such situation,we have witnessed situation like this : \nmodel 1 : cv 0.9895 and lb 0.9811\nmodel 2 : exact same setting as model 1(just augmentaiton change,like from cutmix to cutout only) and result is CV 0.9893 and lb 0.9827 :(\nhard for me to decide </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776409,
          "author_name": "noklamchan",
          "author_url": "",
          "post_date": "03/17/2020 11:20:37",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks, I really missed this part in the forum discussion. I guess I am somehow lucky for not seeing this as it seems the overfitting problem is very serious and the private LB shocked a lot. I did try add grapheme  just for testing, it does not seem to help much for my local CV</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 772127,
      "author_name": "ryunosukeishizaki",
      "author_url": "",
      "post_date": "03/15/2020 04:03:41",
      "content": "<p>My model is about 99.0 without any augmentation :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 772140,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "03/15/2020 04:55:00",
          "content": "<p><a href=\"/ryunosukeishizaki\">@ryunosukeishizaki</a> wouldn't this cause the model to overfit without augmentation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 772144,
          "author_name": "ryunosukeishizaki",
          "author_url": "",
          "post_date": "03/15/2020 05:00:48",
          "content": "<p>I'm not sure,\nI just started this competition 2days ago then I didn't test, but\nnow I'm having CV 0.994 with augmentation, not submitting yet Lol</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775187,
          "author_name": "noklamchan",
          "author_url": "",
          "post_date": "03/16/2020 11:47:44",
          "content": "<p>2 days CV0.994... it is very impressive..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775193,
          "author_name": "ryunosukeishizaki",
          "author_url": "",
          "post_date": "03/16/2020 11:51:19",
          "content": "<p>As you can see, top teams are sharing solutions so openly :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775412,
          "author_name": "noklamchan",
          "author_url": "",
          "post_date": "03/16/2020 16:44:57",
          "content": "<p>I guess I have to wait until the end of the competition, I try using a lot of things mentioned in discussion... so far I couldn't get beyond 0.98, even with the suggestion here saying that I should use the grapheme column for training too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "768379": "My top CVs are around 0.989 and I can't seem to make it go above 0.99. I'm using only cutout as my major augmentation and some other basic augmentations like rotation, zoom, shift, etc. I haven't tried mixup, cutmix, and other augs mentioned on this forum. I don't know whether my problem is due to lack of model tuning or lack of augmentation. So I'm here to ask you guys about your experiences. Even a little bit of hint would be an immense help. Thanks in advance!",
    "768397": "Yes, all of my models have been trained with cutout only (no other augmentation) and all of them achieved around 99.4-99.6 CV so far.",
    "768399": "Thanks heaps!",
    "771789": "that's amazing... I have been trying with various model and cutmix/mixup/gridmask etc... barely get 0.97 on local single model... not sure what's wrong. Seems people are saying that it is easy to get 0.98.",
    "772127": "My model is about 99.0 without any augmentation :)",
    "772140": "ryunosukeishizaki wouldn't this cause the model to overfit without augmentation?",
    "772144": "I'm not sure,\nI just started this competition 2days ago then I didn't test, but\nnow I'm having CV 0.994 with augmentation, not submitting yet Lol",
    "772260": "I believe you are using 3 labels for the training. Using the last column in train.csv as a supplyment is going to help a lot.",
    "772524": "So what magic are you using in your modelling? Hoping to hear about that after the competition ends",
    "772532": "Indeed... so u train with 4 label instead?  I almost forgot there is a last column.",
    "772577": "yep, it will easily get 0.98 with 4 labels",
    "772648": "omg... I am getting this so late.........no way I can retrain my 5-fold to get this right.. but thanks! I will start training.",
    "772722": "the last col is a string though right? curious how you guys are dealing with that. thinking about label encoding or taking the sum of the ord of the chars in the string but not sure if thats a good solution",
    "775187": "2 days CV0.994... it is very impressive..",
    "775193": "As you can see, top teams are sharing solutions so openly :)",
    "775227": "\"easy to get 0.98.\" \n\nit is already mentioned in the forum and shown in the public kernel that just predicting the 4th column and use it decode into \"root,vowel,const = decode(grapheme)\" will get slightly better results (+0.001 to 0.002).\n\nalmost all tricks already revealed in the discussion forum. that should get you at least 0.985",
    "775349": "hengck23  did you get large gap between cv/lb? we are facing large gap between cv/lb, like for cv 0.9895 lb 0.9811\nshould we still trust our cv for private lb? we had 80-20 split for efficientnet-b3 training",
    "775412": "I guess I have to wait until the end of the competition, I try using a lot of things mentioned in discussion... so far I couldn't get beyond 0.98, even with the suggestion here saying that I should use the grapheme column for training too.",
    "775513": "mobassir this is a very reasonable gap actually, considering the unseen grapheme cases in the test dataset. In the 'Best single model' discussion thread, you will find a similar gap even for the ~99 scoring models.",
    "775524": "udaykamal  it's very hard to choose model in such situation,we have witnessed situation like this : \nmodel 1 : cv 0.9895 and lb 0.9811\nmodel 2 : exact same setting as model 1(just augmentaiton change,like from cutmix to cutout only) and result is CV 0.9893 and lb 0.9827 :(\nhard for me to decide",
    "776409": "hengck23 Thanks, I really missed this part in the forum discussion. I guess I am somehow lucky for not seeing this as it seems the overfitting problem is very serious and the private LB shocked a lot. I did try add grapheme  just for testing, it does not seem to help much for my local CV"
  },
  "source": "meta"
}