{
  "id": 321312,
  "title": "Did you choose your best submission ?",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/321312",
  "author_name": "",
  "post_date": "2022-04-26T07:58:33.583514Z",
  "votes": 12,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone, I'd like to thank <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for hosting this competition and congratulate all the winners <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> <a href=\"https://www.kaggle.com/dmitrykonovalov\" target=\"_blank\">@dmitrykonovalov</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>. </p>\n<p>Today, I woke up and i realized that i didn't select my best submissions. Actually i missed 10 of them following CV (should have gone second with 0.53595). I'd like to know if some you experienced the same.</p>",
  "messages": [
    {
      "id": "1768367",
      "postDate": "04/26/2022 07:58:33",
      "content": "<p>Hi everyone, I'd like to thank <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for hosting this competition and congratulate all the winners <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> <a href=\"https://www.kaggle.com/dmitrykonovalov\" target=\"_blank\">@dmitrykonovalov</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>. </p>\n<p>Today, I woke up and i realized that i didn't select my best submissions. Actually i missed 10 of them following CV (should have gone second with 0.53595). I'd like to know if some you experienced the same.</p>",
      "rawMarkdown": "Hi everyone, I'd like to thank @robikscube for hosting this competition and congratulate all the winners @dienhoa @dmitrykonovalov @pheadrus. \n\nToday, I woke up and i realized that i didn't select my best submissions. Actually i missed 10 of them following CV (should have gone second with 0.53595). I'd like to know if some you experienced the same.",
      "votes": null
    },
    {
      "id": "1768411",
      "postDate": "04/26/2022 09:11:24",
      "content": "<p>Nope, I chose like the 15th best on private, trusted CV. :( </p>\n<p>Had 54.xx subs but ignored those since CV was low. </p>",
      "rawMarkdown": "Nope, I chose like the 15th best on private, trusted CV. :( \n\nHad 54.xx subs but ignored those since CV was low.",
      "votes": null
    },
    {
      "id": "1768423",
      "postDate": "04/26/2022 09:24:36",
      "content": "<p>Oops !!! So i'm not alone, i chose the 11th best on private</p>",
      "rawMarkdown": "Oops !!! So i'm not alone, i chose the 11th best on private",
      "votes": null
    },
    {
      "id": "1768455",
      "postDate": "04/26/2022 09:53:33",
      "content": "<p>lol. I am sure there are many others like us in this comp! :) </p>",
      "rawMarkdown": "lol. I am sure there are many others like us in this comp! :)",
      "votes": null
    },
    {
      "id": "1768680",
      "postDate": "04/26/2022 14:53:06",
      "content": "<p>I had a couple better that were trained with mixup, but mixup wasn't showing any benefit to my CV. I would have stayed in 8th place, so not that big of a deal for me. Guessing there were more of the rare music genres in the hidden private dataset.</p>",
      "rawMarkdown": "I had a couple better that were trained with mixup, but mixup wasn't showing any benefit to my CV. I would have stayed in 8th place, so not that big of a deal for me. Guessing there were more of the rare music genres in the hidden private dataset.",
      "votes": null
    },
    {
      "id": "1768910",
      "postDate": "04/26/2022 18:37:14",
      "content": "<p>Can I ask what is your best CV <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>? Mine is ~ 0.615 </p>",
      "rawMarkdown": "Can I ask what is your best CV @ulrich07 @pheadrus? Mine is ~ 0.615",
      "votes": null
    },
    {
      "id": "1768921",
      "postDate": "04/26/2022 18:41:56",
      "content": "<p>Without pseudo labels CV=0.60x, with Pseudo labels CV=0.62x</p>",
      "rawMarkdown": "Without pseudo labels CV=0.60x, with Pseudo labels CV=0.62x",
      "votes": null
    },
    {
      "id": "1768943",
      "postDate": "04/26/2022 19:01:45",
      "content": "<p>Thanks. I didn’t try pseudo labeling, I thought with accuracy ~ 0.6, the pseudo labels are not very reliable, isn’t it?</p>",
      "rawMarkdown": "Thanks. I didn’t try pseudo labeling, I thought with accuracy ~ 0.6, the pseudo labels are not very reliable, isn’t it?",
      "votes": null
    },
    {
      "id": "1769298",
      "postDate": "04/27/2022 06:16:05",
      "content": "<p>Ensemble CV 62.5. Single models best CV was 59.</p>",
      "rawMarkdown": "Ensemble CV 62.5. Single models best CV was 59.",
      "votes": null
    },
    {
      "id": "1769339",
      "postDate": "04/27/2022 06:59:06",
      "content": "<p>In my case pseudo labels works good. It both increases my LB and CV</p>",
      "rawMarkdown": "In my case pseudo labels works good. It both increases my LB and CV",
      "votes": null
    },
    {
      "id": "1769437",
      "postDate": "04/27/2022 09:05:44",
      "content": "<p>How did you select these PLs? </p>",
      "rawMarkdown": "How did you select these PLs?",
      "votes": null
    },
    {
      "id": "1769478",
      "postDate": "04/27/2022 09:59:50",
      "content": "<p>I chose  my best LB model (fastai xresnet50) given that all my models had  approximately same Cvs. This model scored 0.579x.</p>",
      "rawMarkdown": "I chose  my best LB model (fastai xresnet50) given that all my models had  approximately same Cvs. This model scored 0.579x.",
      "votes": null
    },
    {
      "id": "1769594",
      "postDate": "04/27/2022 12:28:30",
      "content": "<p>My best Private LB is from a single model xse_resnext50 with 7 folds, the avg CV ~ 0.606</p>\n<p>After reflection, I think because of lacking experience in Kaggle, I struggled to improve my score on the Public LB. I didn't well prepare the split, then can not calculate the Ensemble CV for multiple models -&gt; can't evaluate the Ensemble</p>\n<p>Didn't try pseudo labeling too because thinking the accuracy is not very reliable. But as you said <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> , I can try anyway to see if the CV or Public LB improves or not. If the Private Dataset is from the same distribution of Public Dataset and Traning Set, I can miss a lot of points here :D</p>\n<p>Luckily,  the Private DataSet is not from the same distribution of Training Set and Public LB Dataset. My lack of experience helps me not overfit the Private LB. </p>\n<p>Anyway, thank you all for participating in this competition <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>  <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> I've learned a lot</p>",
      "rawMarkdown": "My best Private LB is from a single model xse_resnext50 with 7 folds, the avg CV ~ 0.606\n\nAfter reflection, I think because of lacking experience in Kaggle, I struggled to improve my score on the Public LB. I didn't well prepare the split, then can not calculate the Ensemble CV for multiple models -> can't evaluate the Ensemble\n\nDidn't try pseudo labeling too because thinking the accuracy is not very reliable. But as you said @ulrich07 , I can try anyway to see if the CV or Public LB improves or not. If the Private Dataset is from the same distribution of Public Dataset and Traning Set, I can miss a lot of points here :D\n\nLuckily,  the Private DataSet is not from the same distribution of Training Set and Public LB Dataset. My lack of experience helps me not overfit the Private LB. \n\nAnyway, thank you all for participating in this competition @pheadrus  @ulrich07 I've learned a lot",
      "votes": null
    },
    {
      "id": "1769935",
      "postDate": "04/27/2022 19:01:18",
      "content": "<p>But does it also mean that picking the best CV from Ensemble can lead to overfitting the CV too, Same as Early Stopping? The best way to double-check is the improvement of the Public LB. If both of them still lead to a bad Private LB then it is just because of bad luck :D</p>",
      "rawMarkdown": "But does it also mean that picking the best CV from Ensemble can lead to overfitting the CV too, Same as Early Stopping? The best way to double-check is the improvement of the Public LB. If both of them still lead to a bad Private LB then it is just because of bad luck :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1768411,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "04/26/2022 09:11:24",
      "content": "<p>Nope, I chose like the 15th best on private, trusted CV. :( </p>\n<p>Had 54.xx subs but ignored those since CV was low. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1768423,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "04/26/2022 09:24:36",
          "content": "<p>Oops !!! So i'm not alone, i chose the 11th best on private</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1768455,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/26/2022 09:53:33",
          "content": "<p>lol. I am sure there are many others like us in this comp! :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1768680,
      "author_name": "pcmiller",
      "author_url": "",
      "post_date": "04/26/2022 14:53:06",
      "content": "<p>I had a couple better that were trained with mixup, but mixup wasn't showing any benefit to my CV. I would have stayed in 8th place, so not that big of a deal for me. Guessing there were more of the rare music genres in the hidden private dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1768910,
      "author_name": "dienhoa",
      "author_url": "",
      "post_date": "04/26/2022 18:37:14",
      "content": "<p>Can I ask what is your best CV <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>? Mine is ~ 0.615 </p>",
      "votes": null,
      "replies": [
        {
          "id": 1768921,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "04/26/2022 18:41:56",
          "content": "<p>Without pseudo labels CV=0.60x, with Pseudo labels CV=0.62x</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1768943,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/26/2022 19:01:45",
          "content": "<p>Thanks. I didn’t try pseudo labeling, I thought with accuracy ~ 0.6, the pseudo labels are not very reliable, isn’t it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1769298,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/27/2022 06:16:05",
          "content": "<p>Ensemble CV 62.5. Single models best CV was 59.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1769339,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "04/27/2022 06:59:06",
          "content": "<p>In my case pseudo labels works good. It both increases my LB and CV</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1769437,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/27/2022 09:05:44",
          "content": "<p>How did you select these PLs? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1769478,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "04/27/2022 09:59:50",
          "content": "<p>I chose  my best LB model (fastai xresnet50) given that all my models had  approximately same Cvs. This model scored 0.579x.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1769594,
      "author_name": "dienhoa",
      "author_url": "",
      "post_date": "04/27/2022 12:28:30",
      "content": "<p>My best Private LB is from a single model xse_resnext50 with 7 folds, the avg CV ~ 0.606</p>\n<p>After reflection, I think because of lacking experience in Kaggle, I struggled to improve my score on the Public LB. I didn't well prepare the split, then can not calculate the Ensemble CV for multiple models -&gt; can't evaluate the Ensemble</p>\n<p>Didn't try pseudo labeling too because thinking the accuracy is not very reliable. But as you said <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> , I can try anyway to see if the CV or Public LB improves or not. If the Private Dataset is from the same distribution of Public Dataset and Traning Set, I can miss a lot of points here :D</p>\n<p>Luckily,  the Private DataSet is not from the same distribution of Training Set and Public LB Dataset. My lack of experience helps me not overfit the Private LB. </p>\n<p>Anyway, thank you all for participating in this competition <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>  <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> I've learned a lot</p>",
      "votes": null,
      "replies": [
        {
          "id": 1769935,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/27/2022 19:01:18",
          "content": "<p>But does it also mean that picking the best CV from Ensemble can lead to overfitting the CV too, Same as Early Stopping? The best way to double-check is the improvement of the Public LB. If both of them still lead to a bad Private LB then it is just because of bad luck :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1768367": "Hi everyone, I'd like to thank @robikscube for hosting this competition and congratulate all the winners @dienhoa @dmitrykonovalov @pheadrus. \n\nToday, I woke up and i realized that i didn't select my best submissions. Actually i missed 10 of them following CV (should have gone second with 0.53595). I'd like to know if some you experienced the same.",
    "1768411": "Nope, I chose like the 15th best on private, trusted CV. :( \n\nHad 54.xx subs but ignored those since CV was low.",
    "1768423": "Oops !!! So i'm not alone, i chose the 11th best on private",
    "1768455": "lol. I am sure there are many others like us in this comp! :)",
    "1768680": "I had a couple better that were trained with mixup, but mixup wasn't showing any benefit to my CV. I would have stayed in 8th place, so not that big of a deal for me. Guessing there were more of the rare music genres in the hidden private dataset.",
    "1768910": "Can I ask what is your best CV @ulrich07 @pheadrus? Mine is ~ 0.615",
    "1768921": "Without pseudo labels CV=0.60x, with Pseudo labels CV=0.62x",
    "1768943": "Thanks. I didn’t try pseudo labeling, I thought with accuracy ~ 0.6, the pseudo labels are not very reliable, isn’t it?",
    "1769298": "Ensemble CV 62.5. Single models best CV was 59.",
    "1769339": "In my case pseudo labels works good. It both increases my LB and CV",
    "1769437": "How did you select these PLs?",
    "1769478": "I chose  my best LB model (fastai xresnet50) given that all my models had  approximately same Cvs. This model scored 0.579x.",
    "1769594": "My best Private LB is from a single model xse_resnext50 with 7 folds, the avg CV ~ 0.606\n\nAfter reflection, I think because of lacking experience in Kaggle, I struggled to improve my score on the Public LB. I didn't well prepare the split, then can not calculate the Ensemble CV for multiple models -> can't evaluate the Ensemble\n\nDidn't try pseudo labeling too because thinking the accuracy is not very reliable. But as you said @ulrich07 , I can try anyway to see if the CV or Public LB improves or not. If the Private Dataset is from the same distribution of Public Dataset and Traning Set, I can miss a lot of points here :D\n\nLuckily,  the Private DataSet is not from the same distribution of Training Set and Public LB Dataset. My lack of experience helps me not overfit the Private LB. \n\nAnyway, thank you all for participating in this competition @pheadrus  @ulrich07 I've learned a lot",
    "1769935": "But does it also mean that picking the best CV from Ensemble can lead to overfitting the CV too, Same as Early Stopping? The best way to double-check is the improvement of the Public LB. If both of them still lead to a bad Private LB then it is just because of bad luck :D"
  },
  "source": "meta"
}