{
  "id": 219287,
  "title": "Shakeup expectations?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/219287",
  "author_name": "",
  "post_date": "2021-02-14T08:12:16.219186800Z",
  "votes": 2,
  "comment_count": 24,
  "views": 0,
  "content": "<p>My CV and Public LB scores have been reasonably well correlated and close</p>\n<p>What are you noticing? Do you expect a big shakeup? or scores more similar to Public LB based on your submissions?</p>\n<p>Thanks for sharing your thoughts!</p>",
  "messages": [
    {
      "id": "1199901",
      "postDate": "02/14/2021 08:12:16",
      "content": "<p>My CV and Public LB scores have been reasonably well correlated and close</p>\n<p>What are you noticing? Do you expect a big shakeup? or scores more similar to Public LB based on your submissions?</p>\n<p>Thanks for sharing your thoughts!</p>",
      "rawMarkdown": "My CV and Public LB scores have been reasonably well correlated and close\n\nWhat are you noticing? Do you expect a big shakeup? or scores more similar to Public LB based on your submissions?\n\nThanks for sharing your thoughts!",
      "votes": null
    },
    {
      "id": "1199959",
      "postDate": "02/14/2021 09:35:53",
      "content": "<p>Kaggle is \"overfitting test data\". </p>\n<p>Judging at the current public leaderboard, top kagglers are experts hence I would expect little shakeup.</p>\n<p>Top kagglers can actually \"visualize the test data\" … just like by looking at the reflection of the moon on the water, they know what the moon looks like in 3d. </p>",
      "rawMarkdown": "Kaggle is \"overfitting test data\". \n\nJudging at the current public leaderboard, top kagglers are experts hence I would expect little shakeup.\n\nTop kagglers can actually \"visualize the test data\" ... just like by looking at the reflection of the moon on the water, they know what the moon looks like in 3d.",
      "votes": null
    },
    {
      "id": "1199983",
      "postDate": "02/14/2021 10:05:03",
      "content": "<p>The top of the LB is not really crowded so I'd say the shake-up is not going to be really big. Although I'd love our score to go up by 0.01+ I don't think it will happen.</p>",
      "rawMarkdown": "The top of the LB is not really crowded so I'd say the shake-up is not going to be really big. Although I'd love our score to go up by 0.01+ I don't think it will happen.",
      "votes": null
    },
    {
      "id": "1200022",
      "postDate": "02/14/2021 10:46:21",
      "content": "<p>Agree! the top of the LB is not crowded! it's been such  a struggle and learning for me to get to even where I am!! Will be great to see and learn from what you and others in the top have done!</p>",
      "rawMarkdown": "Agree! the top of the LB is not crowded! it's been such  a struggle and learning for me to get to even where I am!! Will be great to see and learn from what you and others in the top have done!",
      "votes": null
    },
    {
      "id": "1200025",
      "postDate": "02/14/2021 10:48:12",
      "content": "<blockquote>\n  <p>Top kagglers can actually \"visualize the test data\"</p>\n</blockquote>\n<p>Hopefully, staying and competing in kaggle competitions will help learn from top kagglers like you and  learn to understand the test data better as well!</p>",
      "rawMarkdown": ">  Top kagglers can actually \"visualize the test data\"\n\nHopefully, staying and competing in kaggle competitions will help learn from top kagglers like you and  learn to understand the test data better as well!",
      "votes": null
    },
    {
      "id": "1200440",
      "postDate": "02/14/2021 16:57:44",
      "content": "<p>Public test set is quite small, and its hard to do proper validation. So I expect some shakeup</p>",
      "rawMarkdown": "Public test set is quite small, and its hard to do proper validation. So I expect some shakeup",
      "votes": null
    },
    {
      "id": "1200489",
      "postDate": "02/14/2021 18:01:14",
      "content": "<p>I made a 5 fold split and labeled one fold completely. Then I started to label the other 4 folds but I got only about 0.87 lwlrap score on the validation. But my LB was under 0.8 . I got ~0.95 on my validation, when I trained on them…</p>",
      "rawMarkdown": "I made a 5 fold split and labeled one fold completely. Then I started to label the other 4 folds but I got only about 0.87 lwlrap score on the validation. But my LB was under 0.8 . I got ~0.95 on my validation, when I trained on them...",
      "votes": null
    },
    {
      "id": "1200522",
      "postDate": "02/14/2021 18:29:00",
      "content": "<p>Yes, the Public LB is only 21% of the test data! Hoping for a shakeup and not getting shaken out!! 😄</p>\n<p>Also, great work!! eagerly waiting to see how you and team managed to get such awesome score! There will be lots of learning when your team shares your approach!!</p>",
      "rawMarkdown": "Yes, the Public LB is only 21% of the test data! Hoping for a shakeup and not getting shaken out!! 😄\n\nAlso, great work!! eagerly waiting to see how you and team managed to get such awesome score! There will be lots of learning when your team shares your approach!!",
      "votes": null
    },
    {
      "id": "1200626",
      "postDate": "02/14/2021 19:44:09",
      "content": "<p>fingers crossed! the private LB is also kind to us!</p>",
      "rawMarkdown": "fingers crossed! the private LB is also kind to us!",
      "votes": null
    },
    {
      "id": "1201057",
      "postDate": "02/15/2021 06:23:47",
      "content": "<p>Surely there will be some shakeup, but I don't think it will be so massive in this case unless some people are greatly overstating their model's capabilities with some leaderboard probing. There is definitely going to be directionality based on quality of model. </p>\n<p>If you are looking for an exciting shake-up terrain you should look at jane street. I have my popcorn ready for the blood bath. </p>",
      "rawMarkdown": "Surely there will be some shakeup, but I don't think it will be so massive in this case unless some people are greatly overstating their model's capabilities with some leaderboard probing. There is definitely going to be directionality based on quality of model. \n\nIf you are looking for an exciting shake-up terrain you should look at jane street. I have my popcorn ready for the blood bath.",
      "votes": null
    },
    {
      "id": "1201219",
      "postDate": "02/15/2021 08:44:08",
      "content": "<p>I think that VinBigData chest X-ray also has a great shakeup potential ;)</p>",
      "rawMarkdown": "I think that VinBigData chest X-ray also has a great shakeup potential ;)",
      "votes": null
    },
    {
      "id": "1201327",
      "postDate": "02/15/2021 10:21:54",
      "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> why do you think so?</p>",
      "rawMarkdown": "indswetrust why do you think so?",
      "votes": null
    },
    {
      "id": "1201360",
      "postDate": "02/15/2021 11:07:36",
      "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <br>\n10/90 public/private test set distribution.<br>\nLocal evaluation on a 10% valid subset shows a significant variation in mAP score. In addition, this mAP does not always correlate well with mAP evaluated on the full validation set - it tends to be higher.<br>\nIf not careful with validation, it is quite easy to overfit on the public data, especially if a model suddenly scores high on the public LB. </p>\n<p>I think this will cause a shakeup, unless everyone has a robust validation scheme and do not trust the public LB. </p>",
      "rawMarkdown": "aerdem4 \n10/90 public/private test set distribution.\nLocal evaluation on a 10% valid subset shows a significant variation in mAP score. In addition, this mAP does not always correlate well with mAP evaluated on the full validation set - it tends to be higher.\nIf not careful with validation, it is quite easy to overfit on the public data, especially if a model suddenly scores high on the public LB. \n\nI think this will cause a shakeup, unless everyone has a robust validation scheme and do not trust the public LB.",
      "votes": null
    },
    {
      "id": "1201460",
      "postDate": "02/15/2021 12:09:51",
      "content": "<p>\" test set is quite small\"</p>\n<p>this reminds me of a recent competition when I was ranked 2nd in public and end up 4th in public.<br>\nSince my target is to get to the top 3 and there is quite a gap to the public 4th lb score, I thought I would be safe. I still have a day left and a couple of submission slots which I didn't make good use of it. I overlooked the fact there that test data is small.</p>\n<p>In the final standing, I dropped from 2nd to 4th, while the 5th rose to 2nd. I compared his solution and my solution and found that I only ensemble about 10 models and he had used an amazing 100 models.</p>\n<p>hence I think \"robustness and generalization\" is something that one has to take care of when the test set is small</p>",
      "rawMarkdown": "\" test set is quite small\"\n\nthis reminds me of a recent competition when I was ranked 2nd in public and end up 4th in public.\nSince my target is to get to the top 3 and there is quite a gap to the public 4th lb score, I thought I would be safe. I still have a day left and a couple of submission slots which I didn't make good use of it. I overlooked the fact there that test data is small.\n\nIn the final standing, I dropped from 2nd to 4th, while the 5th rose to 2nd. I compared his solution and my solution and found that I only ensemble about 10 models and he had used an amazing 100 models.\n\nhence I think \"robustness and generalization\" is something that one has to take care of when the test set is small",
      "votes": null
    },
    {
      "id": "1207196",
      "postDate": "02/17/2021 18:34:42",
      "content": "<p>I expect some shakeup due to small LB test set and lack of proper training data. </p>\n<p>We will see soon!</p>",
      "rawMarkdown": "I expect some shakeup due to small LB test set and lack of proper training data. \n\nWe will see soon!",
      "votes": null
    },
    {
      "id": "1208103",
      "postDate": "02/18/2021 06:52:31",
      "content": "<p>Top kagglers can actually \"visualize the test data\"</p>\n<p>this is what I mean:<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342</a><br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a></p>\n<p>Quote:<br>\nCongratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956… This is crazy!</p>",
      "rawMarkdown": "Top kagglers can actually \"visualize the test data\"\n\nthis is what I mean:\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\nQuote:\nCongratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956… This is crazy!",
      "votes": null
    },
    {
      "id": "1208108",
      "postDate": "02/18/2021 06:56:21",
      "content": "<p>\"…I would expect little shakeup.\"</p>\n<p>this is because most top kagglers would already know by the last week that top solution will include pseudo label. they are basically training with the train (and maybe test set) and training with model uncertainty (embedded as pseudo label). Hence there will be little shakeup. </p>",
      "rawMarkdown": "\"...I would expect little shakeup.\"\n\nthis is because most top kagglers would already know by the last week that top solution will include pseudo label. they are basically training with the train (and maybe test set) and training with model uncertainty (embedded as pseudo label). Hence there will be little shakeup.",
      "votes": null
    },
    {
      "id": "1208231",
      "postDate": "02/18/2021 08:03:32",
      "content": "<p>The test set is not allowed for training, right?<br>\nI mean that is for measuring how well you will do on new data. But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.</p>",
      "rawMarkdown": "The test set is not allowed for training, right?\nI mean that is for measuring how well you will do on new data. But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.",
      "votes": null
    },
    {
      "id": "1208239",
      "postDate": "02/18/2021 08:05:45",
      "content": "<p>you are allowed to train on pseudo labels of test, but you are not allowed to manually label it</p>",
      "rawMarkdown": "you are allowed to train on pseudo labels of test, but you are not allowed to manually label it",
      "votes": null
    },
    {
      "id": "1208248",
      "postDate": "02/18/2021 08:12:57",
      "content": "<p>That's new to me. Is this a competition's or a kaggle's rule?</p>\n<p>Btw., congratulations on 1st place!</p>",
      "rawMarkdown": "That's new to me. Is this a competition's or a kaggle's rule?\n\nBtw., congratulations on 1st place!",
      "votes": null
    },
    {
      "id": "1208250",
      "postDate": "02/18/2021 08:13:41",
      "content": "<p>Kaggle's rule, you see this in many competitions. Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.</p>",
      "rawMarkdown": "Kaggle's rule, you see this in many competitions. Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.",
      "votes": null
    },
    {
      "id": "1208265",
      "postDate": "02/18/2021 08:20:26",
      "content": "<p>\". But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.\"</p>\n<p>you are correct. This is something that we will **not **do in the commercial application because we want train a model that works well and predict its performance in unseen data (black box testing). And in commercial applications, we seek not the best performance, but the most reliable performance.</p>\n<p>but kaggle is \"testdata overfitting\". we just want to see how good we can (it is actually whitish-gray box testing). there is an over-focus on getting the best accuracy. (and this is often not the most reliable one)</p>",
      "rawMarkdown": "\". But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.\"\n\nyou are correct. This is something that we will **not **do in the commercial application because we want train a model that works well and predict its performance in unseen data (black box testing). And in commercial applications, we seek not the best performance, but the most reliable performance.\n\nbut kaggle is \"testdata overfitting\". we just want to see how good we can (it is actually whitish-gray box testing). there is an over-focus on getting the best accuracy. (and this is often not the most reliable one)",
      "votes": null
    },
    {
      "id": "1208275",
      "postDate": "02/18/2021 08:23:51",
      "content": "<p>\"Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.\"</p>\n<p>distilling on model prediction results is dangerous. But I think self-supervised (i.e. ground-truth is in the data itself) is something more workable. e.g. inpainting block-masked spectrograms </p>",
      "rawMarkdown": "\"Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.\"\n\ndistilling on model prediction results is dangerous. But I think self-supervised (i.e. ground-truth is in the data itself) is something more workable. e.g. inpainting block-masked spectrograms",
      "votes": null
    },
    {
      "id": "1208293",
      "postDate": "02/18/2021 08:29:24",
      "content": "<p>I see, thank you for the clarification!! : )</p>",
      "rawMarkdown": "I see, thank you for the clarification!! : )",
      "votes": null
    },
    {
      "id": "1208325",
      "postDate": "02/18/2021 08:45:18",
      "content": "<p>We did use pseudos, but only in-sample pseudos. I think if you use oof or test pseudos, it is harder to break out of the predictions you already have, specifically if there is no real constraint regarding blending, otherwise you can utilize it for ensembling effects.</p>",
      "rawMarkdown": "We did use pseudos, but only in-sample pseudos. I think if you use oof or test pseudos, it is harder to break out of the predictions you already have, specifically if there is no real constraint regarding blending, otherwise you can utilize it for ensembling effects.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1199959,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/14/2021 09:35:53",
      "content": "<p>Kaggle is \"overfitting test data\". </p>\n<p>Judging at the current public leaderboard, top kagglers are experts hence I would expect little shakeup.</p>\n<p>Top kagglers can actually \"visualize the test data\" … just like by looking at the reflection of the moon on the water, they know what the moon looks like in 3d. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1200025,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/14/2021 10:48:12",
          "content": "<blockquote>\n  <p>Top kagglers can actually \"visualize the test data\"</p>\n</blockquote>\n<p>Hopefully, staying and competing in kaggle competitions will help learn from top kagglers like you and  learn to understand the test data better as well!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1200489,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "02/14/2021 18:01:14",
          "content": "<p>I made a 5 fold split and labeled one fold completely. Then I started to label the other 4 folds but I got only about 0.87 lwlrap score on the validation. But my LB was under 0.8 . I got ~0.95 on my validation, when I trained on them…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1200626,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/14/2021 19:44:09",
          "content": "<p>fingers crossed! the private LB is also kind to us!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208103,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 06:52:31",
          "content": "<p>Top kagglers can actually \"visualize the test data\"</p>\n<p>this is what I mean:<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342</a><br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a></p>\n<p>Quote:<br>\nCongratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956… This is crazy!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208108,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 06:56:21",
          "content": "<p>\"…I would expect little shakeup.\"</p>\n<p>this is because most top kagglers would already know by the last week that top solution will include pseudo label. they are basically training with the train (and maybe test set) and training with model uncertainty (embedded as pseudo label). Hence there will be little shakeup. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208231,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "02/18/2021 08:03:32",
          "content": "<p>The test set is not allowed for training, right?<br>\nI mean that is for measuring how well you will do on new data. But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208239,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 08:05:45",
          "content": "<p>you are allowed to train on pseudo labels of test, but you are not allowed to manually label it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208248,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "02/18/2021 08:12:57",
          "content": "<p>That's new to me. Is this a competition's or a kaggle's rule?</p>\n<p>Btw., congratulations on 1st place!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208250,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 08:13:41",
          "content": "<p>Kaggle's rule, you see this in many competitions. Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208265,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 08:20:26",
          "content": "<p>\". But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.\"</p>\n<p>you are correct. This is something that we will **not **do in the commercial application because we want train a model that works well and predict its performance in unseen data (black box testing). And in commercial applications, we seek not the best performance, but the most reliable performance.</p>\n<p>but kaggle is \"testdata overfitting\". we just want to see how good we can (it is actually whitish-gray box testing). there is an over-focus on getting the best accuracy. (and this is often not the most reliable one)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208275,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 08:23:51",
          "content": "<p>\"Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.\"</p>\n<p>distilling on model prediction results is dangerous. But I think self-supervised (i.e. ground-truth is in the data itself) is something more workable. e.g. inpainting block-masked spectrograms </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208293,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "02/18/2021 08:29:24",
          "content": "<p>I see, thank you for the clarification!! : )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208325,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 08:45:18",
          "content": "<p>We did use pseudos, but only in-sample pseudos. I think if you use oof or test pseudos, it is harder to break out of the predictions you already have, specifically if there is no real constraint regarding blending, otherwise you can utilize it for ensembling effects.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1199983,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/14/2021 10:05:03",
      "content": "<p>The top of the LB is not really crowded so I'd say the shake-up is not going to be really big. Although I'd love our score to go up by 0.01+ I don't think it will happen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1200022,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/14/2021 10:46:21",
          "content": "<p>Agree! the top of the LB is not crowded! it's been such  a struggle and learning for me to get to even where I am!! Will be great to see and learn from what you and others in the top have done!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1200440,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "02/14/2021 16:57:44",
      "content": "<p>Public test set is quite small, and its hard to do proper validation. So I expect some shakeup</p>",
      "votes": null,
      "replies": [
        {
          "id": 1200522,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/14/2021 18:29:00",
          "content": "<p>Yes, the Public LB is only 21% of the test data! Hoping for a shakeup and not getting shaken out!! 😄</p>\n<p>Also, great work!! eagerly waiting to see how you and team managed to get such awesome score! There will be lots of learning when your team shares your approach!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201460,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/15/2021 12:09:51",
          "content": "<p>\" test set is quite small\"</p>\n<p>this reminds me of a recent competition when I was ranked 2nd in public and end up 4th in public.<br>\nSince my target is to get to the top 3 and there is quite a gap to the public 4th lb score, I thought I would be safe. I still have a day left and a couple of submission slots which I didn't make good use of it. I overlooked the fact there that test data is small.</p>\n<p>In the final standing, I dropped from 2nd to 4th, while the 5th rose to 2nd. I compared his solution and my solution and found that I only ensemble about 10 models and he had used an amazing 100 models.</p>\n<p>hence I think \"robustness and generalization\" is something that one has to take care of when the test set is small</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1201057,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/15/2021 06:23:47",
      "content": "<p>Surely there will be some shakeup, but I don't think it will be so massive in this case unless some people are greatly overstating their model's capabilities with some leaderboard probing. There is definitely going to be directionality based on quality of model. </p>\n<p>If you are looking for an exciting shake-up terrain you should look at jane street. I have my popcorn ready for the blood bath. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1201219,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "02/15/2021 08:44:08",
          "content": "<p>I think that VinBigData chest X-ray also has a great shakeup potential ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201327,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "02/15/2021 10:21:54",
          "content": "<p><a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> why do you think so?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201360,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "02/15/2021 11:07:36",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <br>\n10/90 public/private test set distribution.<br>\nLocal evaluation on a 10% valid subset shows a significant variation in mAP score. In addition, this mAP does not always correlate well with mAP evaluated on the full validation set - it tends to be higher.<br>\nIf not careful with validation, it is quite easy to overfit on the public data, especially if a model suddenly scores high on the public LB. </p>\n<p>I think this will cause a shakeup, unless everyone has a robust validation scheme and do not trust the public LB. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207196,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "02/17/2021 18:34:42",
      "content": "<p>I expect some shakeup due to small LB test set and lack of proper training data. </p>\n<p>We will see soon!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1199901": "My CV and Public LB scores have been reasonably well correlated and close\n\nWhat are you noticing? Do you expect a big shakeup? or scores more similar to Public LB based on your submissions?\n\nThanks for sharing your thoughts!",
    "1199959": "Kaggle is \"overfitting test data\". \n\nJudging at the current public leaderboard, top kagglers are experts hence I would expect little shakeup.\n\nTop kagglers can actually \"visualize the test data\" ... just like by looking at the reflection of the moon on the water, they know what the moon looks like in 3d.",
    "1199983": "The top of the LB is not really crowded so I'd say the shake-up is not going to be really big. Although I'd love our score to go up by 0.01+ I don't think it will happen.",
    "1200022": "Agree! the top of the LB is not crowded! it's been such  a struggle and learning for me to get to even where I am!! Will be great to see and learn from what you and others in the top have done!",
    "1200025": ">  Top kagglers can actually \"visualize the test data\"\n\nHopefully, staying and competing in kaggle competitions will help learn from top kagglers like you and  learn to understand the test data better as well!",
    "1200440": "Public test set is quite small, and its hard to do proper validation. So I expect some shakeup",
    "1200489": "I made a 5 fold split and labeled one fold completely. Then I started to label the other 4 folds but I got only about 0.87 lwlrap score on the validation. But my LB was under 0.8 . I got ~0.95 on my validation, when I trained on them...",
    "1200522": "Yes, the Public LB is only 21% of the test data! Hoping for a shakeup and not getting shaken out!! 😄\n\nAlso, great work!! eagerly waiting to see how you and team managed to get such awesome score! There will be lots of learning when your team shares your approach!!",
    "1200626": "fingers crossed! the private LB is also kind to us!",
    "1201057": "Surely there will be some shakeup, but I don't think it will be so massive in this case unless some people are greatly overstating their model's capabilities with some leaderboard probing. There is definitely going to be directionality based on quality of model. \n\nIf you are looking for an exciting shake-up terrain you should look at jane street. I have my popcorn ready for the blood bath.",
    "1201219": "I think that VinBigData chest X-ray also has a great shakeup potential ;)",
    "1201327": "indswetrust why do you think so?",
    "1201360": "aerdem4 \n10/90 public/private test set distribution.\nLocal evaluation on a 10% valid subset shows a significant variation in mAP score. In addition, this mAP does not always correlate well with mAP evaluated on the full validation set - it tends to be higher.\nIf not careful with validation, it is quite easy to overfit on the public data, especially if a model suddenly scores high on the public LB. \n\nI think this will cause a shakeup, unless everyone has a robust validation scheme and do not trust the public LB.",
    "1201460": "\" test set is quite small\"\n\nthis reminds me of a recent competition when I was ranked 2nd in public and end up 4th in public.\nSince my target is to get to the top 3 and there is quite a gap to the public 4th lb score, I thought I would be safe. I still have a day left and a couple of submission slots which I didn't make good use of it. I overlooked the fact there that test data is small.\n\nIn the final standing, I dropped from 2nd to 4th, while the 5th rose to 2nd. I compared his solution and my solution and found that I only ensemble about 10 models and he had used an amazing 100 models.\n\nhence I think \"robustness and generalization\" is something that one has to take care of when the test set is small",
    "1207196": "I expect some shakeup due to small LB test set and lack of proper training data. \n\nWe will see soon!",
    "1208103": "Top kagglers can actually \"visualize the test data\"\n\nthis is what I mean:\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\nQuote:\nCongratulations! I multiplied \"s3\" and \"s18\" by 3 of our best submission, and it gave us Public 0.944 and Private 0.956… This is crazy!",
    "1208108": "\"...I would expect little shakeup.\"\n\nthis is because most top kagglers would already know by the last week that top solution will include pseudo label. they are basically training with the train (and maybe test set) and training with model uncertainty (embedded as pseudo label). Hence there will be little shakeup.",
    "1208231": "The test set is not allowed for training, right?\nI mean that is for measuring how well you will do on new data. But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.",
    "1208239": "you are allowed to train on pseudo labels of test, but you are not allowed to manually label it",
    "1208248": "That's new to me. Is this a competition's or a kaggle's rule?\n\nBtw., congratulations on 1st place!",
    "1208250": "Kaggle's rule, you see this in many competitions. Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.",
    "1208265": "\". But if your model trains on the test, even if they are pseudo labels, then you go with your model to the wild and your model suddenly is worse than you expected based on the test results.\"\n\nyou are correct. This is something that we will **not **do in the commercial application because we want train a model that works well and predict its performance in unseen data (black box testing). And in commercial applications, we seek not the best performance, but the most reliable performance.\n\nbut kaggle is \"testdata overfitting\". we just want to see how good we can (it is actually whitish-gray box testing). there is an over-focus on getting the best accuracy. (and this is often not the most reliable one)",
    "1208275": "\"Personally I don't believe too much in using test pseudos as it is mostly a distillation effect though.\"\n\ndistilling on model prediction results is dangerous. But I think self-supervised (i.e. ground-truth is in the data itself) is something more workable. e.g. inpainting block-masked spectrograms",
    "1208293": "I see, thank you for the clarification!! : )",
    "1208325": "We did use pseudos, but only in-sample pseudos. I think if you use oof or test pseudos, it is harder to break out of the predictions you already have, specifically if there is no real constraint regarding blending, otherwise you can utilize it for ensembling effects."
  },
  "source": "meta"
}