{
  "id": 305626,
  "title": "are we overfitting the LB?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/305626",
  "author_name": "",
  "post_date": "2022-02-06T09:00:44.761233300Z",
  "votes": 22,
  "comment_count": 30,
  "views": 0,
  "content": "<p>from my experiments, high resolution will score higher with LB, but not the CV. My CV scores the highest with the resolution that trained with. So i think there will be an earchquake with the private LB.<br>\nLuckly we have 4 condidates selections this time, i'm going to trust CV more than LB, but the highest LB will take one place.</p>\n<p>Here is my CV F2 score:</p>\n<p>video0 ~ 0.658<br>\nvideo1 ~ 0.639<br>\nvideo2 ~ 0.746</p>\n<p>So, what's your plan to handle the final shaking?</p>",
  "messages": [
    {
      "id": "1678121",
      "postDate": "02/06/2022 09:00:44",
      "content": "<p>from my experiments, high resolution will score higher with LB, but not the CV. My CV scores the highest with the resolution that trained with. So i think there will be an earchquake with the private LB.<br>\nLuckly we have 4 condidates selections this time, i'm going to trust CV more than LB, but the highest LB will take one place.</p>\n<p>Here is my CV F2 score:</p>\n<p>video0 ~ 0.658<br>\nvideo1 ~ 0.639<br>\nvideo2 ~ 0.746</p>\n<p>So, what's your plan to handle the final shaking?</p>",
      "rawMarkdown": "from my experiments, high resolution will score higher with LB, but not the CV. My CV scores the highest with the resolution that trained with. So i think there will be an earchquake with the private LB.\nLuckly we have 4 condidates selections this time, i'm going to trust CV more than LB, but the highest LB will take one place.\n\nHere is my CV F2 score:\n\nvideo0 ~ 0.658\nvideo1 ~ 0.639\nvideo2 ~ 0.746\n\nSo, what's your plan to handle the final shaking?",
      "votes": null
    },
    {
      "id": "1678143",
      "postDate": "02/06/2022 09:25:42",
      "content": "<p>(I have lb of .5) Same resolution makes sense, when high resolution like x10 ( 10000 vs 1280 ) makes no sense for me. Can you (or anybody) share your video_x folds cv? My average oof is about .72, video 1 being worst of them.  I cant imagine local cv showing something at x10 res. </p>",
      "rawMarkdown": "(I have lb of .5) Same resolution makes sense, when high resolution like x10 ( 10000 vs 1280 ) makes no sense for me. Can you (or anybody) share your video_x folds cv? My average oof is about .72, video 1 being worst of them.  I cant imagine local cv showing something at x10 res.",
      "votes": null
    },
    {
      "id": "1678177",
      "postDate": "02/06/2022 10:02:01",
      "content": "<p>Yea I have the same issue. The strange thing is the higher my CV is (easier training), the lower the LB score. I think we should have 1 best LB, 1 best CV, the others can be ensemble results.</p>",
      "rawMarkdown": "Yea I have the same issue. The strange thing is the higher my CV is (easier training), the lower the LB score. I think we should have 1 best LB, 1 best CV, the others can be ensemble results.",
      "votes": null
    },
    {
      "id": "1678184",
      "postDate": "02/06/2022 10:22:34",
      "content": "<p>hi, i update the topic content to add my local CV score. Average of mine CV score is 0.681.</p>",
      "rawMarkdown": "hi, i update the topic content to add my local CV score. Average of mine CV score is 0.681.",
      "votes": null
    },
    {
      "id": "1678188",
      "postDate": "02/06/2022 10:26:42",
      "content": "<p>yes, only if we can not train a model that works best both with LB and CV in the same resolution. I think there are already some guys find the secret that \"why higher resolution will benefit the public LB but not for CV\"</p>",
      "rawMarkdown": "yes, only if we can not train a model that works best both with LB and CV in the same resolution. I think there are already some guys find the secret that \"why higher resolution will benefit the public LB but not for CV\"",
      "votes": null
    },
    {
      "id": "1678192",
      "postDate": "02/06/2022 10:33:15",
      "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>  Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5? </p>",
      "rawMarkdown": "snaker  Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5?",
      "votes": null
    },
    {
      "id": "1678199",
      "postDate": "02/06/2022 10:37:15",
      "content": "<p>\"Average of mine CV score is 0.681.\"<br>\nHow much LB was 0.681CV giving you?</p>\n<p>did you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.</p>",
      "rawMarkdown": "\"Average of mine CV score is 0.681.\"\nHow much LB was 0.681CV giving you?\n\ndid you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.",
      "votes": null
    },
    {
      "id": "1678352",
      "postDate": "02/06/2022 13:23:16",
      "content": "<blockquote>\n  <p>Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5?</p>\n</blockquote>\n<p>for video1, it drop down to ~ 0.3, but if add TTA (which is the augment=True in yolov5) with high resolution it will go back to ~0.58.</p>",
      "rawMarkdown": "> Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5?\n\nfor video1, it drop down to ~ 0.3, but if add TTA (which is the augment=True in yolov5) with high resolution it will go back to ~0.58.",
      "votes": null
    },
    {
      "id": "1678354",
      "postDate": "02/06/2022 13:25:10",
      "content": "<blockquote>\n  <p>\"Average of mine CV score is 0.681.\"<br>\n  How much LB was 0.681CV giving you?<br>\n  did you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.</p>\n</blockquote>\n<p>i train only with image that has ground truth. The lb is my current score with high resolution and ensemble.</p>",
      "rawMarkdown": "> \"Average of mine CV score is 0.681.\"\nHow much LB was 0.681CV giving you?\ndid you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.\n\ni train only with image that has ground truth. The lb is my current score with high resolution and ensemble.",
      "votes": null
    },
    {
      "id": "1678369",
      "postDate": "02/06/2022 13:46:19",
      "content": "<p>We truly need advice from our heroes <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>",
      "rawMarkdown": "We truly need advice from our heroes @remekkinas and @hengck23.",
      "votes": null
    },
    {
      "id": "1678373",
      "postDate": "02/06/2022 13:49:31",
      "content": "<p>you forgot <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> </p>",
      "rawMarkdown": "you forgot @steamedsheep",
      "votes": null
    },
    {
      "id": "1678443",
      "postDate": "02/06/2022 14:31:28",
      "content": "<p>I am not a hero 😊 I am like most of people to learn here .. </p>\n<p>Yes, we can see shake up when \"resize all you need\" will be only one indicator to choose final solution. I am sure this will be overfitted solution (explanation later - now I can not say why (I will say after competition end); final score will depend on private dataset - when Kaggle staff take specific starfish (I use specific word now but we will change this word in a week) solution won't work. \"Resize is all you need\" makes model not able to generalize (I am sure …). My suggestion is to choose \"resize\" as one of submit and another … using local cv validated (I know that this is not easy because we do not know disrtibution of private dataset but … we should trust in local score).</p>",
      "rawMarkdown": "I am not a hero 😊 I am like most of people to learn here .. \n\nYes, we can see shake up when \"resize all you need\" will be only one indicator to choose final solution. I am sure this will be overfitted solution (explanation later - now I can not say why (I will say after competition end); final score will depend on private dataset - when Kaggle staff take specific starfish (I use specific word now but we will change this word in a week) solution won't work. \"Resize is all you need\" makes model not able to generalize (I am sure ...). My suggestion is to choose \"resize\" as one of submit and another ... using local cv validated (I know that this is not easy because we do not know disrtibution of private dataset but ... we should trust in local score).",
      "votes": null
    },
    {
      "id": "1678478",
      "postDate": "02/06/2022 14:56:26",
      "content": "<p>i'm thinking the same thing, chose the highest LB, and all of the rests rely on CV.</p>",
      "rawMarkdown": "i'm thinking the same thing, chose the highest LB, and all of the rests rely on CV.",
      "votes": null
    },
    {
      "id": "1678815",
      "postDate": "02/06/2022 19:55:59",
      "content": "<p>I've been using groupkfold on sequence id -&gt; training on only annotated images and using both annotation and non annotated for validation (kept only a random sample of it, about 1800~ with labels + 1300 empty) and using the yolov5 f2 marker (from another post), I can consistently get 0.8X and .7X CV but I haven't been able to reach those levels on public LB. I am guessing that maybe there are cots in the private LB in the first 25% with maybe many empty sequences in private LB. I will trust my CV for now. </p>",
      "rawMarkdown": "I've been using groupkfold on sequence id -> training on only annotated images and using both annotation and non annotated for validation (kept only a random sample of it, about 1800~ with labels + 1300 empty) and using the yolov5 f2 marker (from another post), I can consistently get 0.8X and .7X CV but I haven't been able to reach those levels on public LB. I am guessing that maybe there are cots in the private LB in the first 25% with maybe many empty sequences in private LB. I will trust my CV for now.",
      "votes": null
    },
    {
      "id": "1678825",
      "postDate": "02/06/2022 20:06:45",
      "content": "<p>Now that you said that <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> , I'm starting to understand where I'm going wrong… I won't comment any more so as not to spoil your work, but if that's the case, you have a chance of winning (and I want to see this brilliant solution). I don't think I'll have time to implement it, but I wish you luck.</p>",
      "rawMarkdown": "Now that you said that @remekkinas , I'm starting to understand where I'm going wrong... I won't comment any more so as not to spoil your work, but if that's the case, you have a chance of winning (and I want to see this brilliant solution). I don't think I'll have time to implement it, but I wish you luck.",
      "votes": null
    },
    {
      "id": "1678848",
      "postDate": "02/06/2022 20:24:46",
      "content": "<p>Thank you. I think there is many better candiates to win. We have the same puzzle to solve - choose solution which won't be overfitted :) Where we will be in priv score? For me it does not matter - the bes I can get from this competion was learnind - mission in 100% accomplished.  </p>",
      "rawMarkdown": "Thank you. I think there is many better candiates to win. We have the same puzzle to solve - choose solution which won't be overfitted :) Where we will be in priv score? For me it does not matter - the bes I can get from this competion was learnind - mission in 100% accomplished.",
      "votes": null
    },
    {
      "id": "1678906",
      "postDate": "02/06/2022 20:59:10",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> Do you know how resistant SAHI is to shakeup?</p>",
      "rawMarkdown": "remekkinas Do you know how resistant SAHI is to shakeup?",
      "votes": null
    },
    {
      "id": "1678915",
      "postDate": "02/06/2022 21:06:37",
      "content": "<p>What do you mean?</p>",
      "rawMarkdown": "What do you mean?",
      "votes": null
    },
    {
      "id": "1678925",
      "postDate": "02/06/2022 21:12:59",
      "content": "<p>hm, I should probably rephrase:</p>\n<p>I think its been kind of <em>proven</em> that increasing img_size to 2x or 3x isnt sustainable for private LB because of the large amount of FPs. I wanted to ask if you knew if SAHI was reliable enough that we could use it for private LB.</p>",
      "rawMarkdown": "hm, I should probably rephrase:\n\nI think its been kind of *proven* that increasing img_size to 2x or 3x isnt sustainable for private LB because of the large amount of FPs. I wanted to ask if you knew if SAHI was reliable enough that we could use it for private LB.",
      "votes": null
    },
    {
      "id": "1678930",
      "postDate": "02/06/2022 21:18:13",
      "content": "<p>To be honest I do not play with SAHI after initial tests. Why? <br>\nIt introduced many FN. Second I have no time unfortunately. This week I am going to make some test with SAHI. </p>",
      "rawMarkdown": "To be honest I do not play with SAHI after initial tests. Why? \nIt introduced many FN. Second I have no time unfortunately. This week I am going to make some test with SAHI.",
      "votes": null
    },
    {
      "id": "1678947",
      "postDate": "02/06/2022 21:34:14",
      "content": "<p><code>To be honest I do not play with SAHI after initial tests. Why?</code></p>\n<p>SAHI is the only alternative I know to decrease FNs other than increasing img_size. Just wanted to know if there was a stable alternative to img_size/SAHI</p>",
      "rawMarkdown": "`To be honest I do not play with SAHI after initial tests. Why?`\n\nSAHI is the only alternative I know to decrease FNs other than increasing img_size. Just wanted to know if there was a stable alternative to img_size/SAHI",
      "votes": null
    },
    {
      "id": "1679069",
      "postDate": "02/07/2022 00:24:14",
      "content": "<p>Did you consider data leak? Did you mean \"video0 ~ 0.658\" which was trained on video1 and video2? Additionally, what is your LB score?</p>",
      "rawMarkdown": "Did you consider data leak? Did you mean \"video0 ~ 0.658\" which was trained on video1 and video2? Additionally, what is your LB score?",
      "votes": null
    },
    {
      "id": "1679135",
      "postDate": "02/07/2022 02:31:24",
      "content": "<p>you can use 2-stage or 3-stage to reduce FN.</p>",
      "rawMarkdown": "you can use 2-stage or 3-stage to reduce FN.",
      "votes": null
    },
    {
      "id": "1679147",
      "postDate": "02/07/2022 02:42:46",
      "content": "<p>video0 ~ 0.658 means model trained with video1 and video2 and validate on video0. <br>\nI don't consider data leak cause there are many people scoring better than me.</p>\n<p>My LB score is 0.74.</p>",
      "rawMarkdown": "video0 ~ 0.658 means model trained with video1 and video2 and validate on video0. \nI don't consider data leak cause there are many people scoring better than me.\n\nMy LB score is 0.74.",
      "votes": null
    },
    {
      "id": "1679148",
      "postDate": "02/07/2022 02:45:56",
      "content": "<p>I use video_id to split so there are only 3 folds. The CV should be lower than other split strategy.</p>",
      "rawMarkdown": "I use video_id to split so there are only 3 folds. The CV should be lower than other split strategy.",
      "votes": null
    },
    {
      "id": "1679372",
      "postDate": "02/07/2022 07:14:28",
      "content": "<p>My models(m6,l6) trained with high resolution and bigger batchsize are giving CV of 0.7+ but all are failing on LB (&lt;0.65). Befuddled right now :|</p>",
      "rawMarkdown": "My models(m6,l6) trained with high resolution and bigger batchsize are giving CV of 0.7+ but all are failing on LB (<0.65). Befuddled right now :|",
      "votes": null
    },
    {
      "id": "1679381",
      "postDate": "02/07/2022 07:23:11",
      "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -&gt; better performance?</p>\n<p>Also an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2</p>",
      "rawMarkdown": "snaker your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -> better performance?\n\nAlso an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2",
      "votes": null
    },
    {
      "id": "1679508",
      "postDate": "02/07/2022 09:07:36",
      "content": "<p>Did you use empty images to calculate your F2?</p>",
      "rawMarkdown": "Did you use empty images to calculate your F2?",
      "votes": null
    },
    {
      "id": "1679553",
      "postDate": "02/07/2022 09:33:25",
      "content": "<p>yes, validation is on all of the fold. training is on samples with ground truth only</p>",
      "rawMarkdown": "yes, validation is on all of the fold. training is on samples with ground truth only",
      "votes": null
    },
    {
      "id": "1679555",
      "postDate": "02/07/2022 09:35:22",
      "content": "<p>if you want a high LB, you should train with a lower resolution and inference with higher resolution. This will benefit the small starfishes ( the public LB is full of them )</p>",
      "rawMarkdown": "if you want a high LB, you should train with a lower resolution and inference with higher resolution. This will benefit the small starfishes ( the public LB is full of them )",
      "votes": null
    },
    {
      "id": "1679569",
      "postDate": "02/07/2022 09:44:02",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -&gt; better performance?</p>\n</blockquote>\n<p>everyone split the dataset by <code>video_id</code> should have the same results, video2 will score higher than 0 and 1. and yes it's all because of the data amount and distribution.</p>\n<blockquote>\n  <p>Also an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2</p>\n</blockquote>\n<p>Thanks for point this out, i don't use COCO mAP so i have no idea, but it should help others.</p>",
      "rawMarkdown": "> @snaker your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -> better performance?\n\neveryone split the dataset by `video_id` should have the same results, video2 will score higher than 0 and 1. and yes it's all because of the data amount and distribution.\n\n> Also an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2\n\nThanks for point this out, i don't use COCO mAP so i have no idea, but it should help others.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1678143,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "02/06/2022 09:25:42",
      "content": "<p>(I have lb of .5) Same resolution makes sense, when high resolution like x10 ( 10000 vs 1280 ) makes no sense for me. Can you (or anybody) share your video_x folds cv? My average oof is about .72, video 1 being worst of them.  I cant imagine local cv showing something at x10 res. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1678184,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/06/2022 10:22:34",
          "content": "<p>hi, i update the topic content to add my local CV score. Average of mine CV score is 0.681.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678192,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "02/06/2022 10:33:15",
          "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>  Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678199,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "02/06/2022 10:37:15",
          "content": "<p>\"Average of mine CV score is 0.681.\"<br>\nHow much LB was 0.681CV giving you?</p>\n<p>did you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678352,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/06/2022 13:23:16",
          "content": "<blockquote>\n  <p>Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5?</p>\n</blockquote>\n<p>for video1, it drop down to ~ 0.3, but if add TTA (which is the augment=True in yolov5) with high resolution it will go back to ~0.58.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678354,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/06/2022 13:25:10",
          "content": "<blockquote>\n  <p>\"Average of mine CV score is 0.681.\"<br>\n  How much LB was 0.681CV giving you?<br>\n  did you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.</p>\n</blockquote>\n<p>i train only with image that has ground truth. The lb is my current score with high resolution and ensemble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1678177,
      "author_name": "anhnhunhat",
      "author_url": "",
      "post_date": "02/06/2022 10:02:01",
      "content": "<p>Yea I have the same issue. The strange thing is the higher my CV is (easier training), the lower the LB score. I think we should have 1 best LB, 1 best CV, the others can be ensemble results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1678188,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/06/2022 10:26:42",
          "content": "<p>yes, only if we can not train a model that works best both with LB and CV in the same resolution. I think there are already some guys find the secret that \"why higher resolution will benefit the public LB but not for CV\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1678369,
      "author_name": "aengusng",
      "author_url": "",
      "post_date": "02/06/2022 13:46:19",
      "content": "<p>We truly need advice from our heroes <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1678373,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/06/2022 13:49:31",
          "content": "<p>you forgot <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678443,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/06/2022 14:31:28",
          "content": "<p>I am not a hero 😊 I am like most of people to learn here .. </p>\n<p>Yes, we can see shake up when \"resize all you need\" will be only one indicator to choose final solution. I am sure this will be overfitted solution (explanation later - now I can not say why (I will say after competition end); final score will depend on private dataset - when Kaggle staff take specific starfish (I use specific word now but we will change this word in a week) solution won't work. \"Resize is all you need\" makes model not able to generalize (I am sure …). My suggestion is to choose \"resize\" as one of submit and another … using local cv validated (I know that this is not easy because we do not know disrtibution of private dataset but … we should trust in local score).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678478,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/06/2022 14:56:26",
          "content": "<p>i'm thinking the same thing, chose the highest LB, and all of the rests rely on CV.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678825,
          "author_name": "robsonsan",
          "author_url": "",
          "post_date": "02/06/2022 20:06:45",
          "content": "<p>Now that you said that <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> , I'm starting to understand where I'm going wrong… I won't comment any more so as not to spoil your work, but if that's the case, you have a chance of winning (and I want to see this brilliant solution). I don't think I'll have time to implement it, but I wish you luck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678848,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/06/2022 20:24:46",
          "content": "<p>Thank you. I think there is many better candiates to win. We have the same puzzle to solve - choose solution which won't be overfitted :) Where we will be in priv score? For me it does not matter - the bes I can get from this competion was learnind - mission in 100% accomplished.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1678815,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "02/06/2022 19:55:59",
      "content": "<p>I've been using groupkfold on sequence id -&gt; training on only annotated images and using both annotation and non annotated for validation (kept only a random sample of it, about 1800~ with labels + 1300 empty) and using the yolov5 f2 marker (from another post), I can consistently get 0.8X and .7X CV but I haven't been able to reach those levels on public LB. I am guessing that maybe there are cots in the private LB in the first 25% with maybe many empty sequences in private LB. I will trust my CV for now. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1679148,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/07/2022 02:45:56",
          "content": "<p>I use video_id to split so there are only 3 folds. The CV should be lower than other split strategy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1678906,
      "author_name": "kennyxie",
      "author_url": "",
      "post_date": "02/06/2022 20:59:10",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> Do you know how resistant SAHI is to shakeup?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1678915,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/06/2022 21:06:37",
          "content": "<p>What do you mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678925,
          "author_name": "kennyxie",
          "author_url": "",
          "post_date": "02/06/2022 21:12:59",
          "content": "<p>hm, I should probably rephrase:</p>\n<p>I think its been kind of <em>proven</em> that increasing img_size to 2x or 3x isnt sustainable for private LB because of the large amount of FPs. I wanted to ask if you knew if SAHI was reliable enough that we could use it for private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678930,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/06/2022 21:18:13",
          "content": "<p>To be honest I do not play with SAHI after initial tests. Why? <br>\nIt introduced many FN. Second I have no time unfortunately. This week I am going to make some test with SAHI. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1678947,
          "author_name": "kennyxie",
          "author_url": "",
          "post_date": "02/06/2022 21:34:14",
          "content": "<p><code>To be honest I do not play with SAHI after initial tests. Why?</code></p>\n<p>SAHI is the only alternative I know to decrease FNs other than increasing img_size. Just wanted to know if there was a stable alternative to img_size/SAHI</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1679135,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "02/07/2022 02:31:24",
          "content": "<p>you can use 2-stage or 3-stage to reduce FN.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1679069,
      "author_name": "lixxxxx",
      "author_url": "",
      "post_date": "02/07/2022 00:24:14",
      "content": "<p>Did you consider data leak? Did you mean \"video0 ~ 0.658\" which was trained on video1 and video2? Additionally, what is your LB score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1679147,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/07/2022 02:42:46",
          "content": "<p>video0 ~ 0.658 means model trained with video1 and video2 and validate on video0. <br>\nI don't consider data leak cause there are many people scoring better than me.</p>\n<p>My LB score is 0.74.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1679372,
      "author_name": "sanchitvj",
      "author_url": "",
      "post_date": "02/07/2022 07:14:28",
      "content": "<p>My models(m6,l6) trained with high resolution and bigger batchsize are giving CV of 0.7+ but all are failing on LB (&lt;0.65). Befuddled right now :|</p>",
      "votes": null,
      "replies": [
        {
          "id": 1679555,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/07/2022 09:35:22",
          "content": "<p>if you want a high LB, you should train with a lower resolution and inference with higher resolution. This will benefit the small starfishes ( the public LB is full of them )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1679381,
      "author_name": "alexchwong",
      "author_url": "",
      "post_date": "02/07/2022 07:23:11",
      "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -&gt; better performance?</p>\n<p>Also an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2</p>",
      "votes": null,
      "replies": [
        {
          "id": 1679569,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/07/2022 09:44:02",
          "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -&gt; better performance?</p>\n</blockquote>\n<p>everyone split the dataset by <code>video_id</code> should have the same results, video2 will score higher than 0 and 1. and yes it's all because of the data amount and distribution.</p>\n<blockquote>\n  <p>Also an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2</p>\n</blockquote>\n<p>Thanks for point this out, i don't use COCO mAP so i have no idea, but it should help others.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1679508,
      "author_name": "klawensliu",
      "author_url": "",
      "post_date": "02/07/2022 09:07:36",
      "content": "<p>Did you use empty images to calculate your F2?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1679553,
          "author_name": "snaker",
          "author_url": "",
          "post_date": "02/07/2022 09:33:25",
          "content": "<p>yes, validation is on all of the fold. training is on samples with ground truth only</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1678121": "from my experiments, high resolution will score higher with LB, but not the CV. My CV scores the highest with the resolution that trained with. So i think there will be an earchquake with the private LB.\nLuckly we have 4 condidates selections this time, i'm going to trust CV more than LB, but the highest LB will take one place.\n\nHere is my CV F2 score:\n\nvideo0 ~ 0.658\nvideo1 ~ 0.639\nvideo2 ~ 0.746\n\nSo, what's your plan to handle the final shaking?",
    "1678143": "(I have lb of .5) Same resolution makes sense, when high resolution like x10 ( 10000 vs 1280 ) makes no sense for me. Can you (or anybody) share your video_x folds cv? My average oof is about .72, video 1 being worst of them.  I cant imagine local cv showing something at x10 res.",
    "1678177": "Yea I have the same issue. The strange thing is the higher my CV is (easier training), the lower the LB score. I think we should have 1 best LB, 1 best CV, the others can be ensemble results.",
    "1678184": "hi, i update the topic content to add my local CV score. Average of mine CV score is 0.681.",
    "1678188": "yes, only if we can not train a model that works best both with LB and CV in the same resolution. I think there are already some guys find the secret that \"why higher resolution will benefit the public LB but not for CV\"",
    "1678192": "snaker  Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5?",
    "1678199": "\"Average of mine CV score is 0.681.\"\nHow much LB was 0.681CV giving you?\n\ndid you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.",
    "1678352": "> Thank you! Those certainly are on training resolution, and if you do your high-res trick from lb will they drop down to .5?\n\nfor video1, it drop down to ~ 0.3, but if add TTA (which is the augment=True in yolov5) with high resolution it will go back to ~0.58.",
    "1678354": "> \"Average of mine CV score is 0.681.\"\nHow much LB was 0.681CV giving you?\ndid you filter out the unlabeled images from your training set or you trained on the mixture of both labeled and unlabeled? I think using a mixture of both will be helpful, but could not try that yet.\n\ni train only with image that has ground truth. The lb is my current score with high resolution and ensemble.",
    "1678369": "We truly need advice from our heroes @remekkinas and @hengck23.",
    "1678373": "you forgot @steamedsheep",
    "1678443": "I am not a hero 😊 I am like most of people to learn here .. \n\nYes, we can see shake up when \"resize all you need\" will be only one indicator to choose final solution. I am sure this will be overfitted solution (explanation later - now I can not say why (I will say after competition end); final score will depend on private dataset - when Kaggle staff take specific starfish (I use specific word now but we will change this word in a week) solution won't work. \"Resize is all you need\" makes model not able to generalize (I am sure ...). My suggestion is to choose \"resize\" as one of submit and another ... using local cv validated (I know that this is not easy because we do not know disrtibution of private dataset but ... we should trust in local score).",
    "1678478": "i'm thinking the same thing, chose the highest LB, and all of the rests rely on CV.",
    "1678815": "I've been using groupkfold on sequence id -> training on only annotated images and using both annotation and non annotated for validation (kept only a random sample of it, about 1800~ with labels + 1300 empty) and using the yolov5 f2 marker (from another post), I can consistently get 0.8X and .7X CV but I haven't been able to reach those levels on public LB. I am guessing that maybe there are cots in the private LB in the first 25% with maybe many empty sequences in private LB. I will trust my CV for now.",
    "1678825": "Now that you said that @remekkinas , I'm starting to understand where I'm going wrong... I won't comment any more so as not to spoil your work, but if that's the case, you have a chance of winning (and I want to see this brilliant solution). I don't think I'll have time to implement it, but I wish you luck.",
    "1678848": "Thank you. I think there is many better candiates to win. We have the same puzzle to solve - choose solution which won't be overfitted :) Where we will be in priv score? For me it does not matter - the bes I can get from this competion was learnind - mission in 100% accomplished.",
    "1678906": "remekkinas Do you know how resistant SAHI is to shakeup?",
    "1678915": "What do you mean?",
    "1678925": "hm, I should probably rephrase:\n\nI think its been kind of *proven* that increasing img_size to 2x or 3x isnt sustainable for private LB because of the large amount of FPs. I wanted to ask if you knew if SAHI was reliable enough that we could use it for private LB.",
    "1678930": "To be honest I do not play with SAHI after initial tests. Why? \nIt introduced many FN. Second I have no time unfortunately. This week I am going to make some test with SAHI.",
    "1678947": "`To be honest I do not play with SAHI after initial tests. Why?`\n\nSAHI is the only alternative I know to decrease FNs other than increasing img_size. Just wanted to know if there was a stable alternative to img_size/SAHI",
    "1679069": "Did you consider data leak? Did you mean \"video0 ~ 0.658\" which was trained on video1 and video2? Additionally, what is your LB score?",
    "1679135": "you can use 2-stage or 3-stage to reduce FN.",
    "1679147": "video0 ~ 0.658 means model trained with video1 and video2 and validate on video0. \nI don't consider data leak cause there are many people scoring better than me.\n\nMy LB score is 0.74.",
    "1679148": "I use video_id to split so there are only 3 folds. The CV should be lower than other split strategy.",
    "1679372": "My models(m6,l6) trained with high resolution and bigger batchsize are giving CV of 0.7+ but all are failing on LB (<0.65). Befuddled right now :|",
    "1679381": "snaker your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -> better performance?\n\nAlso an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2",
    "1679508": "Did you use empty images to calculate your F2?",
    "1679553": "yes, validation is on all of the fold. training is on samples with ground truth only",
    "1679555": "if you want a high LB, you should train with a lower resolution and inference with higher resolution. This will benefit the small starfishes ( the public LB is full of them )",
    "1679569": "> @snaker your video2 model is giving you the best CV because it is the smallest video. Since you trained without the smallest video, the model has seen the most amount of data. More data -> better performance?\n\neveryone split the dataset by `video_id` should have the same results, video2 will score higher than 0 and 1. and yes it's all because of the data amount and distribution.\n\n> Also an interesting note that video 2 has smaller COTS than video 0 and 1; COCO mAP large does not output a value, which means it does not have any COTS with area greater than 96^2\n\nThanks for point this out, i don't use COCO mAP so i have no idea, but it should help others."
  },
  "source": "meta"
}