{
  "id": 301132,
  "title": "Training High Res ✅ Inference High Res ⚠️",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/301132",
  "author_name": "",
  "post_date": "2022-01-16T03:36:52.891978300Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>sheep first <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638\" target=\"_blank\">introduced</a> how high res training and using that image size while testing can give a big boost in results. This was infact a useful technique and can be verified during cross validation as well (but only upto certain limits). </p>\n<p>After this on hengck's observations we saw that for yolov5s6 models, using an image size of upto 2.5 times of what was used during training also improved the Public LB scores almost linearly using certain steps. However this observation could not pass the <a href=\"https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res\" target=\"_blank\">cross validation tests</a>(at-least according to my present metric calculations) and instead showed a dip with increasing image size. </p>\n<p>This contradicting observation should cation us to tread lightly while using this trick during inference. I think that certain strong augmentation techniques might help equalise this discrepancy </p>\n<p>UPD: seems like the way sheep splits data (i.e. by video ID) benefits from this up-sizing technique while other split techniques (random or subsequence based or other you may be using) may not (but only upto ~1.5 times not 2.5 as on public LB). </p>\n<p>I do note however that my CV at same image size used during training with subsequences split was comparable to CV at upsized inference for training with video ID split. You can check the different versions of my cross val notebook where I've performed these experiments to verify</p>",
  "messages": [
    {
      "id": "1651747",
      "postDate": "01/16/2022 03:36:52",
      "content": "<p>sheep first <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638\" target=\"_blank\">introduced</a> how high res training and using that image size while testing can give a big boost in results. This was infact a useful technique and can be verified during cross validation as well (but only upto certain limits). </p>\n<p>After this on hengck's observations we saw that for yolov5s6 models, using an image size of upto 2.5 times of what was used during training also improved the Public LB scores almost linearly using certain steps. However this observation could not pass the <a href=\"https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res\" target=\"_blank\">cross validation tests</a>(at-least according to my present metric calculations) and instead showed a dip with increasing image size. </p>\n<p>This contradicting observation should cation us to tread lightly while using this trick during inference. I think that certain strong augmentation techniques might help equalise this discrepancy </p>\n<p>UPD: seems like the way sheep splits data (i.e. by video ID) benefits from this up-sizing technique while other split techniques (random or subsequence based or other you may be using) may not (but only upto ~1.5 times not 2.5 as on public LB). </p>\n<p>I do note however that my CV at same image size used during training with subsequences split was comparable to CV at upsized inference for training with video ID split. You can check the different versions of my cross val notebook where I've performed these experiments to verify</p>",
      "rawMarkdown": "sheep first [introduced](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638) how high res training and using that image size while testing can give a big boost in results. This was infact a useful technique and can be verified during cross validation as well (but only upto certain limits). \n\nAfter this on hengck's observations we saw that for yolov5s6 models, using an image size of upto 2.5 times of what was used during training also improved the Public LB scores almost linearly using certain steps. However this observation could not pass the [cross validation tests](https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res)(at-least according to my present metric calculations) and instead showed a dip with increasing image size. \n\nThis contradicting observation should cation us to tread lightly while using this trick during inference. I think that certain strong augmentation techniques might help equalise this discrepancy \n\nUPD: seems like the way sheep splits data (i.e. by video ID) benefits from this up-sizing technique while other split techniques (random or subsequence based or other you may be using) may not (but only upto ~1.5 times not 2.5 as on public LB). \n\nI do note however that my CV at same image size used during training with subsequences split was comparable to CV at upsized inference for training with video ID split. You can check the different versions of my cross val notebook where I've performed these experiments to verify",
      "votes": null
    },
    {
      "id": "1651790",
      "postDate": "01/16/2022 04:56:07",
      "content": "<p>there are many missed small cots objects or I would say the smaller cots in the public test set are more difficult than the train set. increasing size at inference would help TP, but at the expense of more FP. Since the metric is F2, there is an overall gain in the metric.</p>\n<p>i think the question is not if a bigger inference size would help or not.</p>\n<p>rather, it is how large would we use? i feel that 2.5x is too much.</p>\n<p>also, how to keep the results stable?</p>\n<p>it is difficult to test using local cv because the hidden test and given train set is different.</p>\n<p>i can only think of probing the test set, e.g</p>\n<ul>\n<li>submit at original inference size and bbox of size &gt; xxx</li>\n<li>submit at larger inference size and bbox of size &gt; xxx</li>\n<li>submit at larger inference size and bbox of size &lt; xxx</li>\n<li>submit at original inference size  for all box</li>\n</ul>\n<hr>\n<p>instead of just looking at the numbers, one should also plot the truth and predicted boxes for visualization<br>\nit is easy to observe that during the tracklet, initial cots from afar is not detected. as the camera moves closer and the cots get bigger, it becomes detected.</p>\n<hr>\n<p>it is also not about making the cv split (data are too little), but rather how to simulate all the possible likely cases in cross-validation. maybe your cross validation should also test data that are augmented. e.g use  downsized images, etc</p>",
      "rawMarkdown": "there are many missed small cots objects or I would say the smaller cots in the public test set are more difficult than the train set. increasing size at inference would help TP, but at the expense of more FP. Since the metric is F2, there is an overall gain in the metric.\n\ni think the question is not if a bigger inference size would help or not.\n\nrather, it is how large would we use? i feel that 2.5x is too much.\n\nalso, how to keep the results stable?\n\nit is difficult to test using local cv because the hidden test and given train set is different.\n\ni can only think of probing the test set, e.g\n- submit at original inference size and bbox of size > xxx\n- submit at larger inference size and bbox of size > xxx\n- submit at larger inference size and bbox of size < xxx\n- submit at original inference size  for all box\n\n---\n\ninstead of just looking at the numbers, one should also plot the truth and predicted boxes for visualization\nit is easy to observe that during the tracklet, initial cots from afar is not detected. as the camera moves closer and the cots get bigger, it becomes detected.\n\n---\n\nit is also not about making the cv split (data are too little), but rather how to simulate all the possible likely cases in cross-validation. maybe your cross validation should also test data that are augmented. e.g use  downsized images, etc",
      "votes": null
    },
    {
      "id": "1651791",
      "postDate": "01/16/2022 04:56:40",
      "content": "<p>Hi! How are you splitting your data, to calculate your CV?</p>",
      "rawMarkdown": "Hi! How are you splitting your data, to calculate your CV?",
      "votes": null
    },
    {
      "id": "1651815",
      "postDate": "01/16/2022 05:21:03",
      "content": "<p>you can check my notebook <a href=\"https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res\" target=\"_blank\">here</a>.  I do splitting and creating folds in the notebook using the same code as I did in Colab so that there's no leak in CV from different environments</p>",
      "rawMarkdown": "you can check my notebook [here](https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res).  I do splitting and creating folds in the notebook using the same code as I did in Colab so that there's no leak in CV from different environments",
      "votes": null
    },
    {
      "id": "1651850",
      "postDate": "01/16/2022 05:55:08",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> <br>\nIt has been said that video-based splitting is better than subsequence-based splitting.<br>\nI hope you have seen this topic, it don't really talks about the difference but it talks about video based splitting,</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723\" target=\"_blank\">Be CAREFUL with your TRAIN/VALID splits and avoid LB GAPS</a></li>\n</ul>",
      "rawMarkdown": "Hello @mrinath \nIt has been said that video-based splitting is better than subsequence-based splitting.\nI hope you have seen this topic, it don't really talks about the difference but it talks about video based splitting,\n- [Be CAREFUL with your TRAIN/VALID splits and avoid LB GAPS](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723)",
      "votes": null
    },
    {
      "id": "1651885",
      "postDate": "01/16/2022 06:41:58",
      "content": "<p>Hmm as you rightly stated the topic only mentions the benefits of video based splitting. In my own limited number of experiments, I think that subsequences based splitting fared a bit better than the method used by Adriano in his notebook of using 6 k samples and splitting by video. </p>\n<p>Basically different sequences also contain different kind of features with respect to each other (since they’re part of 3 very big videos) so there shouldn’t be cases such as those mentioned in the Adrianos topic in a large quantity. As hengck has mentioned it’s more about detecting those really small COTS so video vs sub sequence splitting shouldn’t really give that huge a CV difference. Random split will probably be significantly worse but haven’t tried that yet</p>",
      "rawMarkdown": "Hmm as you rightly stated the topic only mentions the benefits of video based splitting. In my own limited number of experiments, I think that subsequences based splitting fared a bit better than the method used by Adriano in his notebook of using 6 k samples and splitting by video. \n\nBasically different sequences also contain different kind of features with respect to each other (since they’re part of 3 very big videos) so there shouldn’t be cases such as those mentioned in the Adrianos topic in a large quantity. As hengck has mentioned it’s more about detecting those really small COTS so video vs sub sequence splitting shouldn’t really give that huge a CV difference. Random split will probably be significantly worse but haven’t tried that yet",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1651790,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/16/2022 04:56:07",
      "content": "<p>there are many missed small cots objects or I would say the smaller cots in the public test set are more difficult than the train set. increasing size at inference would help TP, but at the expense of more FP. Since the metric is F2, there is an overall gain in the metric.</p>\n<p>i think the question is not if a bigger inference size would help or not.</p>\n<p>rather, it is how large would we use? i feel that 2.5x is too much.</p>\n<p>also, how to keep the results stable?</p>\n<p>it is difficult to test using local cv because the hidden test and given train set is different.</p>\n<p>i can only think of probing the test set, e.g</p>\n<ul>\n<li>submit at original inference size and bbox of size &gt; xxx</li>\n<li>submit at larger inference size and bbox of size &gt; xxx</li>\n<li>submit at larger inference size and bbox of size &lt; xxx</li>\n<li>submit at original inference size  for all box</li>\n</ul>\n<hr>\n<p>instead of just looking at the numbers, one should also plot the truth and predicted boxes for visualization<br>\nit is easy to observe that during the tracklet, initial cots from afar is not detected. as the camera moves closer and the cots get bigger, it becomes detected.</p>\n<hr>\n<p>it is also not about making the cv split (data are too little), but rather how to simulate all the possible likely cases in cross-validation. maybe your cross validation should also test data that are augmented. e.g use  downsized images, etc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1651791,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "01/16/2022 04:56:40",
      "content": "<p>Hi! How are you splitting your data, to calculate your CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1651815,
          "author_name": "ferlockx",
          "author_url": "",
          "post_date": "01/16/2022 05:21:03",
          "content": "<p>you can check my notebook <a href=\"https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res\" target=\"_blank\">here</a>.  I do splitting and creating folds in the notebook using the same code as I did in Colab so that there's no leak in CV from different environments</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1651850,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "01/16/2022 05:55:08",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> <br>\nIt has been said that video-based splitting is better than subsequence-based splitting.<br>\nI hope you have seen this topic, it don't really talks about the difference but it talks about video based splitting,</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723\" target=\"_blank\">Be CAREFUL with your TRAIN/VALID splits and avoid LB GAPS</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1651885,
          "author_name": "ferlockx",
          "author_url": "",
          "post_date": "01/16/2022 06:41:58",
          "content": "<p>Hmm as you rightly stated the topic only mentions the benefits of video based splitting. In my own limited number of experiments, I think that subsequences based splitting fared a bit better than the method used by Adriano in his notebook of using 6 k samples and splitting by video. </p>\n<p>Basically different sequences also contain different kind of features with respect to each other (since they’re part of 3 very big videos) so there shouldn’t be cases such as those mentioned in the Adrianos topic in a large quantity. As hengck has mentioned it’s more about detecting those really small COTS so video vs sub sequence splitting shouldn’t really give that huge a CV difference. Random split will probably be significantly worse but haven’t tried that yet</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1651747": "sheep first [introduced](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638) how high res training and using that image size while testing can give a big boost in results. This was infact a useful technique and can be verified during cross validation as well (but only upto certain limits). \n\nAfter this on hengck's observations we saw that for yolov5s6 models, using an image size of upto 2.5 times of what was used during training also improved the Public LB scores almost linearly using certain steps. However this observation could not pass the [cross validation tests](https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res)(at-least according to my present metric calculations) and instead showed a dip with increasing image size. \n\nThis contradicting observation should cation us to tread lightly while using this trick during inference. I think that certain strong augmentation techniques might help equalise this discrepancy \n\nUPD: seems like the way sheep splits data (i.e. by video ID) benefits from this up-sizing technique while other split techniques (random or subsequence based or other you may be using) may not (but only upto ~1.5 times not 2.5 as on public LB). \n\nI do note however that my CV at same image size used during training with subsequences split was comparable to CV at upsized inference for training with video ID split. You can check the different versions of my cross val notebook where I've performed these experiments to verify",
    "1651790": "there are many missed small cots objects or I would say the smaller cots in the public test set are more difficult than the train set. increasing size at inference would help TP, but at the expense of more FP. Since the metric is F2, there is an overall gain in the metric.\n\ni think the question is not if a bigger inference size would help or not.\n\nrather, it is how large would we use? i feel that 2.5x is too much.\n\nalso, how to keep the results stable?\n\nit is difficult to test using local cv because the hidden test and given train set is different.\n\ni can only think of probing the test set, e.g\n- submit at original inference size and bbox of size > xxx\n- submit at larger inference size and bbox of size > xxx\n- submit at larger inference size and bbox of size < xxx\n- submit at original inference size  for all box\n\n---\n\ninstead of just looking at the numbers, one should also plot the truth and predicted boxes for visualization\nit is easy to observe that during the tracklet, initial cots from afar is not detected. as the camera moves closer and the cots get bigger, it becomes detected.\n\n---\n\nit is also not about making the cv split (data are too little), but rather how to simulate all the possible likely cases in cross-validation. maybe your cross validation should also test data that are augmented. e.g use  downsized images, etc",
    "1651791": "Hi! How are you splitting your data, to calculate your CV?",
    "1651815": "you can check my notebook [here](https://www.kaggle.com/ferlockx/f2-cv-0-77-subseq-split-and-higher-res).  I do splitting and creating folds in the notebook using the same code as I did in Colab so that there's no leak in CV from different environments",
    "1651850": "Hello @mrinath \nIt has been said that video-based splitting is better than subsequence-based splitting.\nI hope you have seen this topic, it don't really talks about the difference but it talks about video based splitting,\n- [Be CAREFUL with your TRAIN/VALID splits and avoid LB GAPS](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723)",
    "1651885": "Hmm as you rightly stated the topic only mentions the benefits of video based splitting. In my own limited number of experiments, I think that subsequences based splitting fared a bit better than the method used by Adriano in his notebook of using 6 k samples and splitting by video. \n\nBasically different sequences also contain different kind of features with respect to each other (since they’re part of 3 very big videos) so there shouldn’t be cases such as those mentioned in the Adrianos topic in a large quantity. As hengck has mentioned it’s more about detecting those really small COTS so video vs sub sequence splitting shouldn’t really give that huge a CV difference. Random split will probably be significantly worse but haven’t tried that yet"
  },
  "source": "meta"
}