{
  "id": 293723,
  "title": "Be CAREFUL with your TRAIN/VALID splits and avoid LB GAPS",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/293723",
  "author_name": "",
  "post_date": "2021-12-06T16:27:16.924180900Z",
  "votes": 128,
  "comment_count": 41,
  "views": 0,
  "content": "<p>From what I could tell, many of the public notebooks split the train/valid datasets at random. Keep in mind that video data have stupidly high temporal coherence and by splitting the dataset randomly you will end up with very similar images in both sets.</p>\n<p>Consider the following example:<br>\n<img src=\"https://i.imgur.com/C5eB8bJ.png\" alt=\"\"><br>\n<em>Two consecutive frames of video 2</em></p>\n<p>If you end up having the left frame on the train set and the right one in the validation set it is very likely that the model will correctly predict those detections. However, when running inference, the model will be shown to a different video and it is unlikely to be as good.</p>\n<p>I trained two models to demonstrate this effect. </p>\n<ol>\n<li>Randomly splitting the dataset.</li>\n<li>Using videos 0 and 1 for training and video 2 for validation.</li>\n</ol>\n<p>In both cases, I used 6k images (~5k with labels and 1k~ with background only).</p>\n<p>Although both models trained roughly the same (aka. had the same training loss curve) their validation curves were significantly different.</p>\n<p><img src=\"https://i.imgur.com/mKOPV8t.png\" alt=\"\"><br>\n<em>Training losses</em><br>\n<img src=\"https://i.imgur.com/E9IZSc7.png\" alt=\"\"><br>\n<em>Validation losses</em></p>\n<p>The decrease in performance is easily explained by a substantial drop in <strong>recall</strong><br>\n<img src=\"https://i.imgur.com/4APbcdz.png\" alt=\"\"><br>\n<em>Recall and Precision curves</em></p>\n<p><img src=\"https://i.imgur.com/vBWzAfQ.png\" alt=\"\"><br>\n<em>Precision vs. Recall for both models (model 1 on the left)</em></p>\n<p>That being said, I believe that the best split we could do for this dataset is 3-fold (1 for each video) and do a final ensemble using <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">WBF</a> or some other ensembling technique.</p>\n<p>I made my notebooks public so feel free to use them:</p>\n<ul>\n<li>YoloV5 dataset generator with correct train/test splits: <a href=\"https://www.kaggle.com/coldfir3/efficient-yolov5-dataset-generator\" target=\"_blank\">link</a></li>\n<li>YoloV5 train: <a href=\"https://www.kaggle.com/coldfir3/yolov5-train\" target=\"_blank\">link</a></li>\n<li>YoloV5 inference (WIP): link</li>\n</ul>",
  "messages": [
    {
      "id": "1608778",
      "postDate": "12/06/2021 16:27:16",
      "content": "<p>From what I could tell, many of the public notebooks split the train/valid datasets at random. Keep in mind that video data have stupidly high temporal coherence and by splitting the dataset randomly you will end up with very similar images in both sets.</p>\n<p>Consider the following example:<br>\n<img src=\"https://i.imgur.com/C5eB8bJ.png\" alt=\"\"><br>\n<em>Two consecutive frames of video 2</em></p>\n<p>If you end up having the left frame on the train set and the right one in the validation set it is very likely that the model will correctly predict those detections. However, when running inference, the model will be shown to a different video and it is unlikely to be as good.</p>\n<p>I trained two models to demonstrate this effect. </p>\n<ol>\n<li>Randomly splitting the dataset.</li>\n<li>Using videos 0 and 1 for training and video 2 for validation.</li>\n</ol>\n<p>In both cases, I used 6k images (~5k with labels and 1k~ with background only).</p>\n<p>Although both models trained roughly the same (aka. had the same training loss curve) their validation curves were significantly different.</p>\n<p><img src=\"https://i.imgur.com/mKOPV8t.png\" alt=\"\"><br>\n<em>Training losses</em><br>\n<img src=\"https://i.imgur.com/E9IZSc7.png\" alt=\"\"><br>\n<em>Validation losses</em></p>\n<p>The decrease in performance is easily explained by a substantial drop in <strong>recall</strong><br>\n<img src=\"https://i.imgur.com/4APbcdz.png\" alt=\"\"><br>\n<em>Recall and Precision curves</em></p>\n<p><img src=\"https://i.imgur.com/vBWzAfQ.png\" alt=\"\"><br>\n<em>Precision vs. Recall for both models (model 1 on the left)</em></p>\n<p>That being said, I believe that the best split we could do for this dataset is 3-fold (1 for each video) and do a final ensemble using <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">WBF</a> or some other ensembling technique.</p>\n<p>I made my notebooks public so feel free to use them:</p>\n<ul>\n<li>YoloV5 dataset generator with correct train/test splits: <a href=\"https://www.kaggle.com/coldfir3/efficient-yolov5-dataset-generator\" target=\"_blank\">link</a></li>\n<li>YoloV5 train: <a href=\"https://www.kaggle.com/coldfir3/yolov5-train\" target=\"_blank\">link</a></li>\n<li>YoloV5 inference (WIP): link</li>\n</ul>",
      "rawMarkdown": "From what I could tell, many of the public notebooks split the train/valid datasets at random. Keep in mind that video data have stupidly high temporal coherence and by splitting the dataset randomly you will end up with very similar images in both sets.\n\nConsider the following example:\n![](https://i.imgur.com/C5eB8bJ.png)\n*Two consecutive frames of video 2*\n\nIf you end up having the left frame on the train set and the right one in the validation set it is very likely that the model will correctly predict those detections. However, when running inference, the model will be shown to a different video and it is unlikely to be as good.\n\nI trained two models to demonstrate this effect. \n1. Randomly splitting the dataset.\n2. Using videos 0 and 1 for training and video 2 for validation.\n\nIn both cases, I used 6k images (~5k with labels and 1k~ with background only).\n\nAlthough both models trained roughly the same (aka. had the same training loss curve) their validation curves were significantly different.\n\n![](https://i.imgur.com/mKOPV8t.png)\n*Training losses*\n![](https://i.imgur.com/E9IZSc7.png)\n*Validation losses*\n\nThe decrease in performance is easily explained by a substantial drop in **recall**\n![](https://i.imgur.com/4APbcdz.png)\n*Recall and Precision curves*\n\n![](https://i.imgur.com/vBWzAfQ.png)\n*Precision vs. Recall for both models (model 1 on the left)*\n\nThat being said, I believe that the best split we could do for this dataset is 3-fold (1 for each video) and do a final ensemble using [WBF](https://github.com/ZFTurbo/Weighted-Boxes-Fusion) or some other ensembling technique.\n\nI made my notebooks public so feel free to use them:\n* YoloV5 dataset generator with correct train/test splits: [link](https://www.kaggle.com/coldfir3/efficient-yolov5-dataset-generator)\n* YoloV5 train: [link](https://www.kaggle.com/coldfir3/yolov5-train)\n* YoloV5 inference (WIP): link",
      "votes": null
    },
    {
      "id": "1608788",
      "postDate": "12/06/2021 16:45:30",
      "content": "<p>I think GroupKFold using 'sequence' column as 'group' parameter can be fair too (you'll end up with 3-4 video sequences for each fold, if using 5-fold split), I see the gaps between CV/LB but I rather blame it on difficulty of the videos in validation, my guess is that videos on public LB contain more hard samples than my current validation dataset</p>\n<p>We can treat any unique sequence as an individual video, I think so</p>",
      "rawMarkdown": "I think GroupKFold using 'sequence' column as 'group' parameter can be fair too (you'll end up with 3-4 video sequences for each fold, if using 5-fold split), I see the gaps between CV/LB but I rather blame it on difficulty of the videos in validation, my guess is that videos on public LB contain more hard samples than my current validation dataset\n\nWe can treat any unique sequence as an individual video, I think so",
      "votes": null
    },
    {
      "id": "1608816",
      "postDate": "12/06/2021 17:24:01",
      "content": "<p>That is a good insight, I will try using this split and update my post.</p>",
      "rawMarkdown": "That is a good insight, I will try using this split and update my post.",
      "votes": null
    },
    {
      "id": "1609077",
      "postDate": "12/06/2021 22:46:54",
      "content": "<p>I got some strange results when splitting the images using <code>GroupKFold</code>. I tested with 2 folds and in both cases, the model performed worse when compared to splitting the videos (the validation sizes are roughly the same at 20%). Looking forward to more insights on this.</p>\n<p><img src=\"https://i.imgur.com/M9ZuLs1.png\" alt=\"\"></p>",
      "rawMarkdown": "I got some strange results when splitting the images using `GroupKFold`. I tested with 2 folds and in both cases, the model performed worse when compared to splitting the videos (the validation sizes are roughly the same at 20%). Looking forward to more insights on this.\n\n![](https://i.imgur.com/M9ZuLs1.png)",
      "votes": null
    },
    {
      "id": "1610573",
      "postDate": "12/07/2021 10:24:45",
      "content": "<p>👍👍muito legal</p>",
      "rawMarkdown": "👍👍muito legal",
      "votes": null
    },
    {
      "id": "1611590",
      "postDate": "12/08/2021 04:53:28",
      "content": "<p>great job, I find that my CV is larger than my LB(larger around 0.3), I think also I have data problem.</p>",
      "rawMarkdown": "great job, I find that my CV is larger than my LB(larger around 0.3), I think also I have data problem.",
      "votes": null
    },
    {
      "id": "1611964",
      "postDate": "12/08/2021 12:32:52",
      "content": "<p>Worse or not, the matter is if it's leaky or not<br>\nDifferent validation splits may perform differently on this dataset (easy sample are too far apart from hard samples, as I can see from manually viewing each of the videos and corresponding GTs)</p>\n<p>But these split still should be distinguishable in terms of \"leaky/not leaky\" and from these plots I can say, that group fold is not a leaky strategy, as I thought :)</p>",
      "rawMarkdown": "Worse or not, the matter is if it's leaky or not\nDifferent validation splits may perform differently on this dataset (easy sample are too far apart from hard samples, as I can see from manually viewing each of the videos and corresponding GTs)\n\nBut these split still should be distinguishable in terms of \"leaky/not leaky\" and from these plots I can say, that group fold is not a leaky strategy, as I thought :)",
      "votes": null
    },
    {
      "id": "1612016",
      "postDate": "12/08/2021 13:17:07",
      "content": "<p>You are correct!</p>",
      "rawMarkdown": "You are correct!",
      "votes": null
    },
    {
      "id": "1612783",
      "postDate": "12/09/2021 10:12:51",
      "content": "<p>woah gotta learn a few thigs from you</p>",
      "rawMarkdown": "woah gotta learn a few thigs from you",
      "votes": null
    },
    {
      "id": "1612956",
      "postDate": "12/09/2021 13:50:12",
      "content": "<p>We can all learn with each others =)</p>",
      "rawMarkdown": "We can all learn with each others =)",
      "votes": null
    },
    {
      "id": "1612957",
      "postDate": "12/09/2021 13:50:31",
      "content": "<p>Tnx! let us know if you find out the source of the leak! =)</p>",
      "rawMarkdown": "Tnx! let us know if you find out the source of the leak! =)",
      "votes": null
    },
    {
      "id": "1613058",
      "postDate": "12/09/2021 15:00:58",
      "content": "<p>Thank you for warning!<br>\nI noticed that in video 2 there are only 7.9 percent of images with starfish on it against 25.5 and 32 in video 0 and 1. It may be not correct to validate on that data</p>",
      "rawMarkdown": "Thank you for warning!\nI noticed that in video 2 there are only 7.9 percent of images with starfish on it against 25.5 and 32 in video 0 and 1. It may be not correct to validate on that data",
      "votes": null
    },
    {
      "id": "1613101",
      "postDate": "12/09/2021 15:42:31",
      "content": "<p>I didn't noticed that. On that subject, I think the validation scores should be computed only using frames that has any annotations on it. I will try to investigate this further (video-wise 3-fold) and see if any substantial differences appear </p>",
      "rawMarkdown": "I didn't noticed that. On that subject, I think the validation scores should be computed only using frames that has any annotations on it. I will try to investigate this further (video-wise 3-fold) and see if any substantial differences appear",
      "votes": null
    },
    {
      "id": "1616388",
      "postDate": "12/13/2021 11:46:58",
      "content": "<p>Precission / Recall is better in random split ratio because there is leakage to val set. Almost same image in train and val set. Probability of close frames from same sequence :)</p>",
      "rawMarkdown": "Precission / Recall is better in random split ratio because there is leakage to val set. Almost same image in train and val set. Probability of close frames from same sequence :)",
      "votes": null
    },
    {
      "id": "1617259",
      "postDate": "12/14/2021 01:49:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a>, Do you use below split code ? </p>\n<p><code>kf = GroupKFold(n_splits = 5)</code><br>\n<code>df_train = df_train.reset_index(drop=True)</code><br>\n<code>df_train['fold'] = -1</code><br>\n<code>for fold, (train_idx, val_idx) in enumerate(kf.split(df_train, y = df_train.video_id.tolist(), groups=df_train.sequence)):</code><br>\n<code>df_train.loc[val_idx, 'fold'] = fold</code></p>",
      "rawMarkdown": "Hi @martynoveduard, Do you use below split code ? \n\n`kf = GroupKFold(n_splits = 5)`\n`df_train = df_train.reset_index(drop=True)`\n`df_train['fold'] = -1`\n`for fold, (train_idx, val_idx) in enumerate(kf.split(df_train, y = df_train.video_id.tolist(), groups=df_train.sequence)):`\n`df_train.loc[val_idx, 'fold'] = fold`",
      "votes": null
    },
    {
      "id": "1617967",
      "postDate": "12/14/2021 14:35:44",
      "content": "<p>I used the following</p>\n<pre><code>from sklearn.model_selection import GroupKFold\n\nkf = GroupKFold(n_splits = 5) \ndf['fold'] = -1\nfor fold, (train_idx, val_idx) in enumerate(kf.split(df, y = df.video_id.tolist(), groups=df.sequence)):\n    df.loc[val_idx, 'fold'] = fold\n</code></pre>",
      "rawMarkdown": "I used the following\n```\nfrom sklearn.model_selection import GroupKFold\n\nkf = GroupKFold(n_splits = 5) \ndf['fold'] = -1\nfor fold, (train_idx, val_idx) in enumerate(kf.split(df, y = df.video_id.tolist(), groups=df.sequence)):\n    df.loc[val_idx, 'fold'] = fold\n```",
      "votes": null
    },
    {
      "id": "1618377",
      "postDate": "12/14/2021 23:53:44",
      "content": "<p><a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a> Thanks for reply.</p>",
      "rawMarkdown": "coldfir3 Thanks for reply.",
      "votes": null
    },
    {
      "id": "1618576",
      "postDate": "12/15/2021 05:30:28",
      "content": "<p>Oh. Thanks for your suggestion. :)</p>",
      "rawMarkdown": "Oh. Thanks for your suggestion. :)",
      "votes": null
    },
    {
      "id": "1620302",
      "postDate": "12/16/2021 16:34:28",
      "content": "<p>Thanks for the insight. It greatly helped me!</p>",
      "rawMarkdown": "Thanks for the insight. It greatly helped me!",
      "votes": null
    },
    {
      "id": "1620403",
      "postDate": "12/16/2021 18:00:49",
      "content": "<p>Glad it was helpful =)</p>",
      "rawMarkdown": "Glad it was helpful =)",
      "votes": null
    },
    {
      "id": "1620490",
      "postDate": "12/16/2021 20:42:17",
      "content": "<p>Data into each video is not IID because there is temporal autocorrelation between images. Anyway, even in the same video, when images are far from each other in time they are neglibly correlated. Split each video into enough large chunks, consider each chuck as a group and do GroupKFold cross validation.  </p>",
      "rawMarkdown": "Data into each video is not IID because there is temporal autocorrelation between images. Anyway, even in the same video, when images are far from each other in time they are neglibly correlated. Split each video into enough large chunks, consider each chuck as a group and do GroupKFold cross validation.",
      "votes": null
    },
    {
      "id": "1622669",
      "postDate": "12/19/2021 03:07:55",
      "content": "<p>Thanks for sharing.<br>\nIn order to better understand you analysis, could you share what is the IoU threshold used to measure precision &amp; recall?</p>",
      "rawMarkdown": "Thanks for sharing.\nIn order to better understand you analysis, could you share what is the IoU threshold used to measure precision & recall?",
      "votes": null
    },
    {
      "id": "1622670",
      "postDate": "12/19/2021 03:10:44",
      "content": "<p>Iou = 0.5   </p>",
      "rawMarkdown": "Iou = 0.5",
      "votes": null
    },
    {
      "id": "1623526",
      "postDate": "12/20/2021 01:56:58",
      "content": "<p>What about stratifiedKfold?</p>",
      "rawMarkdown": "What about stratifiedKfold?",
      "votes": null
    },
    {
      "id": "1623549",
      "postDate": "12/20/2021 02:18:22",
      "content": "<p>That would depend on how you would stratify the data <a href=\"https://www.kaggle.com/saberdokmak\" target=\"_blank\">@saberdokmak</a> </p>\n<p>They key idea is that we should not leak information from the test set to the validation set</p>",
      "rawMarkdown": "That would depend on how you would stratify the data @saberdokmak \n\nThey key idea is that we should not leak information from the test set to the validation set",
      "votes": null
    },
    {
      "id": "1625311",
      "postDate": "12/21/2021 17:00:55",
      "content": "<p>Hey! I've prepared a notebook, with the splitting of .csv files by video or length. Maybe it could help somebody)<br>\n<a href=\"https://www.kaggle.com/vadbeg/yolox-dataset\" target=\"_blank\">https://www.kaggle.com/vadbeg/yolox-dataset</a></p>",
      "rawMarkdown": "Hey! I've prepared a notebook, with the splitting of .csv files by video or length. Maybe it could help somebody)\nhttps://www.kaggle.com/vadbeg/yolox-dataset",
      "votes": null
    },
    {
      "id": "1626677",
      "postDate": "12/23/2021 05:23:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a> , great insights! I tried your video split method. I train on Video image of 0 and 1, and val on 2 with labels or all 2 images (also tried val on 0 with labels). CV score is quit high, 0.65+. But LB score is only about 0.426. Still struggle in this situation and don't know why.</p>\n<p>If any insights or tips, it would be nice to discuss.</p>",
      "rawMarkdown": "Hi @coldfir3 , great insights! I tried your video split method. I train on Video image of 0 and 1, and val on 2 with labels or all 2 images (also tried val on 0 with labels). CV score is quit high, 0.65+. But LB score is only about 0.426. Still struggle in this situation and don't know why.\n\nIf any insights or tips, it would be nice to discuss.",
      "votes": null
    },
    {
      "id": "1626729",
      "postDate": "12/23/2021 06:53:00",
      "content": "<p>Have you validated using unlabelled data as well?</p>",
      "rawMarkdown": "Have you validated using unlabelled data as well?",
      "votes": null
    },
    {
      "id": "1626774",
      "postDate": "12/23/2021 07:59:54",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> During training, I use a few unlabeled video 2 images to choose best weights, not all unlabeled video 2. The total video 0+1+2 number is 6000. After training, I compute a F2 score on all video 2 images, F2 score is still high. Therefore, I think whether to validate on unlabeled video 2 during training is not important. </p>\n<p>Kindly correct me if I am wrong.</p>",
      "rawMarkdown": "remekkinas During training, I use a few unlabeled video 2 images to choose best weights, not all unlabeled video 2. The total video 0+1+2 number is 6000. After training, I compute a F2 score on all video 2 images, F2 score is still high. Therefore, I think whether to validate on unlabeled video 2 during training is not important. \n\nKindly correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "1626790",
      "postDate": "12/23/2021 08:07:35",
      "content": "<p>We are almost in the same point now 😄 Last two days we as a team were thinking about this. I think that TOP LB guys shoud help us … but they are focused on competition and winning (it is ok certainly). As I came here to learn and share … we came to the conclusion that - we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score. </p>",
      "rawMarkdown": "We are almost in the same point now 😄 Last two days we as a team were thinking about this. I think that TOP LB guys shoud help us ... but they are focused on competition and winning (it is ok certainly). As I came here to learn and share ... we came to the conclusion that - we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score.",
      "votes": null
    },
    {
      "id": "1626854",
      "postDate": "12/23/2021 09:03:19",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I think you are doing great. At least a small gap between F2 and LB.</p>\n<p>For me, if conservative GroupFold, the F2 LB gap is really small. No one to blame, just because my model is bad. Haha…   Try more image aug and maybe try YOLOv5. </p>",
      "rawMarkdown": "remekkinas I think you are doing great. At least a small gap between F2 and LB.\n\nFor me, if conservative GroupFold, the F2 LB gap is really small. No one to blame, just because my model is bad. Haha...   Try more image aug and maybe try YOLOv5.",
      "votes": null
    },
    {
      "id": "1626860",
      "postDate": "12/23/2021 09:14:56",
      "content": "<p>We are using YoloX, Yolov5 and FasterRCNN. </p>\n<ul>\n<li>YoloX - without problem 0.53-0.536 … and maybe a little bit more … but not so much … so I do not submit more … (We are looking for better model to jump). I know YoloX very very well now …. because I had to introduce many custom changes - custom augumentations (based on Albumentations), custom metric, custom checkpoint save … I can see some improvements which could be introduce to YoloX … and I will contribute to YoloX</li>\n<li>Yolo5 … without success … probably today I jump over 0.5 (probably … because I am testing it locally using f2 score script) … but I am afraid still below YoloX … I asked many people about tips but no help … I understand … this is competition :) … I do not know how many experiments (augumentations, traning parameter change, dataset split - 3 ways) we have done so far … but still looking for one … tip … (I know that there is something easy … which we can not see).</li>\n<li>FasterRCNN similar to Yolov5 </li>\n</ul>\n<p>When I find reasonable model (about 0.57-0.58) I will introduce next tools … TTA/WBF/Tracking etc. Now I am looking for one tip …. Boost Yolo5 ….</p>",
      "rawMarkdown": "We are using YoloX, Yolov5 and FasterRCNN. \n- YoloX - without problem 0.53-0.536 ... and maybe a little bit more ... but not so much ... so I do not submit more ... (We are looking for better model to jump). I know YoloX very very well now .... because I had to introduce many custom changes - custom augumentations (based on Albumentations), custom metric, custom checkpoint save ... I can see some improvements which could be introduce to YoloX ... and I will contribute to YoloX\n- Yolo5 ... without success ... probably today I jump over 0.5 (probably ... because I am testing it locally using f2 score script) ... but I am afraid still below YoloX ... I asked many people about tips but no help ... I understand ... this is competition :) ... I do not know how many experiments (augumentations, traning parameter change, dataset split - 3 ways) we have done so far ... but still looking for one ... tip ... (I know that there is something easy ... which we can not see).\n- FasterRCNN similar to Yolov5 \n\nWhen I find reasonable model (about 0.57-0.58) I will introduce next tools ... TTA/WBF/Tracking etc. Now I am looking for one tip .... Boost Yolo5 ....",
      "votes": null
    },
    {
      "id": "1626873",
      "postDate": "12/23/2021 09:25:48",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> Hi, definately you guys are great! Competition is like this. Just all in, and believe good thing will come. If there is a tip, maybe merge with some solo guy. You are on right track. Focus on single model improvement. Ensemble is last three week task.</p>\n<p>This is my first Object Detection competition. Need to learn from you, haha, since you are so familiar with YOLO family now.</p>",
      "rawMarkdown": "remekkinas Hi, definately you guys are great! Competition is like this. Just all in, and believe good thing will come. If there is a tip, maybe merge with some solo guy. You are on right track. Focus on single model improvement. Ensemble is last three week task.\n\nThis is my first Object Detection competition. Need to learn from you, haha, since you are so familiar with YOLO family now.",
      "votes": null
    },
    {
      "id": "1626892",
      "postDate": "12/23/2021 09:53:36",
      "content": "<p>Thank you! We will be in touch :)</p>",
      "rawMarkdown": "Thank you! We will be in touch :)",
      "votes": null
    },
    {
      "id": "1636875",
      "postDate": "01/03/2022 10:54:36",
      "content": "<p><a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a>  how is correlation. </p>",
      "rawMarkdown": "coldfir3  how is correlation.",
      "votes": null
    },
    {
      "id": "1639421",
      "postDate": "01/05/2022 16:09:53",
      "content": "<p>Can you please share which script you use for computing F2 score on the validation data? I used one available on a discussion  but it doesn't seem accurate</p>",
      "rawMarkdown": "Can you please share which script you use for computing F2 score on the validation data? I used one available on a discussion  but it doesn't seem accurate",
      "votes": null
    },
    {
      "id": "1645539",
      "postDate": "01/11/2022 02:08:58",
      "content": "<p>Yes! I got it.</p>",
      "rawMarkdown": "Yes! I got it.",
      "votes": null
    },
    {
      "id": "1657978",
      "postDate": "01/20/2022 16:07:07",
      "content": "<p>Good idea! Thanks!</p>",
      "rawMarkdown": "Good idea! Thanks!",
      "votes": null
    },
    {
      "id": "1679453",
      "postDate": "02/07/2022 08:41:54",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a> Adriano, thanks for your reminder, I also tried split via video-id or subsequences, hit no luck though. Did you make some progress in the splitting strategy?</p>",
      "rawMarkdown": "Hey @coldfir3 Adriano, thanks for your reminder, I also tried split via video-id or subsequences, hit no luck though. Did you make some progress in the splitting strategy?",
      "votes": null
    },
    {
      "id": "1679482",
      "postDate": "02/07/2022 08:50:18",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> Remek, could you give me a tip on how to define the data.yaml s.t. </p>\n<blockquote>\n  <p>we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score</p>\n</blockquote>\n<p>Did you just pick randomly 1%~3% of empty frames from each video for training validation?<br>\nThanks!</p>",
      "rawMarkdown": "Hey @remekkinas Remek, could you give me a tip on how to define the data.yaml s.t. \n\n> we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score\n\nDid you just pick randomly 1%~3% of empty frames from each video for training validation?\nThanks!",
      "votes": null
    },
    {
      "id": "1679490",
      "postDate": "02/07/2022 08:53:19",
      "content": "<p>Yes. I just use sample(frac=0.01) as I remember. </p>",
      "rawMarkdown": "Yes. I just use sample(frac=0.01) as I remember.",
      "votes": null
    },
    {
      "id": "1681490",
      "postDate": "02/08/2022 13:58:05",
      "content": "<p>Thanks for your clarification!</p>",
      "rawMarkdown": "Thanks for your clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1608788,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "12/06/2021 16:45:30",
      "content": "<p>I think GroupKFold using 'sequence' column as 'group' parameter can be fair too (you'll end up with 3-4 video sequences for each fold, if using 5-fold split), I see the gaps between CV/LB but I rather blame it on difficulty of the videos in validation, my guess is that videos on public LB contain more hard samples than my current validation dataset</p>\n<p>We can treat any unique sequence as an individual video, I think so</p>",
      "votes": null,
      "replies": [
        {
          "id": 1608816,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/06/2021 17:24:01",
          "content": "<p>That is a good insight, I will try using this split and update my post.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1609077,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/06/2021 22:46:54",
          "content": "<p>I got some strange results when splitting the images using <code>GroupKFold</code>. I tested with 2 folds and in both cases, the model performed worse when compared to splitting the videos (the validation sizes are roughly the same at 20%). Looking forward to more insights on this.</p>\n<p><img src=\"https://i.imgur.com/M9ZuLs1.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1611964,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "12/08/2021 12:32:52",
          "content": "<p>Worse or not, the matter is if it's leaky or not<br>\nDifferent validation splits may perform differently on this dataset (easy sample are too far apart from hard samples, as I can see from manually viewing each of the videos and corresponding GTs)</p>\n<p>But these split still should be distinguishable in terms of \"leaky/not leaky\" and from these plots I can say, that group fold is not a leaky strategy, as I thought :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1612016,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/08/2021 13:17:07",
          "content": "<p>You are correct!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1617259,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "12/14/2021 01:49:59",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a>, Do you use below split code ? </p>\n<p><code>kf = GroupKFold(n_splits = 5)</code><br>\n<code>df_train = df_train.reset_index(drop=True)</code><br>\n<code>df_train['fold'] = -1</code><br>\n<code>for fold, (train_idx, val_idx) in enumerate(kf.split(df_train, y = df_train.video_id.tolist(), groups=df_train.sequence)):</code><br>\n<code>df_train.loc[val_idx, 'fold'] = fold</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1617967,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/14/2021 14:35:44",
          "content": "<p>I used the following</p>\n<pre><code>from sklearn.model_selection import GroupKFold\n\nkf = GroupKFold(n_splits = 5) \ndf['fold'] = -1\nfor fold, (train_idx, val_idx) in enumerate(kf.split(df, y = df.video_id.tolist(), groups=df.sequence)):\n    df.loc[val_idx, 'fold'] = fold\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1618377,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "12/14/2021 23:53:44",
          "content": "<p><a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a> Thanks for reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623526,
          "author_name": "saberdokmak",
          "author_url": "",
          "post_date": "12/20/2021 01:56:58",
          "content": "<p>What about stratifiedKfold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623549,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/20/2021 02:18:22",
          "content": "<p>That would depend on how you would stratify the data <a href=\"https://www.kaggle.com/saberdokmak\" target=\"_blank\">@saberdokmak</a> </p>\n<p>They key idea is that we should not leak information from the test set to the validation set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1636875,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/03/2022 10:54:36",
          "content": "<p><a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a>  how is correlation. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1610573,
      "author_name": "orianapalmacalabokis",
      "author_url": "",
      "post_date": "12/07/2021 10:24:45",
      "content": "<p>👍👍muito legal</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1611590,
      "author_name": "yuanzilong",
      "author_url": "",
      "post_date": "12/08/2021 04:53:28",
      "content": "<p>great job, I find that my CV is larger than my LB(larger around 0.3), I think also I have data problem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1612957,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/09/2021 13:50:31",
          "content": "<p>Tnx! let us know if you find out the source of the leak! =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1612783,
      "author_name": "manjunathgb",
      "author_url": "",
      "post_date": "12/09/2021 10:12:51",
      "content": "<p>woah gotta learn a few thigs from you</p>",
      "votes": null,
      "replies": [
        {
          "id": 1612956,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/09/2021 13:50:12",
          "content": "<p>We can all learn with each others =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1613058,
      "author_name": "alexandryu",
      "author_url": "",
      "post_date": "12/09/2021 15:00:58",
      "content": "<p>Thank you for warning!<br>\nI noticed that in video 2 there are only 7.9 percent of images with starfish on it against 25.5 and 32 in video 0 and 1. It may be not correct to validate on that data</p>",
      "votes": null,
      "replies": [
        {
          "id": 1613101,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/09/2021 15:42:31",
          "content": "<p>I didn't noticed that. On that subject, I think the validation scores should be computed only using frames that has any annotations on it. I will try to investigate this further (video-wise 3-fold) and see if any substantial differences appear </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1616388,
      "author_name": "lukaszborecki",
      "author_url": "",
      "post_date": "12/13/2021 11:46:58",
      "content": "<p>Precission / Recall is better in random split ratio because there is leakage to val set. Almost same image in train and val set. Probability of close frames from same sequence :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1618576,
      "author_name": "hsulet",
      "author_url": "",
      "post_date": "12/15/2021 05:30:28",
      "content": "<p>Oh. Thanks for your suggestion. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1620302,
      "author_name": "pritampaul360",
      "author_url": "",
      "post_date": "12/16/2021 16:34:28",
      "content": "<p>Thanks for the insight. It greatly helped me!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1620403,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/16/2021 18:00:49",
          "content": "<p>Glad it was helpful =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1620490,
      "author_name": "lucamassaron",
      "author_url": "",
      "post_date": "12/16/2021 20:42:17",
      "content": "<p>Data into each video is not IID because there is temporal autocorrelation between images. Anyway, even in the same video, when images are far from each other in time they are neglibly correlated. Split each video into enough large chunks, consider each chuck as a group and do GroupKFold cross validation.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1622669,
      "author_name": "diegoalejogm",
      "author_url": "",
      "post_date": "12/19/2021 03:07:55",
      "content": "<p>Thanks for sharing.<br>\nIn order to better understand you analysis, could you share what is the IoU threshold used to measure precision &amp; recall?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1622670,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "12/19/2021 03:10:44",
          "content": "<p>Iou = 0.5   </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1625311,
      "author_name": "vadbeg",
      "author_url": "",
      "post_date": "12/21/2021 17:00:55",
      "content": "<p>Hey! I've prepared a notebook, with the splitting of .csv files by video or length. Maybe it could help somebody)<br>\n<a href=\"https://www.kaggle.com/vadbeg/yolox-dataset\" target=\"_blank\">https://www.kaggle.com/vadbeg/yolox-dataset</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1626677,
      "author_name": "xiaojiu1414",
      "author_url": "",
      "post_date": "12/23/2021 05:23:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a> , great insights! I tried your video split method. I train on Video image of 0 and 1, and val on 2 with labels or all 2 images (also tried val on 0 with labels). CV score is quit high, 0.65+. But LB score is only about 0.426. Still struggle in this situation and don't know why.</p>\n<p>If any insights or tips, it would be nice to discuss.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1626729,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/23/2021 06:53:00",
          "content": "<p>Have you validated using unlabelled data as well?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626774,
          "author_name": "xiaojiu1414",
          "author_url": "",
          "post_date": "12/23/2021 07:59:54",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> During training, I use a few unlabeled video 2 images to choose best weights, not all unlabeled video 2. The total video 0+1+2 number is 6000. After training, I compute a F2 score on all video 2 images, F2 score is still high. Therefore, I think whether to validate on unlabeled video 2 during training is not important. </p>\n<p>Kindly correct me if I am wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626790,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/23/2021 08:07:35",
          "content": "<p>We are almost in the same point now 😄 Last two days we as a team were thinking about this. I think that TOP LB guys shoud help us … but they are focused on competition and winning (it is ok certainly). As I came here to learn and share … we came to the conclusion that - we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626854,
          "author_name": "xiaojiu1414",
          "author_url": "",
          "post_date": "12/23/2021 09:03:19",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I think you are doing great. At least a small gap between F2 and LB.</p>\n<p>For me, if conservative GroupFold, the F2 LB gap is really small. No one to blame, just because my model is bad. Haha…   Try more image aug and maybe try YOLOv5. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626860,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/23/2021 09:14:56",
          "content": "<p>We are using YoloX, Yolov5 and FasterRCNN. </p>\n<ul>\n<li>YoloX - without problem 0.53-0.536 … and maybe a little bit more … but not so much … so I do not submit more … (We are looking for better model to jump). I know YoloX very very well now …. because I had to introduce many custom changes - custom augumentations (based on Albumentations), custom metric, custom checkpoint save … I can see some improvements which could be introduce to YoloX … and I will contribute to YoloX</li>\n<li>Yolo5 … without success … probably today I jump over 0.5 (probably … because I am testing it locally using f2 score script) … but I am afraid still below YoloX … I asked many people about tips but no help … I understand … this is competition :) … I do not know how many experiments (augumentations, traning parameter change, dataset split - 3 ways) we have done so far … but still looking for one … tip … (I know that there is something easy … which we can not see).</li>\n<li>FasterRCNN similar to Yolov5 </li>\n</ul>\n<p>When I find reasonable model (about 0.57-0.58) I will introduce next tools … TTA/WBF/Tracking etc. Now I am looking for one tip …. Boost Yolo5 ….</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626873,
          "author_name": "xiaojiu1414",
          "author_url": "",
          "post_date": "12/23/2021 09:25:48",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> Hi, definately you guys are great! Competition is like this. Just all in, and believe good thing will come. If there is a tip, maybe merge with some solo guy. You are on right track. Focus on single model improvement. Ensemble is last three week task.</p>\n<p>This is my first Object Detection competition. Need to learn from you, haha, since you are so familiar with YOLO family now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626892,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/23/2021 09:53:36",
          "content": "<p>Thank you! We will be in touch :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1639421,
          "author_name": "eliaik",
          "author_url": "",
          "post_date": "01/05/2022 16:09:53",
          "content": "<p>Can you please share which script you use for computing F2 score on the validation data? I used one available on a discussion  but it doesn't seem accurate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1645539,
          "author_name": "tiandaye",
          "author_url": "",
          "post_date": "01/11/2022 02:08:58",
          "content": "<p>Yes! I got it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1679482,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "02/07/2022 08:50:18",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> Remek, could you give me a tip on how to define the data.yaml s.t. </p>\n<blockquote>\n  <p>we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score</p>\n</blockquote>\n<p>Did you just pick randomly 1%~3% of empty frames from each video for training validation?<br>\nThanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1679490,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/07/2022 08:53:19",
          "content": "<p>Yes. I just use sample(frac=0.01) as I remember. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1681490,
          "author_name": "fuckvenkatraman",
          "author_url": "",
          "post_date": "02/08/2022 13:58:05",
          "content": "<p>Thanks for your clarification!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1657978,
      "author_name": "wuhaowang",
      "author_url": "",
      "post_date": "01/20/2022 16:07:07",
      "content": "<p>Good idea! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1679453,
      "author_name": "fuckvenkatraman",
      "author_url": "",
      "post_date": "02/07/2022 08:41:54",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/coldfir3\" target=\"_blank\">@coldfir3</a> Adriano, thanks for your reminder, I also tried split via video-id or subsequences, hit no luck though. Did you make some progress in the splitting strategy?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1608778": "From what I could tell, many of the public notebooks split the train/valid datasets at random. Keep in mind that video data have stupidly high temporal coherence and by splitting the dataset randomly you will end up with very similar images in both sets.\n\nConsider the following example:\n![](https://i.imgur.com/C5eB8bJ.png)\n*Two consecutive frames of video 2*\n\nIf you end up having the left frame on the train set and the right one in the validation set it is very likely that the model will correctly predict those detections. However, when running inference, the model will be shown to a different video and it is unlikely to be as good.\n\nI trained two models to demonstrate this effect. \n1. Randomly splitting the dataset.\n2. Using videos 0 and 1 for training and video 2 for validation.\n\nIn both cases, I used 6k images (~5k with labels and 1k~ with background only).\n\nAlthough both models trained roughly the same (aka. had the same training loss curve) their validation curves were significantly different.\n\n![](https://i.imgur.com/mKOPV8t.png)\n*Training losses*\n![](https://i.imgur.com/E9IZSc7.png)\n*Validation losses*\n\nThe decrease in performance is easily explained by a substantial drop in **recall**\n![](https://i.imgur.com/4APbcdz.png)\n*Recall and Precision curves*\n\n![](https://i.imgur.com/vBWzAfQ.png)\n*Precision vs. Recall for both models (model 1 on the left)*\n\nThat being said, I believe that the best split we could do for this dataset is 3-fold (1 for each video) and do a final ensemble using [WBF](https://github.com/ZFTurbo/Weighted-Boxes-Fusion) or some other ensembling technique.\n\nI made my notebooks public so feel free to use them:\n* YoloV5 dataset generator with correct train/test splits: [link](https://www.kaggle.com/coldfir3/efficient-yolov5-dataset-generator)\n* YoloV5 train: [link](https://www.kaggle.com/coldfir3/yolov5-train)\n* YoloV5 inference (WIP): link",
    "1608788": "I think GroupKFold using 'sequence' column as 'group' parameter can be fair too (you'll end up with 3-4 video sequences for each fold, if using 5-fold split), I see the gaps between CV/LB but I rather blame it on difficulty of the videos in validation, my guess is that videos on public LB contain more hard samples than my current validation dataset\n\nWe can treat any unique sequence as an individual video, I think so",
    "1608816": "That is a good insight, I will try using this split and update my post.",
    "1609077": "I got some strange results when splitting the images using `GroupKFold`. I tested with 2 folds and in both cases, the model performed worse when compared to splitting the videos (the validation sizes are roughly the same at 20%). Looking forward to more insights on this.\n\n![](https://i.imgur.com/M9ZuLs1.png)",
    "1610573": "👍👍muito legal",
    "1611590": "great job, I find that my CV is larger than my LB(larger around 0.3), I think also I have data problem.",
    "1611964": "Worse or not, the matter is if it's leaky or not\nDifferent validation splits may perform differently on this dataset (easy sample are too far apart from hard samples, as I can see from manually viewing each of the videos and corresponding GTs)\n\nBut these split still should be distinguishable in terms of \"leaky/not leaky\" and from these plots I can say, that group fold is not a leaky strategy, as I thought :)",
    "1612016": "You are correct!",
    "1612783": "woah gotta learn a few thigs from you",
    "1612956": "We can all learn with each others =)",
    "1612957": "Tnx! let us know if you find out the source of the leak! =)",
    "1613058": "Thank you for warning!\nI noticed that in video 2 there are only 7.9 percent of images with starfish on it against 25.5 and 32 in video 0 and 1. It may be not correct to validate on that data",
    "1613101": "I didn't noticed that. On that subject, I think the validation scores should be computed only using frames that has any annotations on it. I will try to investigate this further (video-wise 3-fold) and see if any substantial differences appear",
    "1616388": "Precission / Recall is better in random split ratio because there is leakage to val set. Almost same image in train and val set. Probability of close frames from same sequence :)",
    "1617259": "Hi @martynoveduard, Do you use below split code ? \n\n`kf = GroupKFold(n_splits = 5)`\n`df_train = df_train.reset_index(drop=True)`\n`df_train['fold'] = -1`\n`for fold, (train_idx, val_idx) in enumerate(kf.split(df_train, y = df_train.video_id.tolist(), groups=df_train.sequence)):`\n`df_train.loc[val_idx, 'fold'] = fold`",
    "1617967": "I used the following\n```\nfrom sklearn.model_selection import GroupKFold\n\nkf = GroupKFold(n_splits = 5) \ndf['fold'] = -1\nfor fold, (train_idx, val_idx) in enumerate(kf.split(df, y = df.video_id.tolist(), groups=df.sequence)):\n    df.loc[val_idx, 'fold'] = fold\n```",
    "1618377": "coldfir3 Thanks for reply.",
    "1618576": "Oh. Thanks for your suggestion. :)",
    "1620302": "Thanks for the insight. It greatly helped me!",
    "1620403": "Glad it was helpful =)",
    "1620490": "Data into each video is not IID because there is temporal autocorrelation between images. Anyway, even in the same video, when images are far from each other in time they are neglibly correlated. Split each video into enough large chunks, consider each chuck as a group and do GroupKFold cross validation.",
    "1622669": "Thanks for sharing.\nIn order to better understand you analysis, could you share what is the IoU threshold used to measure precision & recall?",
    "1622670": "Iou = 0.5",
    "1623526": "What about stratifiedKfold?",
    "1623549": "That would depend on how you would stratify the data @saberdokmak \n\nThey key idea is that we should not leak information from the test set to the validation set",
    "1625311": "Hey! I've prepared a notebook, with the splitting of .csv files by video or length. Maybe it could help somebody)\nhttps://www.kaggle.com/vadbeg/yolox-dataset",
    "1626677": "Hi @coldfir3 , great insights! I tried your video split method. I train on Video image of 0 and 1, and val on 2 with labels or all 2 images (also tried val on 0 with labels). CV score is quit high, 0.65+. But LB score is only about 0.426. Still struggle in this situation and don't know why.\n\nIf any insights or tips, it would be nice to discuss.",
    "1626729": "Have you validated using unlabelled data as well?",
    "1626774": "remekkinas During training, I use a few unlabeled video 2 images to choose best weights, not all unlabeled video 2. The total video 0+1+2 number is 6000. After training, I compute a F2 score on all video 2 images, F2 score is still high. Therefore, I think whether to validate on unlabeled video 2 during training is not important. \n\nKindly correct me if I am wrong.",
    "1626790": "We are almost in the same point now 😄 Last two days we as a team were thinking about this. I think that TOP LB guys shoud help us ... but they are focused on competition and winning (it is ok certainly). As I came here to learn and share ... we came to the conclusion that - we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score.",
    "1626854": "remekkinas I think you are doing great. At least a small gap between F2 and LB.\n\nFor me, if conservative GroupFold, the F2 LB gap is really small. No one to blame, just because my model is bad. Haha...   Try more image aug and maybe try YOLOv5.",
    "1626860": "We are using YoloX, Yolov5 and FasterRCNN. \n- YoloX - without problem 0.53-0.536 ... and maybe a little bit more ... but not so much ... so I do not submit more ... (We are looking for better model to jump). I know YoloX very very well now .... because I had to introduce many custom changes - custom augumentations (based on Albumentations), custom metric, custom checkpoint save ... I can see some improvements which could be introduce to YoloX ... and I will contribute to YoloX\n- Yolo5 ... without success ... probably today I jump over 0.5 (probably ... because I am testing it locally using f2 score script) ... but I am afraid still below YoloX ... I asked many people about tips but no help ... I understand ... this is competition :) ... I do not know how many experiments (augumentations, traning parameter change, dataset split - 3 ways) we have done so far ... but still looking for one ... tip ... (I know that there is something easy ... which we can not see).\n- FasterRCNN similar to Yolov5 \n\nWhen I find reasonable model (about 0.57-0.58) I will introduce next tools ... TTA/WBF/Tracking etc. Now I am looking for one tip .... Boost Yolo5 ....",
    "1626873": "remekkinas Hi, definately you guys are great! Competition is like this. Just all in, and believe good thing will come. If there is a tip, maybe merge with some solo guy. You are on right track. Focus on single model improvement. Ensemble is last three week task.\n\nThis is my first Object Detection competition. Need to learn from you, haha, since you are so familiar with YOLO family now.",
    "1626892": "Thank you! We will be in touch :)",
    "1636875": "coldfir3  how is correlation.",
    "1639421": "Can you please share which script you use for computing F2 score on the validation data? I used one available on a discussion  but it doesn't seem accurate",
    "1645539": "Yes! I got it.",
    "1657978": "Good idea! Thanks!",
    "1679453": "Hey @coldfir3 Adriano, thanks for your reminder, I also tried split via video-id or subsequences, hit no luck though. Did you make some progress in the splitting strategy?",
    "1679482": "Hey @remekkinas Remek, could you give me a tip on how to define the data.yaml s.t. \n\n> we do validation training on labeled data and some 1%-3% background images (without labels). F2 score we do with all images in fold. As we can see we achieve this way F2 very close to LB score\n\nDid you just pick randomly 1%~3% of empty frames from each video for training validation?\nThanks!",
    "1679490": "Yes. I just use sample(frac=0.01) as I remember.",
    "1681490": "Thanks for your clarification!"
  },
  "source": "meta"
}