{
  "id": 308656,
  "title": "12th place solution (mysteriously disqualified) - Single ConvNeXt CascadeRCNN",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/308656",
  "author_name": "",
  "post_date": "2022-02-19T16:38:38.847881800Z",
  "votes": 35,
  "comment_count": 15,
  "views": 0,
  "content": "<h2>Solution Overview</h2>\n<p>Overall I did not make a lot of experiments as I spent just one week for this challenge. I was demotivated by public LB scores, stopped submitting and decided to wait a few weeks for the shakeup.</p>\n<p>With my experiments using MMdetection the best performing model was ConvNeXt base CascadeRCNN.</p>\n<p>Validation by video splits:</p>\n<ul>\n<li>video 0 - 0.71</li>\n<li>video 1 - 0.67</li>\n<li>video 2 - 0.8 (used this fold for submissions)</li>\n</ul>\n<p><strong>0.632</strong>  public LB / <strong>0.718</strong> private LB</p>\n<h3>Training</h3>\n<p>During training the following modifications were made to standard mmdet training pipeline:</p>\n<p><strong>Random Sized Crops around bboxes</strong><br>\nAs boxes were not dense using simple random crop during training was not optimal and I used random sized crop around bounding box if image is not empty. The crops covered multiple scales (1x-4x)  and also allowed to use quite a big batch size during training.</p>\n<p><strong>Balanced sampling from empty/nonempty images</strong><br>\nTo reduce FP predictions I used all images during training with probability to get a non empty/empty crop 0.6/0.4 respectively.</p>\n<p>Compared to full image training this approach improved CV by 3% and slightly improved LB scores. </p>\n<h3>Inference</h3>\n<p>Because random sized crops covered 1x-4x scale range I could safely use high resolution. Originally the images had  3840x2160 scales and were downscaled to 1280x720 for competition purposes.<br>\nIt was quite safe to assume that 3840x2160 resolution is the best for inference. It also provided the best results on CV/Public/Private splits</p>\n<h3>Postprocessing</h3>\n<p>Removed false positive predictions (boosted CV by 1-2%, LB by 0.3%) by leaving only big clusters of detections</p>\n<pre><code>df = pd.read_csv(\"submission.csv\")\nvals = df.annotations.values\npos = np.array([isinstance(v, str) and len(v) &gt; 2 for v in vals])\npos_dilated = binary_dilation(pos, iterations=1)\npos = remove_small_objects(measure.label(pos_dilated), min_size=10)\nvals[pos == 0] = np.nan\ndf[\"annotations\"] = vals\ndf.to_csv(\"submission.csv\", index=False)\n</code></pre>\n<p>The model is quite heavy so no TTA or ensembling, just simple thresholding based on validation.</p>\n<p>I did not make a lot of submissions and it was a no-brainer to choose the ones with high CV and public LB</p>\n<h2>Disqualification</h2>\n<p>Was really disappointed to see that I was removed from the leaderboard.<br>\nNo notification, no reasons, just removed. </p>\n<p>Input:</p>\n<ul>\n<li>not using anything from public kernels</li>\n<li>competing solo</li>\n<li>did not create fake accounts. That would be really a blatant action for an experienced kaggler.</li>\n<li>did not share anything with other competitors</li>\n</ul>\n<p>Output:</p>\n<ul>\n<li><strong>cheater</strong>! </li>\n</ul>\n<p>Waiting for the answer from the compliance team and <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
  "messages": [
    {
      "id": "1697472",
      "postDate": "02/19/2022 16:38:38",
      "content": "<h2>Solution Overview</h2>\n<p>Overall I did not make a lot of experiments as I spent just one week for this challenge. I was demotivated by public LB scores, stopped submitting and decided to wait a few weeks for the shakeup.</p>\n<p>With my experiments using MMdetection the best performing model was ConvNeXt base CascadeRCNN.</p>\n<p>Validation by video splits:</p>\n<ul>\n<li>video 0 - 0.71</li>\n<li>video 1 - 0.67</li>\n<li>video 2 - 0.8 (used this fold for submissions)</li>\n</ul>\n<p><strong>0.632</strong>  public LB / <strong>0.718</strong> private LB</p>\n<h3>Training</h3>\n<p>During training the following modifications were made to standard mmdet training pipeline:</p>\n<p><strong>Random Sized Crops around bboxes</strong><br>\nAs boxes were not dense using simple random crop during training was not optimal and I used random sized crop around bounding box if image is not empty. The crops covered multiple scales (1x-4x)  and also allowed to use quite a big batch size during training.</p>\n<p><strong>Balanced sampling from empty/nonempty images</strong><br>\nTo reduce FP predictions I used all images during training with probability to get a non empty/empty crop 0.6/0.4 respectively.</p>\n<p>Compared to full image training this approach improved CV by 3% and slightly improved LB scores. </p>\n<h3>Inference</h3>\n<p>Because random sized crops covered 1x-4x scale range I could safely use high resolution. Originally the images had  3840x2160 scales and were downscaled to 1280x720 for competition purposes.<br>\nIt was quite safe to assume that 3840x2160 resolution is the best for inference. It also provided the best results on CV/Public/Private splits</p>\n<h3>Postprocessing</h3>\n<p>Removed false positive predictions (boosted CV by 1-2%, LB by 0.3%) by leaving only big clusters of detections</p>\n<pre><code>df = pd.read_csv(\"submission.csv\")\nvals = df.annotations.values\npos = np.array([isinstance(v, str) and len(v) &gt; 2 for v in vals])\npos_dilated = binary_dilation(pos, iterations=1)\npos = remove_small_objects(measure.label(pos_dilated), min_size=10)\nvals[pos == 0] = np.nan\ndf[\"annotations\"] = vals\ndf.to_csv(\"submission.csv\", index=False)\n</code></pre>\n<p>The model is quite heavy so no TTA or ensembling, just simple thresholding based on validation.</p>\n<p>I did not make a lot of submissions and it was a no-brainer to choose the ones with high CV and public LB</p>\n<h2>Disqualification</h2>\n<p>Was really disappointed to see that I was removed from the leaderboard.<br>\nNo notification, no reasons, just removed. </p>\n<p>Input:</p>\n<ul>\n<li>not using anything from public kernels</li>\n<li>competing solo</li>\n<li>did not create fake accounts. That would be really a blatant action for an experienced kaggler.</li>\n<li>did not share anything with other competitors</li>\n</ul>\n<p>Output:</p>\n<ul>\n<li><strong>cheater</strong>! </li>\n</ul>\n<p>Waiting for the answer from the compliance team and <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "rawMarkdown": "## Solution Overview\n\nOverall I did not make a lot of experiments as I spent just one week for this challenge. I was demotivated by public LB scores, stopped submitting and decided to wait a few weeks for the shakeup.\n\nWith my experiments using MMdetection the best performing model was ConvNeXt base CascadeRCNN.\n\nValidation by video splits:\n- video 0 - 0.71\n- video 1 - 0.67\n- video 2 - 0.8 (used this fold for submissions)\n\n**0.632**  public LB / **0.718** private LB\n\n### Training \nDuring training the following modifications were made to standard mmdet training pipeline:\n\n**Random Sized Crops around bboxes**\nAs boxes were not dense using simple random crop during training was not optimal and I used random sized crop around bounding box if image is not empty. The crops covered multiple scales (1x-4x)  and also allowed to use quite a big batch size during training.\n\n**Balanced sampling from empty/nonempty images**\nTo reduce FP predictions I used all images during training with probability to get a non empty/empty crop 0.6/0.4 respectively.\n\nCompared to full image training this approach improved CV by 3% and slightly improved LB scores. \n\n### Inference\nBecause random sized crops covered 1x-4x scale range I could safely use high resolution. Originally the images had  3840x2160 scales and were downscaled to 1280x720 for competition purposes.\nIt was quite safe to assume that 3840x2160 resolution is the best for inference. It also provided the best results on CV/Public/Private splits\n\n### Postprocessing \n\nRemoved false positive predictions (boosted CV by 1-2%, LB by 0.3%) by leaving only big clusters of detections\n```\ndf = pd.read_csv(\"submission.csv\")\nvals = df.annotations.values\npos = np.array([isinstance(v, str) and len(v) > 2 for v in vals])\npos_dilated = binary_dilation(pos, iterations=1)\npos = remove_small_objects(measure.label(pos_dilated), min_size=10)\nvals[pos == 0] = np.nan\ndf[\"annotations\"] = vals\ndf.to_csv(\"submission.csv\", index=False)\n```\n\nThe model is quite heavy so no TTA or ensembling, just simple thresholding based on validation.\n \nI did not make a lot of submissions and it was a no-brainer to choose the ones with high CV and public LB\n\n## Disqualification\nWas really disappointed to see that I was removed from the leaderboard.\nNo notification, no reasons, just removed. \n\nInput:\n- not using anything from public kernels\n- competing solo\n- did not create fake accounts. That would be really a blatant action for an experienced kaggler.\n- did not share anything with other competitors\n\nOutput:\n- **cheater**! \n\nWaiting for the answer from the compliance team and @addisonhoward",
      "votes": null
    },
    {
      "id": "1697569",
      "postDate": "02/19/2022 17:34:38",
      "content": "<p>Sorry to hear that this happened to you <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>! </p>\n<p>I didn't take part in the competition but read a thread where someone mentioned a LARGE (50+?) Teams got removed this time from the LB? Maybe there were false +ves? I hope they revert the decision for your case. </p>\n<p>QQ: Do you follow progressive learning or is it just training on a smaller resolution and infering on a larger res?</p>",
      "rawMarkdown": "Sorry to hear that this happened to you @selimsef! \n\nI didn't take part in the competition but read a thread where someone mentioned a LARGE (50+?) Teams got removed this time from the LB? Maybe there were false +ves? I hope they revert the decision for your case. \n\nQQ: Do you follow progressive learning or is it just training on a smaller resolution and infering on a larger res?",
      "votes": null
    },
    {
      "id": "1697579",
      "postDate": "02/19/2022 17:40:20",
      "content": "<p>The resolution is quite high, it is like to resize full images to 1280-5120 range and then make a crop 800x800. <br>\nFor fully convolutional networks it is often not a problem to train on small crops and make inference on large images, the key here is to keep similar scales on train/inference.</p>",
      "rawMarkdown": "The resolution is quite high, it is like to resize full images to 1280-5120 range and then make a crop 800x800. \nFor fully convolutional networks it is often not a problem to train on small crops and make inference on large images, the key here is to keep similar scales on train/inference.",
      "votes": null
    },
    {
      "id": "1697668",
      "postDate": "02/19/2022 18:55:24",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> My guess you cant touch sumbission.csv, competition is supposed to be real-time data simulation. Feel bad for you thou.</p>",
      "rawMarkdown": "selimsef My guess you cant touch sumbission.csv, competition is supposed to be real-time data simulation. Feel bad for you thou.",
      "votes": null
    },
    {
      "id": "1697701",
      "postDate": "02/19/2022 19:39:57",
      "content": "<p>It sounds as a possibility, but shouldn't in such a case the system show Submission Error already at the time of submission processing?</p>",
      "rawMarkdown": "It sounds as a possibility, but shouldn't in such a case the system show Submission Error already at the time of submission processing?",
      "votes": null
    },
    {
      "id": "1697715",
      "postDate": "02/19/2022 19:49:01",
      "content": "<p>I have another sub without postprocessing, so this doesn't seem to be the case. <br>\nupd: rechecked kernels code, yeah, unfortunately I've chosen the wrong kernel version.<br>\nStrange though that it did not fail straight away or before LB finalization. <br>\nBut that could be an explanation. </p>",
      "rawMarkdown": "I have another sub without postprocessing, so this doesn't seem to be the case. \nupd: rechecked kernels code, yeah, unfortunately I've chosen the wrong kernel version.\nStrange though that it did not fail straight away or before LB finalization. \nBut that could be an explanation.",
      "votes": null
    },
    {
      "id": "1697813",
      "postDate": "02/19/2022 21:18:12",
      "content": "<p>If it really is the case it will be quite scary for future competitions - I mean the fact, that you have no idea that your submissions are not valid till it is too late to do anything about it. It would be nice to see the system working better.</p>",
      "rawMarkdown": "If it really is the case it will be quite scary for future competitions - I mean the fact, that you have no idea that your submissions are not valid till it is too late to do anything about it. It would be nice to see the system working better.",
      "votes": null
    },
    {
      "id": "1697883",
      "postDate": "02/19/2022 23:47:14",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> Nice solution, in particular the random crops around bboxes! And sorry to hear about your hopefully only temporal disqualification! Do you think you could share the modification in mmdet / code with respect to the random crops? I am working on a similar problem and it would really help me!</p>",
      "rawMarkdown": "selimsef Nice solution, in particular the random crops around bboxes! And sorry to hear about your hopefully only temporal disqualification! Do you think you could share the modification in mmdet / code with respect to the random crops? I am working on a similar problem and it would really help me!",
      "votes": null
    },
    {
      "id": "1698080",
      "postDate": "02/20/2022 05:56:26",
      "content": "<p>I tried crops in lot of ways<br>\n1) 440 square crops no resize on original image 720/1280  YoloM- best score 55.6 public LB<br>\n2) 3200  img size , crops 1480*1856  result were not good . </p>",
      "rawMarkdown": "I tried crops in lot of ways\n1) 440 square crops no resize on original image 720/1280  YoloM- best score 55.6 public LB\n2) 3200  img size , crops 1480*1856  result were not good .",
      "votes": null
    },
    {
      "id": "1698502",
      "postDate": "02/20/2022 12:41:27",
      "content": "<p>From albumentation part you can implement something like <a href=\"https://github.com/selimsef/xview2_solution/blob/5d0caba9c7a9c2707565a189f1a091c86d26b546/augs.py#L160\" target=\"_blank\">this code</a> just change rectangles param to bboxes and parse bbox as xmin, ymin, xmax, ymax  </p>",
      "rawMarkdown": "From albumentation part you can implement something like [this code](https://github.com/selimsef/xview2_solution/blob/5d0caba9c7a9c2707565a189f1a091c86d26b546/augs.py#L160) just change rectangles param to bboxes and parse bbox as xmin, ymin, xmax, ymax",
      "votes": null
    },
    {
      "id": "1698893",
      "postDate": "02/20/2022 18:40:45",
      "content": "<p>This is interesting. </p>\n<p>According to <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/evaluation</a> : </p>\n<blockquote>\n  <p>You must submit to this competition using the provided python time-series API, which ensures that models do not peek forward in time.</p>\n</blockquote>\n<p>Now I'm not sure whether your post-processing peeks forward in time, if it does that's probably be the reason why you got removed. However …</p>\n<p>From  <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/rules\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/rules</a>, using ctrl+f and the word \"submission\", I found these : </p>\n<blockquote>\n  <p>Your Submissions must be made in the manner and format, and in compliance with all other requirements, stated on the Competition Website (the \"Requirements\")</p>\n</blockquote>\n<p>And : </p>\n<blockquote>\n  <p>Your competition submissions (\"Submissions\") must conform to the requirements stated on the Competition Website.</p>\n</blockquote>\n<p>Now as far as I understand, requirements are mentioned here : <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/code-requirements</a> - and do not state that you cannot overwrite the submission.csv file. I agree that if overwriting is not allowed, the submission should fail immediately, and I think that's a fairly easy thing to check.</p>\n<p>Anyways, we need clarification on that. <br>\n<a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  - I'm not sure who I should tag.</p>",
      "rawMarkdown": "This is interesting. \n\nAccording to https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/evaluation : \n> You must submit to this competition using the provided python time-series API, which ensures that models do not peek forward in time.\n\nNow I'm not sure whether your post-processing peeks forward in time, if it does that's probably be the reason why you got removed. However ...\n\nFrom  https://www.kaggle.com/c/tensorflow-great-barrier-reef/rules, using ctrl+f and the word \"submission\", I found these : \n> Your Submissions must be made in the manner and format, and in compliance with all other requirements, stated on the Competition Website (the \"Requirements\")\n\nAnd : \n> Your competition submissions (\"Submissions\") must conform to the requirements stated on the Competition Website.\n\nNow as far as I understand, requirements are mentioned here : https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/code-requirements - and do not state that you cannot overwrite the submission.csv file. I agree that if overwriting is not allowed, the submission should fail immediately, and I think that's a fairly easy thing to check.\n\nAnyways, we need clarification on that. \n@addisonhoward @philculliton @inversion @sohier  - I'm not sure who I should tag.",
      "votes": null
    },
    {
      "id": "1698931",
      "postDate": "02/20/2022 18:58:18",
      "content": "<p>It was mentioned here that it is not allowed:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/297774#1638408\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/297774#1638408</a></p>\n<p>But I agree that it shouldn't be even possible in the first place.</p>",
      "rawMarkdown": "It was mentioned here that it is not allowed:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/297774#1638408\n\nBut I agree that it shouldn't be even possible in the first place.",
      "votes": null
    },
    {
      "id": "1698934",
      "postDate": "02/20/2022 19:00:08",
      "content": "<p>Thanks for sharing, interesting that you got ConvNext to work well, we didn't have any luck with it and the inference runtime was really slow.</p>",
      "rawMarkdown": "Thanks for sharing, interesting that you got ConvNext to work well, we didn't have any luck with it and the inference runtime was really slow.",
      "votes": null
    },
    {
      "id": "1698975",
      "postDate": "02/20/2022 19:49:58",
      "content": "<p>Nice, thanks! That should work!</p>",
      "rawMarkdown": "Nice, thanks! That should work!",
      "votes": null
    },
    {
      "id": "1699038",
      "postDate": "02/20/2022 21:15:43",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  Thanks for the info, that makes it clear! As I joined late I missed these discussions and suggested that once submission passed it should work. <br>\nStill a good achievement for me - was banned the first time 💪😄 </p>",
      "rawMarkdown": "philippsinger  Thanks for the info, that makes it clear! As I joined late I missed these discussions and suggested that once submission passed it should work. \nStill a good achievement for me - was banned the first time 💪😄",
      "votes": null
    },
    {
      "id": "1699039",
      "postDate": "02/20/2022 21:17:55",
      "content": "<p>It performs really great on other tasks as well. But it is very slow compared to original Yolo models, that's true.  <br>\nFor this dataset one needs to tune loss weights, I looked at absolute loss values and added some multipliers, for RPN head it was something like 100 due to rare boxes.  <br>\nIn general losses for RPN head are overwhelmed by negative samples which leads to bad recall if image is really big like 3k and it has just a single box.  <br>\nPlayed around with loss balancing, hard negative mining, focal loss etc. but did not have any improvements. <br>\nSo overall this dataset is not very RCNN friendly. </p>",
      "rawMarkdown": "It performs really great on other tasks as well. But it is very slow compared to original Yolo models, that's true.  \nFor this dataset one needs to tune loss weights, I looked at absolute loss values and added some multipliers, for RPN head it was something like 100 due to rare boxes.  \nIn general losses for RPN head are overwhelmed by negative samples which leads to bad recall if image is really big like 3k and it has just a single box.  \nPlayed around with loss balancing, hard negative mining, focal loss etc. but did not have any improvements. \nSo overall this dataset is not very RCNN friendly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1697569,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/19/2022 17:34:38",
      "content": "<p>Sorry to hear that this happened to you <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>! </p>\n<p>I didn't take part in the competition but read a thread where someone mentioned a LARGE (50+?) Teams got removed this time from the LB? Maybe there were false +ves? I hope they revert the decision for your case. </p>\n<p>QQ: Do you follow progressive learning or is it just training on a smaller resolution and infering on a larger res?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1697579,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2022 17:40:20",
          "content": "<p>The resolution is quite high, it is like to resize full images to 1280-5120 range and then make a crop 800x800. <br>\nFor fully convolutional networks it is often not a problem to train on small crops and make inference on large images, the key here is to keep similar scales on train/inference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698080,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/20/2022 05:56:26",
          "content": "<p>I tried crops in lot of ways<br>\n1) 440 square crops no resize on original image 720/1280  YoloM- best score 55.6 public LB<br>\n2) 3200  img size , crops 1480*1856  result were not good . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1697668,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "02/19/2022 18:55:24",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> My guess you cant touch sumbission.csv, competition is supposed to be real-time data simulation. Feel bad for you thou.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1697701,
          "author_name": "blankaf",
          "author_url": "",
          "post_date": "02/19/2022 19:39:57",
          "content": "<p>It sounds as a possibility, but shouldn't in such a case the system show Submission Error already at the time of submission processing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697715,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2022 19:49:01",
          "content": "<p>I have another sub without postprocessing, so this doesn't seem to be the case. <br>\nupd: rechecked kernels code, yeah, unfortunately I've chosen the wrong kernel version.<br>\nStrange though that it did not fail straight away or before LB finalization. <br>\nBut that could be an explanation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697813,
          "author_name": "blankaf",
          "author_url": "",
          "post_date": "02/19/2022 21:18:12",
          "content": "<p>If it really is the case it will be quite scary for future competitions - I mean the fact, that you have no idea that your submissions are not valid till it is too late to do anything about it. It would be nice to see the system working better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698893,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/20/2022 18:40:45",
          "content": "<p>This is interesting. </p>\n<p>According to <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/evaluation</a> : </p>\n<blockquote>\n  <p>You must submit to this competition using the provided python time-series API, which ensures that models do not peek forward in time.</p>\n</blockquote>\n<p>Now I'm not sure whether your post-processing peeks forward in time, if it does that's probably be the reason why you got removed. However …</p>\n<p>From  <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/rules\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/rules</a>, using ctrl+f and the word \"submission\", I found these : </p>\n<blockquote>\n  <p>Your Submissions must be made in the manner and format, and in compliance with all other requirements, stated on the Competition Website (the \"Requirements\")</p>\n</blockquote>\n<p>And : </p>\n<blockquote>\n  <p>Your competition submissions (\"Submissions\") must conform to the requirements stated on the Competition Website.</p>\n</blockquote>\n<p>Now as far as I understand, requirements are mentioned here : <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/code-requirements</a> - and do not state that you cannot overwrite the submission.csv file. I agree that if overwriting is not allowed, the submission should fail immediately, and I think that's a fairly easy thing to check.</p>\n<p>Anyways, we need clarification on that. <br>\n<a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  - I'm not sure who I should tag.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698931,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/20/2022 18:58:18",
          "content": "<p>It was mentioned here that it is not allowed:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/297774#1638408\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/297774#1638408</a></p>\n<p>But I agree that it shouldn't be even possible in the first place.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699038,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2022 21:15:43",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  Thanks for the info, that makes it clear! As I joined late I missed these discussions and suggested that once submission passed it should work. <br>\nStill a good achievement for me - was banned the first time 💪😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1697883,
      "author_name": "hannes82",
      "author_url": "",
      "post_date": "02/19/2022 23:47:14",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> Nice solution, in particular the random crops around bboxes! And sorry to hear about your hopefully only temporal disqualification! Do you think you could share the modification in mmdet / code with respect to the random crops? I am working on a similar problem and it would really help me!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1698502,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2022 12:41:27",
          "content": "<p>From albumentation part you can implement something like <a href=\"https://github.com/selimsef/xview2_solution/blob/5d0caba9c7a9c2707565a189f1a091c86d26b546/augs.py#L160\" target=\"_blank\">this code</a> just change rectangles param to bboxes and parse bbox as xmin, ymin, xmax, ymax  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698975,
          "author_name": "hannes82",
          "author_url": "",
          "post_date": "02/20/2022 19:49:58",
          "content": "<p>Nice, thanks! That should work!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1698934,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "02/20/2022 19:00:08",
      "content": "<p>Thanks for sharing, interesting that you got ConvNext to work well, we didn't have any luck with it and the inference runtime was really slow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1699039,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2022 21:17:55",
          "content": "<p>It performs really great on other tasks as well. But it is very slow compared to original Yolo models, that's true.  <br>\nFor this dataset one needs to tune loss weights, I looked at absolute loss values and added some multipliers, for RPN head it was something like 100 due to rare boxes.  <br>\nIn general losses for RPN head are overwhelmed by negative samples which leads to bad recall if image is really big like 3k and it has just a single box.  <br>\nPlayed around with loss balancing, hard negative mining, focal loss etc. but did not have any improvements. <br>\nSo overall this dataset is not very RCNN friendly. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1697472": "## Solution Overview\n\nOverall I did not make a lot of experiments as I spent just one week for this challenge. I was demotivated by public LB scores, stopped submitting and decided to wait a few weeks for the shakeup.\n\nWith my experiments using MMdetection the best performing model was ConvNeXt base CascadeRCNN.\n\nValidation by video splits:\n- video 0 - 0.71\n- video 1 - 0.67\n- video 2 - 0.8 (used this fold for submissions)\n\n**0.632**  public LB / **0.718** private LB\n\n### Training \nDuring training the following modifications were made to standard mmdet training pipeline:\n\n**Random Sized Crops around bboxes**\nAs boxes were not dense using simple random crop during training was not optimal and I used random sized crop around bounding box if image is not empty. The crops covered multiple scales (1x-4x)  and also allowed to use quite a big batch size during training.\n\n**Balanced sampling from empty/nonempty images**\nTo reduce FP predictions I used all images during training with probability to get a non empty/empty crop 0.6/0.4 respectively.\n\nCompared to full image training this approach improved CV by 3% and slightly improved LB scores. \n\n### Inference\nBecause random sized crops covered 1x-4x scale range I could safely use high resolution. Originally the images had  3840x2160 scales and were downscaled to 1280x720 for competition purposes.\nIt was quite safe to assume that 3840x2160 resolution is the best for inference. It also provided the best results on CV/Public/Private splits\n\n### Postprocessing \n\nRemoved false positive predictions (boosted CV by 1-2%, LB by 0.3%) by leaving only big clusters of detections\n```\ndf = pd.read_csv(\"submission.csv\")\nvals = df.annotations.values\npos = np.array([isinstance(v, str) and len(v) > 2 for v in vals])\npos_dilated = binary_dilation(pos, iterations=1)\npos = remove_small_objects(measure.label(pos_dilated), min_size=10)\nvals[pos == 0] = np.nan\ndf[\"annotations\"] = vals\ndf.to_csv(\"submission.csv\", index=False)\n```\n\nThe model is quite heavy so no TTA or ensembling, just simple thresholding based on validation.\n \nI did not make a lot of submissions and it was a no-brainer to choose the ones with high CV and public LB\n\n## Disqualification\nWas really disappointed to see that I was removed from the leaderboard.\nNo notification, no reasons, just removed. \n\nInput:\n- not using anything from public kernels\n- competing solo\n- did not create fake accounts. That would be really a blatant action for an experienced kaggler.\n- did not share anything with other competitors\n\nOutput:\n- **cheater**! \n\nWaiting for the answer from the compliance team and @addisonhoward",
    "1697569": "Sorry to hear that this happened to you @selimsef! \n\nI didn't take part in the competition but read a thread where someone mentioned a LARGE (50+?) Teams got removed this time from the LB? Maybe there were false +ves? I hope they revert the decision for your case. \n\nQQ: Do you follow progressive learning or is it just training on a smaller resolution and infering on a larger res?",
    "1697579": "The resolution is quite high, it is like to resize full images to 1280-5120 range and then make a crop 800x800. \nFor fully convolutional networks it is often not a problem to train on small crops and make inference on large images, the key here is to keep similar scales on train/inference.",
    "1697668": "selimsef My guess you cant touch sumbission.csv, competition is supposed to be real-time data simulation. Feel bad for you thou.",
    "1697701": "It sounds as a possibility, but shouldn't in such a case the system show Submission Error already at the time of submission processing?",
    "1697715": "I have another sub without postprocessing, so this doesn't seem to be the case. \nupd: rechecked kernels code, yeah, unfortunately I've chosen the wrong kernel version.\nStrange though that it did not fail straight away or before LB finalization. \nBut that could be an explanation.",
    "1697813": "If it really is the case it will be quite scary for future competitions - I mean the fact, that you have no idea that your submissions are not valid till it is too late to do anything about it. It would be nice to see the system working better.",
    "1697883": "selimsef Nice solution, in particular the random crops around bboxes! And sorry to hear about your hopefully only temporal disqualification! Do you think you could share the modification in mmdet / code with respect to the random crops? I am working on a similar problem and it would really help me!",
    "1698080": "I tried crops in lot of ways\n1) 440 square crops no resize on original image 720/1280  YoloM- best score 55.6 public LB\n2) 3200  img size , crops 1480*1856  result were not good .",
    "1698502": "From albumentation part you can implement something like [this code](https://github.com/selimsef/xview2_solution/blob/5d0caba9c7a9c2707565a189f1a091c86d26b546/augs.py#L160) just change rectangles param to bboxes and parse bbox as xmin, ymin, xmax, ymax",
    "1698893": "This is interesting. \n\nAccording to https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/evaluation : \n> You must submit to this competition using the provided python time-series API, which ensures that models do not peek forward in time.\n\nNow I'm not sure whether your post-processing peeks forward in time, if it does that's probably be the reason why you got removed. However ...\n\nFrom  https://www.kaggle.com/c/tensorflow-great-barrier-reef/rules, using ctrl+f and the word \"submission\", I found these : \n> Your Submissions must be made in the manner and format, and in compliance with all other requirements, stated on the Competition Website (the \"Requirements\")\n\nAnd : \n> Your competition submissions (\"Submissions\") must conform to the requirements stated on the Competition Website.\n\nNow as far as I understand, requirements are mentioned here : https://www.kaggle.com/c/tensorflow-great-barrier-reef/overview/code-requirements - and do not state that you cannot overwrite the submission.csv file. I agree that if overwriting is not allowed, the submission should fail immediately, and I think that's a fairly easy thing to check.\n\nAnyways, we need clarification on that. \n@addisonhoward @philculliton @inversion @sohier  - I'm not sure who I should tag.",
    "1698931": "It was mentioned here that it is not allowed:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/297774#1638408\n\nBut I agree that it shouldn't be even possible in the first place.",
    "1698934": "Thanks for sharing, interesting that you got ConvNext to work well, we didn't have any luck with it and the inference runtime was really slow.",
    "1698975": "Nice, thanks! That should work!",
    "1699038": "philippsinger  Thanks for the info, that makes it clear! As I joined late I missed these discussions and suggested that once submission passed it should work. \nStill a good achievement for me - was banned the first time 💪😄",
    "1699039": "It performs really great on other tasks as well. But it is very slow compared to original Yolo models, that's true.  \nFor this dataset one needs to tune loss weights, I looked at absolute loss values and added some multipliers, for RPN head it was something like 100 due to rare boxes.  \nIn general losses for RPN head are overwhelmed by negative samples which leads to bad recall if image is really big like 3k and it has just a single box.  \nPlayed around with loss balancing, hard negative mining, focal loss etc. but did not have any improvements. \nSo overall this dataset is not very RCNN friendly."
  },
  "source": "meta"
}