{
  "id": 264287,
  "title": "11th Place Solution - My Part",
  "url": "/competitions/siim-covid19-detection/discussion/264287",
  "author_name": "",
  "post_date": "2021-08-11T15:39:39.629081300Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, I would like to thank the host team for all your support and a great competition. Thank you <a href=\"https://www.kaggle.com/yingpengchen\" target=\"_blank\">@yingpengchen</a> <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> <a href=\"https://www.kaggle.com/socom20\" target=\"_blank\">@socom20</a>, I have had a great experience and learned a lot working with you as a team!</p>\n<p>Full solution of my team can be found at: <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263701\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/263701</a></p>\n<p>During this competition and after team merging, I mainly focused on image-level and post-processing, here is what I did for my part:</p>\n<p><strong>Final result</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>none</td>\n<td>0.134</td>\n<td>--</td>\n</tr>\n<tr>\n<td>opacity</td>\n<td>0.100</td>\n<td>--</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Cross-validation</strong><br>\nStratified K Fold by StudyID</p>\n<p><strong>None class prediction</strong><br>\n<code>none_probbility = np.prod(1 - box_conf_i)</code></p>\n<p><strong>Modeling</strong><br>\n<strong>Detectors trained with competition train data only</strong></p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>image size</th>\n<th>batch size</th>\n<th>epochs</th>\n<th>TTA</th>\n<th>iou</th>\n<th>conf</th>\n<th>CV opacity</th>\n<th>CV none</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>VFNetr50</td>\n<td>640</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.48358</td>\n<td>0.23121</td>\n</tr>\n<tr>\n<td>Yolov5m*</td>\n<td>1024</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.48148</td>\n<td>0.76216</td>\n</tr>\n<tr>\n<td>Yolov5x*</td>\n<td>640</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.50930</td>\n<td>--</td>\n</tr>\n<tr>\n<td>Yolov5x</td>\n<td>512</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.51690</td>\n<td>0.78192</td>\n</tr>\n<tr>\n<td>Yolov5l6</td>\n<td>512</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.51650</td>\n<td>0.78190</td>\n</tr>\n<tr>\n<td>Yolov5x6</td>\n<td>512</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.51754</td>\n<td>0.77820</td>\n</tr>\n</tbody>\n</table>\n<p>*: trained with different hyperparameter config</p>\n<p><strong>Detectors trained with pseudo data</strong></p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>image size</th>\n<th>batch size</th>\n<th>epochs</th>\n<th>TTA</th>\n<th>iou</th>\n<th>conf</th>\n<th>CV opacity</th>\n<th>CV none</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Yolov5x</td>\n<td>512</td>\n<td>8</td>\n<td>50</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.53870</td>\n<td>0.79028</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Pseudo labels - training process</strong><br>\nDatasets: Public test set + BIMCV + RICORD</p>\n<ul>\n<li>For BIMCV, the dataset contains a lot of images which are taken for the left/right side of human body. In order to reduce noise, my teammate <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> and I manually removed them from the dataset. And since both training and test data in this competition are drawn from this dataset, to avoid leakage in validation, I removed all of the duplicate images and images that have the same StudyID with these duplicates.</li>\n</ul>\n<p>Making pseudo labels:</p>\n<ul>\n<li>Label images with <code>none_probability &gt; 0.6</code> as none class images</li>\n<li>For those have <code>none_probability &lt;= 0.6</code>, keep boxes with confident &gt; 0.095<br>\nThese thresholds are chosen in order to maximize the f1 score.</li>\n</ul>\n<p>Training:<br>\nAll datasets are merged together and used to train with the same procedure as without pseudo data.</p>\n<p><strong>Post-processing</strong>    </p>\n<ul>\n<li>Weighted boxes fusion with <code>iou_thr=0.6</code> and <code>conf_thr=0.0001</code> as boxes fusion method</li>\n<li><code>box_conf = box_conf**0.84 * (1 - none_probability)**0.16</code></li>\n<li><code>none_probability = none_probability*0.5 + negative_probability*0.5</code></li>\n<li><code>negative_probability = none_probability*0.3 + negative_probability*0.7</code></li>\n</ul>\n<p><strong>Final Submission</strong><br>\nFor final submission, we used Yolotrs-384 + Yolov5x-640 + Yolov5x-512-pseudo labels, all with TTA.</p>\n<p>P/s: Few days before the deadline, I trained several detectors (mentioned above) to improve the quality of pseudo labels and the quality of predictions in general and planned to finish training final models with pseudo labels, the one we used in final submission is actually just for my experiment. Unfortunately, in the last 2 days before the deadline, my laptop had a problem and stopped working. As social distancing is happening in my city, I could not have any help from outside, so my work had to be stopped. This forced us to use whatever we had then. Final result could have been improved without my personal problem. I would like to apologize to my teammates for this. Also, thank you for helping me to finish my part of final submission on the very last day.</p>",
  "messages": [
    {
      "id": "1466741",
      "postDate": "08/11/2021 15:39:39",
      "content": "<p>First of all, I would like to thank the host team for all your support and a great competition. Thank you <a href=\"https://www.kaggle.com/yingpengchen\" target=\"_blank\">@yingpengchen</a> <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> <a href=\"https://www.kaggle.com/socom20\" target=\"_blank\">@socom20</a>, I have had a great experience and learned a lot working with you as a team!</p>\n<p>Full solution of my team can be found at: <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263701\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/263701</a></p>\n<p>During this competition and after team merging, I mainly focused on image-level and post-processing, here is what I did for my part:</p>\n<p><strong>Final result</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>none</td>\n<td>0.134</td>\n<td>--</td>\n</tr>\n<tr>\n<td>opacity</td>\n<td>0.100</td>\n<td>--</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Cross-validation</strong><br>\nStratified K Fold by StudyID</p>\n<p><strong>None class prediction</strong><br>\n<code>none_probbility = np.prod(1 - box_conf_i)</code></p>\n<p><strong>Modeling</strong><br>\n<strong>Detectors trained with competition train data only</strong></p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>image size</th>\n<th>batch size</th>\n<th>epochs</th>\n<th>TTA</th>\n<th>iou</th>\n<th>conf</th>\n<th>CV opacity</th>\n<th>CV none</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>VFNetr50</td>\n<td>640</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.48358</td>\n<td>0.23121</td>\n</tr>\n<tr>\n<td>Yolov5m*</td>\n<td>1024</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.48148</td>\n<td>0.76216</td>\n</tr>\n<tr>\n<td>Yolov5x*</td>\n<td>640</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.50930</td>\n<td>--</td>\n</tr>\n<tr>\n<td>Yolov5x</td>\n<td>512</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.51690</td>\n<td>0.78192</td>\n</tr>\n<tr>\n<td>Yolov5l6</td>\n<td>512</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.51650</td>\n<td>0.78190</td>\n</tr>\n<tr>\n<td>Yolov5x6</td>\n<td>512</td>\n<td>8</td>\n<td>35</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.51754</td>\n<td>0.77820</td>\n</tr>\n</tbody>\n</table>\n<p>*: trained with different hyperparameter config</p>\n<p><strong>Detectors trained with pseudo data</strong></p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>image size</th>\n<th>batch size</th>\n<th>epochs</th>\n<th>TTA</th>\n<th>iou</th>\n<th>conf</th>\n<th>CV opacity</th>\n<th>CV none</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Yolov5x</td>\n<td>512</td>\n<td>8</td>\n<td>50</td>\n<td>Y</td>\n<td>0.5</td>\n<td>0.001</td>\n<td>0.53870</td>\n<td>0.79028</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Pseudo labels - training process</strong><br>\nDatasets: Public test set + BIMCV + RICORD</p>\n<ul>\n<li>For BIMCV, the dataset contains a lot of images which are taken for the left/right side of human body. In order to reduce noise, my teammate <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> and I manually removed them from the dataset. And since both training and test data in this competition are drawn from this dataset, to avoid leakage in validation, I removed all of the duplicate images and images that have the same StudyID with these duplicates.</li>\n</ul>\n<p>Making pseudo labels:</p>\n<ul>\n<li>Label images with <code>none_probability &gt; 0.6</code> as none class images</li>\n<li>For those have <code>none_probability &lt;= 0.6</code>, keep boxes with confident &gt; 0.095<br>\nThese thresholds are chosen in order to maximize the f1 score.</li>\n</ul>\n<p>Training:<br>\nAll datasets are merged together and used to train with the same procedure as without pseudo data.</p>\n<p><strong>Post-processing</strong>    </p>\n<ul>\n<li>Weighted boxes fusion with <code>iou_thr=0.6</code> and <code>conf_thr=0.0001</code> as boxes fusion method</li>\n<li><code>box_conf = box_conf**0.84 * (1 - none_probability)**0.16</code></li>\n<li><code>none_probability = none_probability*0.5 + negative_probability*0.5</code></li>\n<li><code>negative_probability = none_probability*0.3 + negative_probability*0.7</code></li>\n</ul>\n<p><strong>Final Submission</strong><br>\nFor final submission, we used Yolotrs-384 + Yolov5x-640 + Yolov5x-512-pseudo labels, all with TTA.</p>\n<p>P/s: Few days before the deadline, I trained several detectors (mentioned above) to improve the quality of pseudo labels and the quality of predictions in general and planned to finish training final models with pseudo labels, the one we used in final submission is actually just for my experiment. Unfortunately, in the last 2 days before the deadline, my laptop had a problem and stopped working. As social distancing is happening in my city, I could not have any help from outside, so my work had to be stopped. This forced us to use whatever we had then. Final result could have been improved without my personal problem. I would like to apologize to my teammates for this. Also, thank you for helping me to finish my part of final submission on the very last day.</p>",
      "rawMarkdown": "First of all, I would like to thank the host team for all your support and a great competition. Thank you @yingpengchen @morizin @joven1997 @socom20, I have had a great experience and learned a lot working with you as a team!\n\nFull solution of my team can be found at: https://www.kaggle.com/c/siim-covid19-detection/discussion/263701\n\nDuring this competition and after team merging, I mainly focused on image-level and post-processing, here is what I did for my part:\n\n**Final result**\n| | Public LB | Private LB |\n| --- | --- |\n| none | 0.134 | -- |\n| opacity | 0.100 | -- |\n\n**Cross-validation**\nStratified K Fold by StudyID\n\n**None class prediction**\n`none_probbility = np.prod(1 - box_conf_i)`\n\n**Modeling**\n**Detectors trained with competition train data only**\n| backbone | image size | batch size | epochs | TTA | iou | conf | CV opacity | CV none\n| --- | --- |\n| VFNetr50| 640 | 8 | 35 | Y | 0.5 | 0.001 | 0.48358 | 0.23121\n| Yolov5m* | 1024 | 8 | 35 | Y | 0.5 | 0.001 | 0.48148 | 0.76216\n| Yolov5x* | 640 | 8 | 35 | Y | 0.5 | 0.001 | 0.50930| --\n| Yolov5x | 512 | 8 | 35 | Y | 0.5 | 0.001 | 0.51690 | 0.78192\n| Yolov5l6 | 512 | 8 | 35 | Y | 0.5 | 0.001 | 0.51650 | 0.78190\n| Yolov5x6 | 512 | 8 | 35 | Y | 0.5 | 0.001 | 0.51754 | 0.77820\n\n*: trained with different hyperparameter config\n\n**Detectors trained with pseudo data**\n| backbone | image size | batch size | epochs | TTA | iou | conf | CV opacity | CV none\n| --- | --- |\n| Yolov5x | 512| 8 | 50 | Y | 0.5 | 0.001 | 0.53870 | 0.79028\n\n**Pseudo labels - training process**\nDatasets: Public test set + BIMCV + RICORD\n- For BIMCV, the dataset contains a lot of images which are taken for the left/right side of human body. In order to reduce noise, my teammate @joven1997 and I manually removed them from the dataset. And since both training and test data in this competition are drawn from this dataset, to avoid leakage in validation, I removed all of the duplicate images and images that have the same StudyID with these duplicates.\n\nMaking pseudo labels:\n- Label images with `none_probability > 0.6` as none class images\n- For those have `none_probability <= 0.6`, keep boxes with confident > 0.095\nThese thresholds are chosen in order to maximize the f1 score.\n\nTraining:\nAll datasets are merged together and used to train with the same procedure as without pseudo data.\n\n**Post-processing**\t\n- Weighted boxes fusion with `iou_thr=0.6` and `conf_thr=0.0001` as boxes fusion method\n- `box_conf = box_conf**0.84 * (1 - none_probability)**0.16`\n- `none_probability = none_probability*0.5 + negative_probability*0.5`\n- `negative_probability = none_probability*0.3 + negative_probability*0.7`\n\n**Final Submission**\nFor final submission, we used Yolotrs-384 + Yolov5x-640 + Yolov5x-512-pseudo labels, all with TTA.\n\nP/s: Few days before the deadline, I trained several detectors (mentioned above) to improve the quality of pseudo labels and the quality of predictions in general and planned to finish training final models with pseudo labels, the one we used in final submission is actually just for my experiment. Unfortunately, in the last 2 days before the deadline, my laptop had a problem and stopped working. As social distancing is happening in my city, I could not have any help from outside, so my work had to be stopped. This forced us to use whatever we had then. Final result could have been improved without my personal problem. I would like to apologize to my teammates for this. Also, thank you for helping me to finish my part of final submission on the very last day.",
      "votes": null
    },
    {
      "id": "1466746",
      "postDate": "08/11/2021 15:42:12",
      "content": "<p>Hi, I think it would be <code>none_probability</code> instead of <code>bone_probability</code>.</p>",
      "rawMarkdown": "Hi, I think it would be `none_probability` instead of `bone_probability`.",
      "votes": null
    },
    {
      "id": "1466753",
      "postDate": "08/11/2021 15:44:52",
      "content": "<p>What does the <code>Yolotrs</code> mean?</p>",
      "rawMarkdown": "What does the `Yolotrs` mean?",
      "votes": null
    },
    {
      "id": "1466754",
      "postDate": "08/11/2021 15:44:54",
      "content": "<p>My mistake 😅 Thank you for pointing it out</p>",
      "rawMarkdown": "My mistake 😅 Thank you for pointing it out",
      "votes": null
    },
    {
      "id": "1466760",
      "postDate": "08/11/2021 15:48:50",
      "content": "<p>It is Yolo-transformer-s</p>",
      "rawMarkdown": "It is Yolo-transformer-s",
      "votes": null
    },
    {
      "id": "1467375",
      "postDate": "08/12/2021 00:34:05",
      "content": "<p>yolov5s+vit(transformer): <a href=\"https://www.gitmemory.com/issue/ultralytics/yolov5/2329/808852867\" target=\"_blank\">https://www.gitmemory.com/issue/ultralytics/yolov5/2329/808852867</a></p>",
      "rawMarkdown": "yolov5s+vit(transformer): https://www.gitmemory.com/issue/ultralytics/yolov5/2329/808852867",
      "votes": null
    },
    {
      "id": "1467474",
      "postDate": "08/12/2021 02:24:00",
      "content": "<p>Great solution, your work is very similar to mine in this competition. Congrats for the gold medal. How can you 'manually removed'  duplicate images ? This action required domain knowledge ?</p>",
      "rawMarkdown": "Great solution, your work is very similar to mine in this competition. Congrats for the gold medal. How can you 'manually removed'  duplicate images ? This action required domain knowledge ?",
      "votes": null
    },
    {
      "id": "1471824",
      "postDate": "08/14/2021 13:39:49",
      "content": "<p>Precisely, we only manually removed images which show the left or right side of human lung since they are not relevant to the competition data. For duplicate images, there are several ways to handle them, for example using perceptual hashing and Hamming distance. However,  in our case we extracted them based on this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/255260\" target=\"_blank\">post</a> with duplicates already computed and mapped into a .csv file.</p>",
      "rawMarkdown": "Precisely, we only manually removed images which show the left or right side of human lung since they are not relevant to the competition data. For duplicate images, there are several ways to handle them, for example using perceptual hashing and Hamming distance. However,  in our case we extracted them based on this [post](https://www.kaggle.com/c/siim-covid19-detection/discussion/255260) with duplicates already computed and mapped into a .csv file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1466746,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "08/11/2021 15:42:12",
      "content": "<p>Hi, I think it would be <code>none_probability</code> instead of <code>bone_probability</code>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1466754,
          "author_name": "quochungto",
          "author_url": "",
          "post_date": "08/11/2021 15:44:54",
          "content": "<p>My mistake 😅 Thank you for pointing it out</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1466753,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "08/11/2021 15:44:52",
      "content": "<p>What does the <code>Yolotrs</code> mean?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1466760,
          "author_name": "quochungto",
          "author_url": "",
          "post_date": "08/11/2021 15:48:50",
          "content": "<p>It is Yolo-transformer-s</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1467375,
          "author_name": "yingpengchen",
          "author_url": "",
          "post_date": "08/12/2021 00:34:05",
          "content": "<p>yolov5s+vit(transformer): <a href=\"https://www.gitmemory.com/issue/ultralytics/yolov5/2329/808852867\" target=\"_blank\">https://www.gitmemory.com/issue/ultralytics/yolov5/2329/808852867</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1467474,
      "author_name": "researchbntz",
      "author_url": "",
      "post_date": "08/12/2021 02:24:00",
      "content": "<p>Great solution, your work is very similar to mine in this competition. Congrats for the gold medal. How can you 'manually removed'  duplicate images ? This action required domain knowledge ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1471824,
          "author_name": "quochungto",
          "author_url": "",
          "post_date": "08/14/2021 13:39:49",
          "content": "<p>Precisely, we only manually removed images which show the left or right side of human lung since they are not relevant to the competition data. For duplicate images, there are several ways to handle them, for example using perceptual hashing and Hamming distance. However,  in our case we extracted them based on this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/255260\" target=\"_blank\">post</a> with duplicates already computed and mapped into a .csv file.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1466741": "First of all, I would like to thank the host team for all your support and a great competition. Thank you @yingpengchen @morizin @joven1997 @socom20, I have had a great experience and learned a lot working with you as a team!\n\nFull solution of my team can be found at: https://www.kaggle.com/c/siim-covid19-detection/discussion/263701\n\nDuring this competition and after team merging, I mainly focused on image-level and post-processing, here is what I did for my part:\n\n**Final result**\n| | Public LB | Private LB |\n| --- | --- |\n| none | 0.134 | -- |\n| opacity | 0.100 | -- |\n\n**Cross-validation**\nStratified K Fold by StudyID\n\n**None class prediction**\n`none_probbility = np.prod(1 - box_conf_i)`\n\n**Modeling**\n**Detectors trained with competition train data only**\n| backbone | image size | batch size | epochs | TTA | iou | conf | CV opacity | CV none\n| --- | --- |\n| VFNetr50| 640 | 8 | 35 | Y | 0.5 | 0.001 | 0.48358 | 0.23121\n| Yolov5m* | 1024 | 8 | 35 | Y | 0.5 | 0.001 | 0.48148 | 0.76216\n| Yolov5x* | 640 | 8 | 35 | Y | 0.5 | 0.001 | 0.50930| --\n| Yolov5x | 512 | 8 | 35 | Y | 0.5 | 0.001 | 0.51690 | 0.78192\n| Yolov5l6 | 512 | 8 | 35 | Y | 0.5 | 0.001 | 0.51650 | 0.78190\n| Yolov5x6 | 512 | 8 | 35 | Y | 0.5 | 0.001 | 0.51754 | 0.77820\n\n*: trained with different hyperparameter config\n\n**Detectors trained with pseudo data**\n| backbone | image size | batch size | epochs | TTA | iou | conf | CV opacity | CV none\n| --- | --- |\n| Yolov5x | 512| 8 | 50 | Y | 0.5 | 0.001 | 0.53870 | 0.79028\n\n**Pseudo labels - training process**\nDatasets: Public test set + BIMCV + RICORD\n- For BIMCV, the dataset contains a lot of images which are taken for the left/right side of human body. In order to reduce noise, my teammate @joven1997 and I manually removed them from the dataset. And since both training and test data in this competition are drawn from this dataset, to avoid leakage in validation, I removed all of the duplicate images and images that have the same StudyID with these duplicates.\n\nMaking pseudo labels:\n- Label images with `none_probability > 0.6` as none class images\n- For those have `none_probability <= 0.6`, keep boxes with confident > 0.095\nThese thresholds are chosen in order to maximize the f1 score.\n\nTraining:\nAll datasets are merged together and used to train with the same procedure as without pseudo data.\n\n**Post-processing**\t\n- Weighted boxes fusion with `iou_thr=0.6` and `conf_thr=0.0001` as boxes fusion method\n- `box_conf = box_conf**0.84 * (1 - none_probability)**0.16`\n- `none_probability = none_probability*0.5 + negative_probability*0.5`\n- `negative_probability = none_probability*0.3 + negative_probability*0.7`\n\n**Final Submission**\nFor final submission, we used Yolotrs-384 + Yolov5x-640 + Yolov5x-512-pseudo labels, all with TTA.\n\nP/s: Few days before the deadline, I trained several detectors (mentioned above) to improve the quality of pseudo labels and the quality of predictions in general and planned to finish training final models with pseudo labels, the one we used in final submission is actually just for my experiment. Unfortunately, in the last 2 days before the deadline, my laptop had a problem and stopped working. As social distancing is happening in my city, I could not have any help from outside, so my work had to be stopped. This forced us to use whatever we had then. Final result could have been improved without my personal problem. I would like to apologize to my teammates for this. Also, thank you for helping me to finish my part of final submission on the very last day.",
    "1466746": "Hi, I think it would be `none_probability` instead of `bone_probability`.",
    "1466753": "What does the `Yolotrs` mean?",
    "1466754": "My mistake 😅 Thank you for pointing it out",
    "1466760": "It is Yolo-transformer-s",
    "1467375": "yolov5s+vit(transformer): https://www.gitmemory.com/issue/ultralytics/yolov5/2329/808852867",
    "1467474": "Great solution, your work is very similar to mine in this competition. Congrats for the gold medal. How can you 'manually removed'  duplicate images ? This action required domain knowledge ?",
    "1471824": "Precisely, we only manually removed images which show the left or right side of human lung since they are not relevant to the competition data. For duplicate images, there are several ways to handle them, for example using perceptual hashing and Hamming distance. However,  in our case we extracted them based on this [post](https://www.kaggle.com/c/siim-covid19-detection/discussion/255260) with duplicates already computed and mapped into a .csv file."
  },
  "source": "meta"
}