{
  "id": 243273,
  "title": "How To Predict The Correct Image To Perform Object Detection On [Sometimes]",
  "url": "/competitions/siim-covid19-detection/discussion/243273",
  "author_name": "",
  "post_date": "2021-06-01T21:01:19.008808500Z",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>I will throw this out there… not sure how much value it has as the organizers will be fixing things soon (hopefully).</p>\n<p>If a study has multiple images, you can usually tell the image that you should predict the bounding boxes for by looking at the dicom metadata field <strong><code>SeriesNumber</code></strong>. </p>\n<p><strong>Whichever image has the lowest</strong> <strong><code>SeriesNumber</code></strong> <strong>in the study is the one that you will need to predict bounding boxes on.</strong></p>\n<p>I confirmed this theory using many studies and it holds true for most of them.</p>\n<hr>\n<p><strong>CAVEATS</strong></p>\n<ul>\n<li>Some images don't have a listed <strong><code>SeriesNumber</code></strong> (<strong><code>None</code></strong>). There can be multiple images in a study with this value (<strong><code>None</code></strong>) and as such we can't tell which image is the one we should predict on.</li>\n<li>There are odd cases where the above rule doesn't hold true</li>\n<li>I have no idea if this will hold true in the test set and I don't have a model to give it a shot with.</li>\n</ul>\n<hr>\n<p>I would normally put together a big post showing the evidence to support this… but I don't have time and if it's being fixed soon I don't see the value. That being said, I thought I would share.</p>\n<p>Hope this helps!</p>",
  "messages": [
    {
      "id": "1332006",
      "postDate": "06/01/2021 21:01:19",
      "content": "<p>Hi there,</p>\n<p>I will throw this out there… not sure how much value it has as the organizers will be fixing things soon (hopefully).</p>\n<p>If a study has multiple images, you can usually tell the image that you should predict the bounding boxes for by looking at the dicom metadata field <strong><code>SeriesNumber</code></strong>. </p>\n<p><strong>Whichever image has the lowest</strong> <strong><code>SeriesNumber</code></strong> <strong>in the study is the one that you will need to predict bounding boxes on.</strong></p>\n<p>I confirmed this theory using many studies and it holds true for most of them.</p>\n<hr>\n<p><strong>CAVEATS</strong></p>\n<ul>\n<li>Some images don't have a listed <strong><code>SeriesNumber</code></strong> (<strong><code>None</code></strong>). There can be multiple images in a study with this value (<strong><code>None</code></strong>) and as such we can't tell which image is the one we should predict on.</li>\n<li>There are odd cases where the above rule doesn't hold true</li>\n<li>I have no idea if this will hold true in the test set and I don't have a model to give it a shot with.</li>\n</ul>\n<hr>\n<p>I would normally put together a big post showing the evidence to support this… but I don't have time and if it's being fixed soon I don't see the value. That being said, I thought I would share.</p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hi there,\n\nI will throw this out there... not sure how much value it has as the organizers will be fixing things soon (hopefully).\n\nIf a study has multiple images, you can usually tell the image that you should predict the bounding boxes for by looking at the dicom metadata field **`SeriesNumber`**. \n\n**Whichever image has the lowest** **`SeriesNumber`** **in the study is the one that you will need to predict bounding boxes on.**\n\nI confirmed this theory using many studies and it holds true for most of them.\n\n---\n\n**CAVEATS**\n\n* Some images don't have a listed **`SeriesNumber`** (**`None`**). There can be multiple images in a study with this value (**`None`**) and as such we can't tell which image is the one we should predict on.\n* There are odd cases where the above rule doesn't hold true\n* I have no idea if this will hold true in the test set and I don't have a model to give it a shot with.\n\n--- \n\nI would normally put together a big post showing the evidence to support this... but I don't have time and if it's being fixed soon I don't see the value. That being said, I thought I would share.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "1332012",
      "postDate": "06/01/2021 21:03:22",
      "content": "<p>So, it's kind of leak then?</p>",
      "rawMarkdown": "So, it's kind of leak then?",
      "votes": null
    },
    {
      "id": "1332014",
      "postDate": "06/01/2021 21:08:26",
      "content": "<p>Not really. Well kinda. It would be considered a leak if the organizers never fix it… because for those that know the trick they <strong><em>may</em></strong> be able to get a better score for <strong>image level</strong> predictions.</p>\n<p>My understanding is this:</p>\n<ol>\n<li>For a given study we predict one of four labels</li>\n<li>For a given study, we will also bound opacities on a single anterior radiograph. The other images in the study may be included in the <strong>image</strong> set predictions, but they should have no bounding boxes predicted (even though the images may be identical).</li>\n<li>In studies with multiple anterior radiographs (sometimes exact duplicates) we cannot figure out why we are bounding opacities on one image and not another… </li>\n</ol>\n<p>The technique above gives us a glimpse into why they pick one image over another to bound opacities with (the <strong><code>SeriesNumber</code></strong> metadata). I'd guess they picked the <strong><em>first</em></strong> image in the study. First, in this case, is probably the first image taken in the <strong><em>series</em></strong>… hence the lowest <strong><code>SeriesNumber</code></strong> often coinciding with the requirement to have opacities bounded.</p>\n<p>Hope this makes sense. Sorry if my terminology is a bit rough. </p>",
      "rawMarkdown": "Not really. Well kinda. It would be considered a leak if the organizers never fix it... because for those that know the trick they ***may*** be able to get a better score for **image level** predictions.\n\nMy understanding is this:\n\n1. For a given study we predict one of four labels\n2. For a given study, we will also bound opacities on a single anterior radiograph. The other images in the study may be included in the **image** set predictions, but they should have no bounding boxes predicted (even though the images may be identical).\n3. In studies with multiple anterior radiographs (sometimes exact duplicates) we cannot figure out why we are bounding opacities on one image and not another... \n\nThe technique above gives us a glimpse into why they pick one image over another to bound opacities with (the **`SeriesNumber`** metadata). I'd guess they picked the ***first*** image in the study. First, in this case, is probably the first image taken in the ***series***... hence the lowest **`SeriesNumber`** often coinciding with the requirement to have opacities bounded.\n\nHope this makes sense. Sorry if my terminology is a bit rough.",
      "votes": null
    },
    {
      "id": "1332019",
      "postDate": "06/01/2021 21:14:01",
      "content": "<p>nice finding :D</p>",
      "rawMarkdown": "nice finding :D",
      "votes": null
    },
    {
      "id": "1337965",
      "postDate": "06/06/2021 04:27:49",
      "content": "<p>Thanks a lot for your information.</p>\n<ul>\n<li>I checked how many training data are true in this <a href=\"https://www.kaggle.com/tt195361/siim-covid-19-eda-for-train-csv-files\" target=\"_blank\">notebook</a> at \"DICOM Series Number\". For total 6054 studies, true samples are 6046, and false are 8.</li>\n<li>I trained a model for study level, then submitted the result. The LB score is improved from 0.368 to 0.374.<ul>\n<li>I assigned study label to only lowest series number image. The label assigned to the other images is 'Negative'.</li>\n<li>For inference, I used the result from the lowest series number image.</li></ul></li>\n</ul>",
      "rawMarkdown": "Thanks a lot for your information.\n- I checked how many training data are true in this [notebook](https://www.kaggle.com/tt195361/siim-covid-19-eda-for-train-csv-files) at \"DICOM Series Number\". For total 6054 studies, true samples are 6046, and false are 8.\n- I trained a model for study level, then submitted the result. The LB score is improved from 0.368 to 0.374.\n    - I assigned study label to only lowest series number image. The label assigned to the other images is 'Negative'.\n    - For inference, I used the result from the lowest series number image.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1332012,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "06/01/2021 21:03:22",
      "content": "<p>So, it's kind of leak then?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332014,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "06/01/2021 21:08:26",
          "content": "<p>Not really. Well kinda. It would be considered a leak if the organizers never fix it… because for those that know the trick they <strong><em>may</em></strong> be able to get a better score for <strong>image level</strong> predictions.</p>\n<p>My understanding is this:</p>\n<ol>\n<li>For a given study we predict one of four labels</li>\n<li>For a given study, we will also bound opacities on a single anterior radiograph. The other images in the study may be included in the <strong>image</strong> set predictions, but they should have no bounding boxes predicted (even though the images may be identical).</li>\n<li>In studies with multiple anterior radiographs (sometimes exact duplicates) we cannot figure out why we are bounding opacities on one image and not another… </li>\n</ol>\n<p>The technique above gives us a glimpse into why they pick one image over another to bound opacities with (the <strong><code>SeriesNumber</code></strong> metadata). I'd guess they picked the <strong><em>first</em></strong> image in the study. First, in this case, is probably the first image taken in the <strong><em>series</em></strong>… hence the lowest <strong><code>SeriesNumber</code></strong> often coinciding with the requirement to have opacities bounded.</p>\n<p>Hope this makes sense. Sorry if my terminology is a bit rough. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332019,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/01/2021 21:14:01",
          "content": "<p>nice finding :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337965,
      "author_name": "tt195361",
      "author_url": "",
      "post_date": "06/06/2021 04:27:49",
      "content": "<p>Thanks a lot for your information.</p>\n<ul>\n<li>I checked how many training data are true in this <a href=\"https://www.kaggle.com/tt195361/siim-covid-19-eda-for-train-csv-files\" target=\"_blank\">notebook</a> at \"DICOM Series Number\". For total 6054 studies, true samples are 6046, and false are 8.</li>\n<li>I trained a model for study level, then submitted the result. The LB score is improved from 0.368 to 0.374.<ul>\n<li>I assigned study label to only lowest series number image. The label assigned to the other images is 'Negative'.</li>\n<li>For inference, I used the result from the lowest series number image.</li></ul></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1332006": "Hi there,\n\nI will throw this out there... not sure how much value it has as the organizers will be fixing things soon (hopefully).\n\nIf a study has multiple images, you can usually tell the image that you should predict the bounding boxes for by looking at the dicom metadata field **`SeriesNumber`**. \n\n**Whichever image has the lowest** **`SeriesNumber`** **in the study is the one that you will need to predict bounding boxes on.**\n\nI confirmed this theory using many studies and it holds true for most of them.\n\n---\n\n**CAVEATS**\n\n* Some images don't have a listed **`SeriesNumber`** (**`None`**). There can be multiple images in a study with this value (**`None`**) and as such we can't tell which image is the one we should predict on.\n* There are odd cases where the above rule doesn't hold true\n* I have no idea if this will hold true in the test set and I don't have a model to give it a shot with.\n\n--- \n\nI would normally put together a big post showing the evidence to support this... but I don't have time and if it's being fixed soon I don't see the value. That being said, I thought I would share.\n\nHope this helps!",
    "1332012": "So, it's kind of leak then?",
    "1332014": "Not really. Well kinda. It would be considered a leak if the organizers never fix it... because for those that know the trick they ***may*** be able to get a better score for **image level** predictions.\n\nMy understanding is this:\n\n1. For a given study we predict one of four labels\n2. For a given study, we will also bound opacities on a single anterior radiograph. The other images in the study may be included in the **image** set predictions, but they should have no bounding boxes predicted (even though the images may be identical).\n3. In studies with multiple anterior radiographs (sometimes exact duplicates) we cannot figure out why we are bounding opacities on one image and not another... \n\nThe technique above gives us a glimpse into why they pick one image over another to bound opacities with (the **`SeriesNumber`** metadata). I'd guess they picked the ***first*** image in the study. First, in this case, is probably the first image taken in the ***series***... hence the lowest **`SeriesNumber`** often coinciding with the requirement to have opacities bounded.\n\nHope this makes sense. Sorry if my terminology is a bit rough.",
    "1332019": "nice finding :D",
    "1337965": "Thanks a lot for your information.\n- I checked how many training data are true in this [notebook](https://www.kaggle.com/tt195361/siim-covid-19-eda-for-train-csv-files) at \"DICOM Series Number\". For total 6054 studies, true samples are 6046, and false are 8.\n- I trained a model for study level, then submitted the result. The LB score is improved from 0.368 to 0.374.\n    - I assigned study label to only lowest series number image. The label assigned to the other images is 'Negative'.\n    - For inference, I used the result from the lowest series number image."
  },
  "source": "meta"
}