{
  "id": 244126,
  "title": "Up to 75 studies in public test have multiple labels",
  "url": "/competitions/siim-covid19-detection/discussion/244126",
  "author_name": "",
  "post_date": "2021-06-05T09:27:24.440619300Z",
  "votes": 12,
  "comment_count": 16,
  "views": 0,
  "content": "<p>If all studies have single labels as in the training dataset, the LB score would be 0.166 if all four classes were submitted with all confidence set to 1, but the actual LB would be 0.176.<br>\nThe evaluation is mean average precision, so if all the confidence scores are the same, this 0.176 is the percentage of positive samples of studies. In the 1214 studies, there were up to 1289 positives.</p>\n<pre><code>1214*0.177*3/2*4 = 1289.268\n</code></pre>\n<p>We don't know the exact number because we don't know the exact LB score and there could be more than three classes in one study, but it is reasonable to assume that about 70 studies have multi label.<br>\nMany of the current notes use only one image per study to predict, but after the data is corrected, all images should be predicted.</p>",
  "messages": [
    {
      "id": "1336910",
      "postDate": "06/05/2021 09:27:24",
      "content": "<p>If all studies have single labels as in the training dataset, the LB score would be 0.166 if all four classes were submitted with all confidence set to 1, but the actual LB would be 0.176.<br>\nThe evaluation is mean average precision, so if all the confidence scores are the same, this 0.176 is the percentage of positive samples of studies. In the 1214 studies, there were up to 1289 positives.</p>\n<pre><code>1214*0.177*3/2*4 = 1289.268\n</code></pre>\n<p>We don't know the exact number because we don't know the exact LB score and there could be more than three classes in one study, but it is reasonable to assume that about 70 studies have multi label.<br>\nMany of the current notes use only one image per study to predict, but after the data is corrected, all images should be predicted.</p>",
      "rawMarkdown": "If all studies have single labels as in the training dataset, the LB score would be 0.166 if all four classes were submitted with all confidence set to 1, but the actual LB would be 0.176.\nThe evaluation is mean average precision, so if all the confidence scores are the same, this 0.176 is the percentage of positive samples of studies. In the 1214 studies, there were up to 1289 positives.\n```\n1214*0.177*3/2*4 = 1289.268\n```\nWe don't know the exact number because we don't know the exact LB score and there could be more than three classes in one study, but it is reasonable to assume that about 70 studies have multi label.\nMany of the current notes use only one image per study to predict, but after the data is corrected, all images should be predicted.",
      "votes": null
    },
    {
      "id": "1337400",
      "postDate": "06/05/2021 15:28:14",
      "content": "<p>I was under the impression that the classes were mutually exclusive. I don't think there is any multi-label sample in study-level images? Also, the description of the classes suggests that they have to be mutually exclusive because the position of the opacity in the x-ray determines the sample's class. </p>\n<p>The position of the opacity except in atypical is different for each label, so if that is the case atypical might not be mutually exclusive when I think about it but I guess more investigation needs to be done.</p>\n<p>Thanks for the investigation and sharing the results!!</p>",
      "rawMarkdown": "I was under the impression that the classes were mutually exclusive. I don't think there is any multi-label sample in study-level images? Also, the description of the classes suggests that they have to be mutually exclusive because the position of the opacity in the x-ray determines the sample's class. \n\nThe position of the opacity except in atypical is different for each label, so if that is the case atypical might not be mutually exclusive when I think about it but I guess more investigation needs to be done.\n\nThanks for the investigation and sharing the results!!",
      "votes": null
    },
    {
      "id": "1337402",
      "postDate": "06/05/2021 15:33:17",
      "content": "<p>I still wonder how this competition works, please look at my topic <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/241238</a><br>\nhere Phil from Kaggle wrote \"Hi Jacek - all studies should have one study-level label.\"</p>",
      "rawMarkdown": "I still wonder how this competition works, please look at my topic https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\nhere Phil from Kaggle wrote \"Hi Jacek - all studies should have one study-level label.\"",
      "votes": null
    },
    {
      "id": "1337709",
      "postDate": "06/05/2021 20:03:49",
      "content": "<p>I am getting <code>0.1715</code> (using pycocotools) <a href=\"https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929647\" target=\"_blank\">here</a> when predicting all classes for each study in training data, which has all single labels. How do you calculate 0.166?</p>",
      "rawMarkdown": "I am getting `0.1715` (using pycocotools) [here](https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929647) when predicting all classes for each study in training data, which has all single labels. How do you calculate 0.166?",
      "votes": null
    },
    {
      "id": "1337771",
      "postDate": "06/05/2021 22:12:42",
      "content": "<p>I used sklearn.metrics.average_precision_score.</p>",
      "rawMarkdown": "I used sklearn.metrics.average_precision_score.",
      "votes": null
    },
    {
      "id": "1337777",
      "postDate": "06/05/2021 22:22:08",
      "content": "<p>Hmmm, it sure says that. Am I taking the results the wrong way?</p>",
      "rawMarkdown": "Hmmm, it sure says that. Am I taking the results the wrong way?",
      "votes": null
    },
    {
      "id": "1337779",
      "postDate": "06/05/2021 22:23:42",
      "content": "<p>The 6 classes of image level are exclusive, but a study may have more than one image.</p>",
      "rawMarkdown": "The 6 classes of image level are exclusive, but a study may have more than one image.",
      "votes": null
    },
    {
      "id": "1337801",
      "postDate": "06/05/2021 22:37:17",
      "content": "<p>I was first tempted to use it too after seeing in discussions. But after reading the documentation it seemed different than PASCAL VOC@0.5. It doesn't seem to use interpolation for precision which is not the case with PASCAL VOC. </p>\n<pre><code>This implementation is not interpolated and is different from computing the area under the precision-recall curve with the trapezoidal rule, which uses linear interpolation and can be too optimistic.\n</code></pre>\n<p>They might be very highly correlated though so should be fine I guess.</p>",
      "rawMarkdown": "I was first tempted to use it too after seeing in discussions. But after reading the documentation it seemed different than PASCAL VOC@0.5. It doesn't seem to use interpolation for precision which is not the case with PASCAL VOC. \n\n```\nThis implementation is not interpolated and is different from computing the area under the precision-recall curve with the trapezoidal rule, which uses linear interpolation and can be too optimistic.\n```\n\nThey might be very highly correlated though so should be fine I guess.",
      "votes": null
    },
    {
      "id": "1337815",
      "postDate": "06/05/2021 22:57:27",
      "content": "<p>They're exclusive classes, so I think it would be natural for it to be 0.1666 at 1/6, but maybe I'm just not understanding it well enough.<br>\nAlso, the evaluation page says that study level is multi label, but the host seems to say that it is single label. It would be nice to get some clarification in this area as well.</p>",
      "rawMarkdown": "They're exclusive classes, so I think it would be natural for it to be 0.1666 at 1/6, but maybe I'm just not understanding it well enough.\nAlso, the evaluation page says that study level is multi label, but the host seems to say that it is single label. It would be nice to get some clarification in this area as well.",
      "votes": null
    },
    {
      "id": "1337825",
      "postDate": "06/05/2021 23:19:49",
      "content": "<p>Yes, I agree its a bit confusing :D But I think latest confirmation was that classes are mutually exclusive, most studies anyway have 1 image and the ones with multiple images (&gt;1) are just different images taken from the same patient while they are diagnosed as same class. So the way I look at it is that its not a multilabel classification problem and we have free data augmentations for some studies :D</p>\n<p>FYI. sklearn and coco gives similar results on my local validation. </p>\n<p>For example, after minor training:</p>\n<pre><code>sklearn AP per 4 class: [0.5348972510026182, 0.7081703557338961, 0.23515254543913836, 0.0991632183601788] (neg,typ,ind,aty)\n\nsklean mAP 0.26289, coco mAP 0.26770\n</code></pre>\n<p>Since AP calculation requires ordering (similar to AUC) and calculating the area under the curve I am not able to wrap my head around and say mAP should be 0.1666.</p>",
      "rawMarkdown": "Yes, I agree its a bit confusing :D But I think latest confirmation was that classes are mutually exclusive, most studies anyway have 1 image and the ones with multiple images (>1) are just different images taken from the same patient while they are diagnosed as same class. So the way I look at it is that its not a multilabel classification problem and we have free data augmentations for some studies :D\n\nFYI. sklearn and coco gives similar results on my local validation. \n\nFor example, after minor training:\n```\nsklearn AP per 4 class: [0.5348972510026182, 0.7081703557338961, 0.23515254543913836, 0.0991632183601788] (neg,typ,ind,aty)\n\nsklean mAP 0.26289, coco mAP 0.26770\n```\n\nSince AP calculation requires ordering (similar to AUC) and calculating the area under the curve I am not able to wrap my head around and say mAP should be 0.1666.",
      "votes": null
    },
    {
      "id": "1337964",
      "postDate": "06/06/2021 04:26:57",
      "content": "<p>I assumed that every study had a single label, like a training dataset. That means you have to think about multiple labels. I am troubled.</p>",
      "rawMarkdown": "I assumed that every study had a single label, like a training dataset. That means you have to think about multiple labels. I am troubled.",
      "votes": null
    },
    {
      "id": "1337988",
      "postDate": "06/06/2021 04:59:24",
      "content": "<p>The current data seems to indicate that all images related to a single study are the same, so the only way to create a post is to assume a single label. If the data is updated, it would be good to consider multi-labeling only then.</p>",
      "rawMarkdown": "The current data seems to indicate that all images related to a single study are the same, so the only way to create a post is to assume a single label. If the data is updated, it would be good to consider multi-labeling only then.",
      "votes": null
    },
    {
      "id": "1338199",
      "postDate": "06/06/2021 08:34:23",
      "content": "<p>okay, are you using softmax or sigmoid in your models?</p>",
      "rawMarkdown": "okay, are you using softmax or sigmoid in your models?",
      "votes": null
    },
    {
      "id": "1338235",
      "postDate": "06/06/2021 09:34:44",
      "content": "<p>Of course Softmax</p>",
      "rawMarkdown": "Of course Softmax",
      "votes": null
    },
    {
      "id": "1338940",
      "postDate": "06/06/2021 21:50:37",
      "content": "<p>Actually, the result you are getting is due to an error on the host's part I came across this discussion <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1322940\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1322940</a> and the hosts did say that there are multiple copies of few images with different labels and they are working on resolving the issue.</p>",
      "rawMarkdown": "Actually, the result you are getting is due to an error on the host's part I came across this discussion https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1322940 and the hosts did say that there are multiple copies of few images with different labels and they are working on resolving the issue.",
      "votes": null
    },
    {
      "id": "1339398",
      "postDate": "06/07/2021 08:11:24",
      "content": "<p>Yes, you're right. I wanted to show the idea, not the specific numbers.</p>",
      "rawMarkdown": "Yes, you're right. I wanted to show the idea, not the specific numbers.",
      "votes": null
    },
    {
      "id": "1339680",
      "postDate": "06/07/2021 11:48:31",
      "content": "<p>yep. all the images below one study are the same. hope there will be some differences assumed that one patient took several different images after upating.</p>",
      "rawMarkdown": "yep. all the images below one study are the same. hope there will be some differences assumed that one patient took several different images after upating.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1337400,
      "author_name": "varundutt9213",
      "author_url": "",
      "post_date": "06/05/2021 15:28:14",
      "content": "<p>I was under the impression that the classes were mutually exclusive. I don't think there is any multi-label sample in study-level images? Also, the description of the classes suggests that they have to be mutually exclusive because the position of the opacity in the x-ray determines the sample's class. </p>\n<p>The position of the opacity except in atypical is different for each label, so if that is the case atypical might not be mutually exclusive when I think about it but I guess more investigation needs to be done.</p>\n<p>Thanks for the investigation and sharing the results!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337779,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/05/2021 22:23:42",
          "content": "<p>The 6 classes of image level are exclusive, but a study may have more than one image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1338199,
          "author_name": "varundutt9213",
          "author_url": "",
          "post_date": "06/06/2021 08:34:23",
          "content": "<p>okay, are you using softmax or sigmoid in your models?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1338235,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/06/2021 09:34:44",
          "content": "<p>Of course Softmax</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1338940,
          "author_name": "varundutt9213",
          "author_url": "",
          "post_date": "06/06/2021 21:50:37",
          "content": "<p>Actually, the result you are getting is due to an error on the host's part I came across this discussion <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1322940\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1322940</a> and the hosts did say that there are multiple copies of few images with different labels and they are working on resolving the issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1339398,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/07/2021 08:11:24",
          "content": "<p>Yes, you're right. I wanted to show the idea, not the specific numbers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337402,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "06/05/2021 15:33:17",
      "content": "<p>I still wonder how this competition works, please look at my topic <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/241238</a><br>\nhere Phil from Kaggle wrote \"Hi Jacek - all studies should have one study-level label.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337777,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/05/2021 22:22:08",
          "content": "<p>Hmmm, it sure says that. Am I taking the results the wrong way?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337709,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "06/05/2021 20:03:49",
      "content": "<p>I am getting <code>0.1715</code> (using pycocotools) <a href=\"https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929647\" target=\"_blank\">here</a> when predicting all classes for each study in training data, which has all single labels. How do you calculate 0.166?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337771,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/05/2021 22:12:42",
          "content": "<p>I used sklearn.metrics.average_precision_score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337801,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "06/05/2021 22:37:17",
          "content": "<p>I was first tempted to use it too after seeing in discussions. But after reading the documentation it seemed different than PASCAL VOC@0.5. It doesn't seem to use interpolation for precision which is not the case with PASCAL VOC. </p>\n<pre><code>This implementation is not interpolated and is different from computing the area under the precision-recall curve with the trapezoidal rule, which uses linear interpolation and can be too optimistic.\n</code></pre>\n<p>They might be very highly correlated though so should be fine I guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337815,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/05/2021 22:57:27",
          "content": "<p>They're exclusive classes, so I think it would be natural for it to be 0.1666 at 1/6, but maybe I'm just not understanding it well enough.<br>\nAlso, the evaluation page says that study level is multi label, but the host seems to say that it is single label. It would be nice to get some clarification in this area as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337825,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "06/05/2021 23:19:49",
          "content": "<p>Yes, I agree its a bit confusing :D But I think latest confirmation was that classes are mutually exclusive, most studies anyway have 1 image and the ones with multiple images (&gt;1) are just different images taken from the same patient while they are diagnosed as same class. So the way I look at it is that its not a multilabel classification problem and we have free data augmentations for some studies :D</p>\n<p>FYI. sklearn and coco gives similar results on my local validation. </p>\n<p>For example, after minor training:</p>\n<pre><code>sklearn AP per 4 class: [0.5348972510026182, 0.7081703557338961, 0.23515254543913836, 0.0991632183601788] (neg,typ,ind,aty)\n\nsklean mAP 0.26289, coco mAP 0.26770\n</code></pre>\n<p>Since AP calculation requires ordering (similar to AUC) and calculating the area under the curve I am not able to wrap my head around and say mAP should be 0.1666.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337964,
      "author_name": "tensorchoko",
      "author_url": "",
      "post_date": "06/06/2021 04:26:57",
      "content": "<p>I assumed that every study had a single label, like a training dataset. That means you have to think about multiple labels. I am troubled.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337988,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/06/2021 04:59:24",
          "content": "<p>The current data seems to indicate that all images related to a single study are the same, so the only way to create a post is to assume a single label. If the data is updated, it would be good to consider multi-labeling only then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1339680,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "06/07/2021 11:48:31",
          "content": "<p>yep. all the images below one study are the same. hope there will be some differences assumed that one patient took several different images after upating.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1336910": "If all studies have single labels as in the training dataset, the LB score would be 0.166 if all four classes were submitted with all confidence set to 1, but the actual LB would be 0.176.\nThe evaluation is mean average precision, so if all the confidence scores are the same, this 0.176 is the percentage of positive samples of studies. In the 1214 studies, there were up to 1289 positives.\n```\n1214*0.177*3/2*4 = 1289.268\n```\nWe don't know the exact number because we don't know the exact LB score and there could be more than three classes in one study, but it is reasonable to assume that about 70 studies have multi label.\nMany of the current notes use only one image per study to predict, but after the data is corrected, all images should be predicted.",
    "1337400": "I was under the impression that the classes were mutually exclusive. I don't think there is any multi-label sample in study-level images? Also, the description of the classes suggests that they have to be mutually exclusive because the position of the opacity in the x-ray determines the sample's class. \n\nThe position of the opacity except in atypical is different for each label, so if that is the case atypical might not be mutually exclusive when I think about it but I guess more investigation needs to be done.\n\nThanks for the investigation and sharing the results!!",
    "1337402": "I still wonder how this competition works, please look at my topic https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\nhere Phil from Kaggle wrote \"Hi Jacek - all studies should have one study-level label.\"",
    "1337709": "I am getting `0.1715` (using pycocotools) [here](https://www.kaggle.com/keremt/competition-metric-map-0-5?scriptVersionId=64929647) when predicting all classes for each study in training data, which has all single labels. How do you calculate 0.166?",
    "1337771": "I used sklearn.metrics.average_precision_score.",
    "1337777": "Hmmm, it sure says that. Am I taking the results the wrong way?",
    "1337779": "The 6 classes of image level are exclusive, but a study may have more than one image.",
    "1337801": "I was first tempted to use it too after seeing in discussions. But after reading the documentation it seemed different than PASCAL VOC@0.5. It doesn't seem to use interpolation for precision which is not the case with PASCAL VOC. \n\n```\nThis implementation is not interpolated and is different from computing the area under the precision-recall curve with the trapezoidal rule, which uses linear interpolation and can be too optimistic.\n```\n\nThey might be very highly correlated though so should be fine I guess.",
    "1337815": "They're exclusive classes, so I think it would be natural for it to be 0.1666 at 1/6, but maybe I'm just not understanding it well enough.\nAlso, the evaluation page says that study level is multi label, but the host seems to say that it is single label. It would be nice to get some clarification in this area as well.",
    "1337825": "Yes, I agree its a bit confusing :D But I think latest confirmation was that classes are mutually exclusive, most studies anyway have 1 image and the ones with multiple images (>1) are just different images taken from the same patient while they are diagnosed as same class. So the way I look at it is that its not a multilabel classification problem and we have free data augmentations for some studies :D\n\nFYI. sklearn and coco gives similar results on my local validation. \n\nFor example, after minor training:\n```\nsklearn AP per 4 class: [0.5348972510026182, 0.7081703557338961, 0.23515254543913836, 0.0991632183601788] (neg,typ,ind,aty)\n\nsklean mAP 0.26289, coco mAP 0.26770\n```\n\nSince AP calculation requires ordering (similar to AUC) and calculating the area under the curve I am not able to wrap my head around and say mAP should be 0.1666.",
    "1337964": "I assumed that every study had a single label, like a training dataset. That means you have to think about multiple labels. I am troubled.",
    "1337988": "The current data seems to indicate that all images related to a single study are the same, so the only way to create a post is to assume a single label. If the data is updated, it would be good to consider multi-labeling only then.",
    "1338199": "okay, are you using softmax or sigmoid in your models?",
    "1338235": "Of course Softmax",
    "1338940": "Actually, the result you are getting is due to an error on the host's part I came across this discussion https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1322940 and the hosts did say that there are multiple copies of few images with different labels and they are working on resolving the issue.",
    "1339398": "Yes, you're right. I wanted to show the idea, not the specific numbers.",
    "1339680": "yep. all the images below one study are the same. hope there will be some differences assumed that one patient took several different images after upating."
  },
  "source": "meta"
}