{
  "id": 241238,
  "title": "please clarify my understanding of this competition",
  "url": "/competitions/siim-covid19-detection/discussion/241238",
  "author_name": "Jacek Poplawski",
  "post_date": "2021-05-23T17:08:46.765000",
  "votes": 10,
  "comment_count": 23,
  "views": 0,
  "content": "<p>This is what I understand after reading the competition description and some discussions:</p>\n<ol>\n<li>For each image we can have 0, 1 or 2 boxes. Each box is defined as pixel coordinates on the image.</li>\n<li>0 boxes means there are no lung opacities and no atypical findings.</li>\n<li>2 boxes means there are two long opacities or atypical findings - does order of boxes matter? is score same for both ways?</li>\n<li>Each image is classified as one of four category: Typical Appearance, Indeterminate Appearance, Atypical Appearance, Negative for Pneumonia, result of this classification is not in the image row but in the study row</li>\n<li>Each study contains 1 or more images. We should submit category for each study and the bounding box is always 0 0 1 1 (which is understable because we don't tell which image we mean). It is possible to submit multiple categories for each study - does order of categories matter?</li>\n</ol>\n<p>What do we need:</p>\n<ul>\n<li>model A to find bounding boxes on the image (0, 1 or 2)</li>\n<li>model B to find category for the image (4 labels possible)</li>\n</ul>\n<p>But do we need:</p>\n<ul>\n<li>model C to find categories for the study?</li>\n</ul>\n<p>Or should we just add categories from the images and submit that as study categories?</p>\n<p>Additional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?</p>\n<p>Could you share the code to calculate score?</p>",
  "messages": [
    {
      "id": 1320021,
      "postDate": "2021-05-23T17:08:46.767Z",
      "content": "<p>This is what I understand after reading the competition description and some discussions:</p>\n<ol>\n<li>For each image we can have 0, 1 or 2 boxes. Each box is defined as pixel coordinates on the image.</li>\n<li>0 boxes means there are no lung opacities and no atypical findings.</li>\n<li>2 boxes means there are two long opacities or atypical findings - does order of boxes matter? is score same for both ways?</li>\n<li>Each image is classified as one of four category: Typical Appearance, Indeterminate Appearance, Atypical Appearance, Negative for Pneumonia, result of this classification is not in the image row but in the study row</li>\n<li>Each study contains 1 or more images. We should submit category for each study and the bounding box is always 0 0 1 1 (which is understable because we don't tell which image we mean). It is possible to submit multiple categories for each study - does order of categories matter?</li>\n</ol>\n<p>What do we need:</p>\n<ul>\n<li>model A to find bounding boxes on the image (0, 1 or 2)</li>\n<li>model B to find category for the image (4 labels possible)</li>\n</ul>\n<p>But do we need:</p>\n<ul>\n<li>model C to find categories for the study?</li>\n</ul>\n<p>Or should we just add categories from the images and submit that as study categories?</p>\n<p>Additional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?</p>\n<p>Could you share the code to calculate score?</p>",
      "rawMarkdown": "This is what I understand after reading the competition description and some discussions:\n\n1. For each image we can have 0, 1 or 2 boxes. Each box is defined as pixel coordinates on the image.\n2. 0 boxes means there are no lung opacities and no atypical findings.\n3. 2 boxes means there are two long opacities or atypical findings - does order of boxes matter? is score same for both ways?\n4. Each image is classified as one of four category: Typical Appearance, Indeterminate Appearance, Atypical Appearance, Negative for Pneumonia, result of this classification is not in the image row but in the study row\n5. Each study contains 1 or more images. We should submit category for each study and the bounding box is always 0 0 1 1 (which is understable because we don't tell which image we mean). It is possible to submit multiple categories for each study - does order of categories matter?\n\nWhat do we need:\n- model A to find bounding boxes on the image (0, 1 or 2)\n- model B to find category for the image (4 labels possible)\n\nBut do we need:\n- model C to find categories for the study?\n\nOr should we just add categories from the images and submit that as study categories?\n\nAdditional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?\n\nCould you share the code to calculate score?\n\n\n\n",
      "votes": 10
    },
    {
      "id": 1320040,
      "postDate": "2021-05-23T17:27:08.607Z",
      "content": "<p>Thank you for the topic.<br>\nAlmost the same as my comprehension except for one thing.<br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240250</a><br>\nAccording to the post from the host, there is no BBox on pleural effusion or pneumothorax, which are included in \"Atypical\". <br>\nSo I think 'no BBox image' is possible if there is no opacities, or just pleural effusion or pneumothorax.</p>",
      "rawMarkdown": "Thank you for the topic.\nAlmost the same as my comprehension except for one thing.\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/240250\nAccording to the post from the host, there is no BBox on pleural effusion or pneumothorax, which are included in \"Atypical\". \nSo I think 'no BBox image' is possible if there is no opacities, or just pleural effusion or pneumothorax.",
      "votes": 3,
      "replies": [
        {
          "id": 1320077,
          "postDate": "2021-05-23T18:12:43.490Z",
          "content": "<p>Yes, I read this topic too, I agree boxes are not on all atypical. However I wonder what are the answers for the questions, I understand you are confused by this competition too? It looks interesting but I need to understand what is going on here :)</p>",
          "rawMarkdown": "Yes, I read this topic too, I agree boxes are not on all atypical. However I wonder what are the answers for the questions, I understand you are confused by this competition too? It looks interesting but I need to understand what is going on here :)",
          "votes": 2
        },
        {
          "id": 1320088,
          "postDate": "2021-05-23T18:23:02.847Z",
          "content": "<p>Yes, very confused.<br>\nFor BBox inference, my model (maybe our models?) looks like LB score 0.001, so now I cannot answer the question.<br>\nFor study level inference, I guess the order is irrelevant. I just added predicted labels from images (and calculate the mean per study) and submitted.</p>",
          "rawMarkdown": "Yes, very confused.\nFor BBox inference, my model (maybe our models?) looks like LB score 0.001, so now I cannot answer the question.\nFor study level inference, I guess the order is irrelevant. I just added predicted labels from images (and calculate the mean per study) and submitted.",
          "votes": 1
        },
        {
          "id": 1320110,
          "postDate": "2021-05-23T18:42:23.543Z",
          "content": "<p>Do you have any threshold for the label prediction or do you always submit all 4 labels?<br>\nAdded also this to the post.</p>",
          "rawMarkdown": "Do you have any threshold for the label prediction or do you always submit all 4 labels?\nAdded also this to the post."
        },
        {
          "id": 1320256,
          "postDate": "2021-05-24T00:14:31.150Z",
          "content": "<p>No threshold (actually 1e-15) because metric is mAP :)</p>",
          "rawMarkdown": "No threshold (actually 1e-15) because metric is mAP :)"
        },
        {
          "id": 1320263,
          "postDate": "2021-05-24T00:36:45.317Z",
          "content": "<p>so in your submission there are always 4 labels?</p>",
          "rawMarkdown": "so in your submission there are always 4 labels?"
        },
        {
          "id": 1320272,
          "postDate": "2021-05-24T00:44:06.317Z",
          "content": "<p>Yeah, like the public notebook, although I'm a PyTorch user.</p>",
          "rawMarkdown": "Yeah, like the public notebook, although I'm a PyTorch user."
        },
        {
          "id": 1320280,
          "postDate": "2021-05-24T01:07:46.917Z",
          "content": "<p>I love PyTorch, why do you mention PyTorch here? do you mean there are some disadvantages?</p>",
          "rawMarkdown": "I love PyTorch, why do you mention PyTorch here? do you mean there are some disadvantages?"
        },
        {
          "id": 1320287,
          "postDate": "2021-05-24T01:21:50.917Z",
          "content": "<p>No, there seems to be many TF users in the LB (same scores lol), so now I'm trying to beat them with my own way :)</p>",
          "rawMarkdown": "No, there seems to be many TF users in the LB (same scores lol), so now I'm trying to beat them with my own way :)"
        },
        {
          "id": 1320293,
          "postDate": "2021-05-24T01:28:16.780Z",
          "content": "<p>do you use GPU or TPU? do you train model also on local machine?</p>",
          "rawMarkdown": "do you use GPU or TPU? do you train model also on local machine?",
          "votes": 1
        },
        {
          "id": 1324252,
          "postDate": "2021-05-26T18:29:22.687Z",
          "content": "<p>If submissions are scored using regular PASCAL VOC function then my intuition would be the following. Since we are submitting studies and images as separate rows/ids metric won’t care whether an image belongs to a specific study, it will consider each row as a separate id for ground truths. For studies we have 4 options for classes and for images we have 2 (opacity and none). Again, metric function doesn’t care if a negative study have images predicted as none or vice versa it will consider all these 4+2 options as 6 different classes for matching the ground truth with predictions. One can build 4-class classification model to predict study level and a binary opacity detector for image level predictions. You can also do a single object detection model by using classes from studies for image bboxes (which I did) but this requires some post processing to conver image bbox predictions to study level preds, which I think is not ideal.</p>\n<p>Real mystery is why image level predictilns only add 0.001 to most LB scores.</p>\n<p>Correct me if my intuition regarding the metric is flawed.</p>",
          "rawMarkdown": "If submissions are scored using regular PASCAL VOC function then my intuition would be the following. Since we are submitting studies and images as separate rows/ids metric won’t care whether an image belongs to a specific study, it will consider each row as a separate id for ground truths. For studies we have 4 options for classes and for images we have 2 (opacity and none). Again, metric function doesn’t care if a negative study have images predicted as none or vice versa it will consider all these 4+2 options as 6 different classes for matching the ground truth with predictions. One can build 4-class classification model to predict study level and a binary opacity detector for image level predictions. You can also do a single object detection model by using classes from studies for image bboxes (which I did) but this requires some post processing to conver image bbox predictions to study level preds, which I think is not ideal.\n\nReal mystery is why image level predictilns only add 0.001 to most LB scores.\n\nCorrect me if my intuition regarding the metric is flawed."
        },
        {
          "id": 1324265,
          "postDate": "2021-05-26T18:44:57.780Z",
          "content": "<p>\"Real mystery is why image level predictilns only add 0.001 to most LB scores.\"</p>\n<p>Do you mean the difference between no image predictions and all image predictions is so small? Maybe it means they are predicted incorrectly, we need bounding boxes, right? So it is not just label.</p>",
          "rawMarkdown": "\"Real mystery is why image level predictilns only add 0.001 to most LB scores.\"\n\nDo you mean the difference between no image predictions and all image predictions is so small? Maybe it means they are predicted incorrectly, we need bounding boxes, right? So it is not just label."
        },
        {
          "id": 1324281,
          "postDate": "2021-05-26T19:10:16.527Z",
          "content": "<p>Yes. I meant the difference between only study vs study+image preds. There are many discussions regarding that. Image and study level preds are independent in terms of score calculation. There is however a dependence when someone post-processes an image to have none 0 0 1 1 prediction if that image belongs to a study which is predicted negative by classification model.</p>",
          "rawMarkdown": "Yes. I meant the difference between only study vs study+image preds. There are many discussions regarding that. Image and study level preds are independent in terms of score calculation. There is however a dependence when someone post-processes an image to have none 0 0 1 1 prediction if that image belongs to a study which is predicted negative by classification model."
        }
      ]
    },
    {
      "id": 1321404,
      "postDate": "2021-05-24T17:03:55.520Z",
      "content": "<p>Hi! The code to calculate the score is essentially the same as the code provided in the VOC 2010 / 2012 devkits, but in C# form. The C# code itself is unfortunately extremely complex.</p>\n<p>If you predict a label that does not exist in a study's ground truth (i.e. labels with low confidence) it will count as a false positive. While predicting all labels may provide a score improvement over predicting single labels that may be wrong, to maximize your score you'll need to eliminate those false positives.</p>",
      "rawMarkdown": "Hi! The code to calculate the score is essentially the same as the code provided in the VOC 2010 / 2012 devkits, but in C# form. The C# code itself is unfortunately extremely complex.\n\nIf you predict a label that does not exist in a study's ground truth (i.e. labels with low confidence) it will count as a false positive. While predicting all labels may provide a score improvement over predicting single labels that may be wrong, to maximize your score you'll need to eliminate those false positives.",
      "votes": 4,
      "replies": [
        {
          "id": 1321450,
          "postDate": "2021-05-24T17:30:33.860Z",
          "content": "<p>hello Phil, thanks for the answer,</p>\n<p>please clarify \"to maximize your score you'll need to eliminate those false positives\", does it mean I should use threshold and for instance when confidence is small we should change it to zero? or should we do a vote and choose only best classes?</p>\n<p>it is not clear to me how many labels can be in the ground truth, if only one we could just take best confidence but after reading competition description and discussions I assumed each study have multiple labels and in this case submission should always contain all 4 labels</p>\n<p>first example: label A is 0.8, label B is 0.7, label C is 0.6 <br>\nshould we submit A 0.8, B 0.7, C 0.6, D 0.0?<br>\nor should we say two classes are max and just submit A 0.8, B 0.7</p>\n<p>second example: label A 0.9 label B 0.1 label C 0.01<br>\nshould we submit: A 0.9, B 0.1, C 0.01, D 0.0?<br>\nor should we say it must be A and submit A 0.9?</p>",
          "rawMarkdown": "hello Phil, thanks for the answer,\n\nplease clarify \"to maximize your score you'll need to eliminate those false positives\", does it mean I should use threshold and for instance when confidence is small we should change it to zero? or should we do a vote and choose only best classes?\n\nit is not clear to me how many labels can be in the ground truth, if only one we could just take best confidence but after reading competition description and discussions I assumed each study have multiple labels and in this case submission should always contain all 4 labels\n\nfirst example: label A is 0.8, label B is 0.7, label C is 0.6 \nshould we submit A 0.8, B 0.7, C 0.6, D 0.0?\nor should we say two classes are max and just submit A 0.8, B 0.7\n\nsecond example: label A 0.9 label B 0.1 label C 0.01\nshould we submit: A 0.9, B 0.1, C 0.01, D 0.0?\nor should we say it must be A and submit A 0.9?\n"
        },
        {
          "id": 1321481,
          "postDate": "2021-05-24T17:50:44.023Z",
          "content": "<p>Hi Jacek - for studies you should ideally be making one correct prediction. Making more, regardless of confidence score, will result in false positives counting against your overall score. For images you should predict as many bounding boxes as you think exist based on the training data - but once again, false positives (bounding boxes with no matching ground truth) will be counted against your overall score.</p>",
          "rawMarkdown": "Hi Jacek - for studies you should ideally be making one correct prediction. Making more, regardless of confidence score, will result in false positives counting against your overall score. For images you should predict as many bounding boxes as you think exist based on the training data - but once again, false positives (bounding boxes with no matching ground truth) will be counted against your overall score.",
          "votes": 4
        },
        {
          "id": 1321590,
          "postDate": "2021-05-24T19:25:11.753Z",
          "content": "<p>Does it mean each study should have only one label? I think I read that multiple labels are possible in test, so looks like it was just a rumour?</p>",
          "rawMarkdown": "Does it mean each study should have only one label? I think I read that multiple labels are possible in test, so looks like it was just a rumour?"
        },
        {
          "id": 1322748,
          "postDate": "2021-05-25T16:39:32.553Z",
          "content": "<p>Hi Jacek - all studies should have one study-level label.</p>",
          "rawMarkdown": "Hi Jacek - all studies should have one study-level label.",
          "votes": 3
        },
        {
          "id": 1322802,
          "postDate": "2021-05-25T17:31:57.947Z",
          "content": "<p>Thank you, this changes my understanding of the competition.</p>",
          "rawMarkdown": "Thank you, this changes my understanding of the competition."
        }
      ]
    },
    {
      "id": 1320587,
      "postDate": "2021-05-24T07:19:08.637Z",
      "content": "<pre><code>Additional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?\n\nand another question about sorting the predictions by confidence\n</code></pre>\n<p>When calculating mAP:</p>\n<ol>\n<li>They sort the prediction in descending order of confidence. so even if you sort your submission file it wont affect much</li>\n<li>for the Additional Question, Having 0.0001 confidence in study level image for a label, only improves your score. In my last win in the VinBigData Xray Competition ( here is the <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/231511\" target=\"_blank\">writeup</a> ), it was a key to improve our submission. Having low confidence prediction will improve your recall hence improve your score</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> had an implementation for the metric<br>\nhere is it <a href=\"https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\" target=\"_blank\">https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes</a></p>\n<p>I hope you guys got it?. Correct me if I am wrong</p>",
      "rawMarkdown": "```\nAdditional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?\n\nand another question about sorting the predictions by confidence\n\n```\nWhen calculating mAP:\n1. They sort the prediction in descending order of confidence. so even if you sort your submission file it wont affect much\n2. for the Additional Question, Having 0.0001 confidence in study level image for a label, only improves your score. In my last win in the VinBigData Xray Competition ( here is the [writeup](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/231511) ), it was a key to improve our submission. Having low confidence prediction will improve your recall hence improve your score\n\n@zfturbo had an implementation for the metric\nhere is it https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\n\nI hope you guys got it?. Correct me if I am wrong",
      "votes": 1,
      "replies": [
        {
          "id": 1336946,
          "postDate": "2021-06-05T09:57:30.037Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> during submission are we predicting classes also..?<br>\namong 4 .<br>\nor Just opacity or None</p>",
          "rawMarkdown": "@morizin during submission are we predicting classes also..?\namong 4 .\nor Just opacity or None",
          "votes": 1
        },
        {
          "id": 1337026,
          "postDate": "2021-06-05T11:07:40.150Z",
          "content": "<p>For image level ( detection ) I think only two classes (opacity or None) and for Study level ( Multi Label ) there are classes. </p>",
          "rawMarkdown": "For image level ( detection ) I think only two classes (opacity or None) and for Study level ( Multi Label ) there are classes. "
        }
      ]
    },
    {
      "id": 1331266,
      "postDate": "2021-06-01T11:22:05.917Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1320040,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-05-23T17:27:08.607000",
      "content": "<p>Thank you for the topic.<br>\nAlmost the same as my comprehension except for one thing.<br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240250</a><br>\nAccording to the post from the host, there is no BBox on pleural effusion or pneumothorax, which are included in \"Atypical\". <br>\nSo I think 'no BBox image' is possible if there is no opacities, or just pleural effusion or pneumothorax.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1320077,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-23T18:12:43.490000",
          "content": "<p>Yes, I read this topic too, I agree boxes are not on all atypical. However I wonder what are the answers for the questions, I understand you are confused by this competition too? It looks interesting but I need to understand what is going on here :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1320088,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-23T18:23:02.847000",
          "content": "<p>Yes, very confused.<br>\nFor BBox inference, my model (maybe our models?) looks like LB score 0.001, so now I cannot answer the question.<br>\nFor study level inference, I guess the order is irrelevant. I just added predicted labels from images (and calculate the mean per study) and submitted.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1320110,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-23T18:42:23.543000",
          "content": "<p>Do you have any threshold for the label prediction or do you always submit all 4 labels?<br>\nAdded also this to the post.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320256,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-24T00:14:31.150000",
          "content": "<p>No threshold (actually 1e-15) because metric is mAP :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320263,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-24T00:36:45.317000",
          "content": "<p>so in your submission there are always 4 labels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320272,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-24T00:44:06.317000",
          "content": "<p>Yeah, like the public notebook, although I'm a PyTorch user.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320280,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-24T01:07:46.917000",
          "content": "<p>I love PyTorch, why do you mention PyTorch here? do you mean there are some disadvantages?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320287,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-24T01:21:50.917000",
          "content": "<p>No, there seems to be many TF users in the LB (same scores lol), so now I'm trying to beat them with my own way :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320293,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-24T01:28:16.780000",
          "content": "<p>do you use GPU or TPU? do you train model also on local machine?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1324252,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-05-26T18:29:22.687000",
          "content": "<p>If submissions are scored using regular PASCAL VOC function then my intuition would be the following. Since we are submitting studies and images as separate rows/ids metric won’t care whether an image belongs to a specific study, it will consider each row as a separate id for ground truths. For studies we have 4 options for classes and for images we have 2 (opacity and none). Again, metric function doesn’t care if a negative study have images predicted as none or vice versa it will consider all these 4+2 options as 6 different classes for matching the ground truth with predictions. One can build 4-class classification model to predict study level and a binary opacity detector for image level predictions. You can also do a single object detection model by using classes from studies for image bboxes (which I did) but this requires some post processing to conver image bbox predictions to study level preds, which I think is not ideal.</p>\n<p>Real mystery is why image level predictilns only add 0.001 to most LB scores.</p>\n<p>Correct me if my intuition regarding the metric is flawed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1324265,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-26T18:44:57.780000",
          "content": "<p>\"Real mystery is why image level predictilns only add 0.001 to most LB scores.\"</p>\n<p>Do you mean the difference between no image predictions and all image predictions is so small? Maybe it means they are predicted incorrectly, we need bounding boxes, right? So it is not just label.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1324281,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-05-26T19:10:16.527000",
          "content": "<p>Yes. I meant the difference between only study vs study+image preds. There are many discussions regarding that. Image and study level preds are independent in terms of score calculation. There is however a dependence when someone post-processes an image to have none 0 0 1 1 prediction if that image belongs to a study which is predicted negative by classification model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1321404,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2021-05-24T17:03:55.520000",
      "content": "<p>Hi! The code to calculate the score is essentially the same as the code provided in the VOC 2010 / 2012 devkits, but in C# form. The C# code itself is unfortunately extremely complex.</p>\n<p>If you predict a label that does not exist in a study's ground truth (i.e. labels with low confidence) it will count as a false positive. While predicting all labels may provide a score improvement over predicting single labels that may be wrong, to maximize your score you'll need to eliminate those false positives.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1321450,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-24T17:30:33.860000",
          "content": "<p>hello Phil, thanks for the answer,</p>\n<p>please clarify \"to maximize your score you'll need to eliminate those false positives\", does it mean I should use threshold and for instance when confidence is small we should change it to zero? or should we do a vote and choose only best classes?</p>\n<p>it is not clear to me how many labels can be in the ground truth, if only one we could just take best confidence but after reading competition description and discussions I assumed each study have multiple labels and in this case submission should always contain all 4 labels</p>\n<p>first example: label A is 0.8, label B is 0.7, label C is 0.6 <br>\nshould we submit A 0.8, B 0.7, C 0.6, D 0.0?<br>\nor should we say two classes are max and just submit A 0.8, B 0.7</p>\n<p>second example: label A 0.9 label B 0.1 label C 0.01<br>\nshould we submit: A 0.9, B 0.1, C 0.01, D 0.0?<br>\nor should we say it must be A and submit A 0.9?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321481,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2021-05-24T17:50:44.023000",
          "content": "<p>Hi Jacek - for studies you should ideally be making one correct prediction. Making more, regardless of confidence score, will result in false positives counting against your overall score. For images you should predict as many bounding boxes as you think exist based on the training data - but once again, false positives (bounding boxes with no matching ground truth) will be counted against your overall score.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1321590,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-24T19:25:11.753000",
          "content": "<p>Does it mean each study should have only one label? I think I read that multiple labels are possible in test, so looks like it was just a rumour?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1322748,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2021-05-25T16:39:32.553000",
          "content": "<p>Hi Jacek - all studies should have one study-level label.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1322802,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2021-05-25T17:31:57.947000",
          "content": "<p>Thank you, this changes my understanding of the competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1320587,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-05-24T07:19:08.637000",
      "content": "<pre><code>Additional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?\n\nand another question about sorting the predictions by confidence\n</code></pre>\n<p>When calculating mAP:</p>\n<ol>\n<li>They sort the prediction in descending order of confidence. so even if you sort your submission file it wont affect much</li>\n<li>for the Additional Question, Having 0.0001 confidence in study level image for a label, only improves your score. In my last win in the VinBigData Xray Competition ( here is the <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/231511\" target=\"_blank\">writeup</a> ), it was a key to improve our submission. Having low confidence prediction will improve your recall hence improve your score</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> had an implementation for the metric<br>\nhere is it <a href=\"https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\" target=\"_blank\">https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes</a></p>\n<p>I hope you guys got it?. Correct me if I am wrong</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1336946,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-06-05T09:57:30.037000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> during submission are we predicting classes also..?<br>\namong 4 .<br>\nor Just opacity or None</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1337026,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-06-05T11:07:40.150000",
          "content": "<p>For image level ( detection ) I think only two classes (opacity or None) and for Study level ( Multi Label ) there are classes. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1331266,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-01T11:22:05.917000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1320021": "This is what I understand after reading the competition description and some discussions:\n\n1. For each image we can have 0, 1 or 2 boxes. Each box is defined as pixel coordinates on the image.\n2. 0 boxes means there are no lung opacities and no atypical findings.\n3. 2 boxes means there are two long opacities or atypical findings - does order of boxes matter? is score same for both ways?\n4. Each image is classified as one of four category: Typical Appearance, Indeterminate Appearance, Atypical Appearance, Negative for Pneumonia, result of this classification is not in the image row but in the study row\n5. Each study contains 1 or more images. We should submit category for each study and the bounding box is always 0 0 1 1 (which is understable because we don't tell which image we mean). It is possible to submit multiple categories for each study - does order of categories matter?\n\nWhat do we need:\n- model A to find bounding boxes on the image (0, 1 or 2)\n- model B to find category for the image (4 labels possible)\n\nBut do we need:\n- model C to find categories for the study?\n\nOr should we just add categories from the images and submit that as study categories?\n\nAdditional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?\n\nCould you share the code to calculate score?\n\n\n\n",
    "1320040": "Thank you for the topic.\nAlmost the same as my comprehension except for one thing.\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/240250\nAccording to the post from the host, there is no BBox on pleural effusion or pneumothorax, which are included in \"Atypical\". \nSo I think 'no BBox image' is possible if there is no opacities, or just pleural effusion or pneumothorax.",
    "1321404": "Hi! The code to calculate the score is essentially the same as the code provided in the VOC 2010 / 2012 devkits, but in C# form. The C# code itself is unfortunately extremely complex.\n\nIf you predict a label that does not exist in a study's ground truth (i.e. labels with low confidence) it will count as a false positive. While predicting all labels may provide a score improvement over predicting single labels that may be wrong, to maximize your score you'll need to eliminate those false positives.",
    "1320587": "```\nAdditional question: should we always predict all 4 labels per study just with different confidences? For example our model could predict something with 0.0001 confidence, should we add it to the submission or skip it? Why?\n\nand another question about sorting the predictions by confidence\n\n```\nWhen calculating mAP:\n1. They sort the prediction in descending order of confidence. so even if you sort your submission file it wont affect much\n2. for the Additional Question, Having 0.0001 confidence in study level image for a label, only improves your score. In my last win in the VinBigData Xray Competition ( here is the [writeup](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/231511) ), it was a key to improve our submission. Having low confidence prediction will improve your recall hence improve your score\n\n@zfturbo had an implementation for the metric\nhere is it https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\n\nI hope you guys got it?. Correct me if I am wrong",
    "1331266": ""
  }
}