{
  "id": 244393,
  "title": "Clarifications about the competition",
  "url": "/competitions/siim-covid19-detection/discussion/244393",
  "author_name": "",
  "post_date": "2021-06-06T15:34:29.935092500Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I am new to Kaggle and this is my first real competition. I have few things to clarify about this competition.<br>\nBased on my understating, this competition consists of two parts,<br>\nimage classification and object detection. </p>\n<p>The CSV file \"train_image_level.csv\" contains the folder names we have to read for the object detection task and the CSV file \"train_study_level.csv\" contains the folder names require for the classification task. <br>\nAll the images for training are stored in the path :<br>\n/kaggle/input/siim-covid19-detection/train<br>\nand for testing:<br>\n/kaggle/input/siim-covid19-detection/test</p>\n<p>My questions are:</p>\n<ol>\n<li>Currently all the images are in different folders. Can we use the same folder structure for the model training or can I create my own folders for \"study\" and \"image\" in the kaggle folder structure and copy the relevant images and perform the training?</li>\n<li>Are all the images available for detection are also a classification problem? </li>\n<li>Should our submission file contains the results for the images in the \"/kaggle/input/siim-covid19-detection/test\" folder? But we do not know which images to perform classification and which images to perform detection. </li>\n<li>Do we have to perform classification and detection using one ML model or can we have two different models for the each task?</li>\n</ol>\n<p>Thanks a lot!!</p>",
  "messages": [
    {
      "id": "1338620",
      "postDate": "06/06/2021 15:34:29",
      "content": "<p>Hi All,</p>\n<p>I am new to Kaggle and this is my first real competition. I have few things to clarify about this competition.<br>\nBased on my understating, this competition consists of two parts,<br>\nimage classification and object detection. </p>\n<p>The CSV file \"train_image_level.csv\" contains the folder names we have to read for the object detection task and the CSV file \"train_study_level.csv\" contains the folder names require for the classification task. <br>\nAll the images for training are stored in the path :<br>\n/kaggle/input/siim-covid19-detection/train<br>\nand for testing:<br>\n/kaggle/input/siim-covid19-detection/test</p>\n<p>My questions are:</p>\n<ol>\n<li>Currently all the images are in different folders. Can we use the same folder structure for the model training or can I create my own folders for \"study\" and \"image\" in the kaggle folder structure and copy the relevant images and perform the training?</li>\n<li>Are all the images available for detection are also a classification problem? </li>\n<li>Should our submission file contains the results for the images in the \"/kaggle/input/siim-covid19-detection/test\" folder? But we do not know which images to perform classification and which images to perform detection. </li>\n<li>Do we have to perform classification and detection using one ML model or can we have two different models for the each task?</li>\n</ol>\n<p>Thanks a lot!!</p>",
      "rawMarkdown": "Hi All,\n\nI am new to Kaggle and this is my first real competition. I have few things to clarify about this competition.\nBased on my understating, this competition consists of two parts,\nimage classification and object detection. \n\nThe CSV file \"train_image_level.csv\" contains the folder names we have to read for the object detection task and the CSV file \"train_study_level.csv\" contains the folder names require for the classification task. \nAll the images for training are stored in the path :\n/kaggle/input/siim-covid19-detection/train\nand for testing:\n/kaggle/input/siim-covid19-detection/test\n\nMy questions are:\n1. Currently all the images are in different folders. Can we use the same folder structure for the model training or can I create my own folders for \"study\" and \"image\" in the kaggle folder structure and copy the relevant images and perform the training?\n2. Are all the images available for detection are also a classification problem? \n3. Should our submission file contains the results for the images in the \"/kaggle/input/siim-covid19-detection/test\" folder? But we do not know which images to perform classification and which images to perform detection. \n4. Do we have to perform classification and detection using one ML model or can we have two different models for the each task?\n\nThanks a lot!!",
      "votes": null
    },
    {
      "id": "1338661",
      "postDate": "06/06/2021 16:08:37",
      "content": "<ol>\n<li>Yes you can use any way you want for training ! make it separate folders for image level  and study level or a single folder its al your choice </li>\n<li>There are total of 6 labels , 4 labels in study level i.e negative for pneumonia , atypical , intermediate and typical so a classic multiclass classification or you can make it a binary classification by using 4 models for each class  and for image level there are 2 labels opacity or none and you have to predict the bounding boxes in it [ object detection and images can have multiple bbox ] </li>\n<li>yes it should contains our predictions in the prediction string column, we have to do classifcation on the  id which has _study in it and object detection in the  id which has _image in it </li>\n<li>you can use mutiple models for each task it depends upon you because ulitmately you have to be at the top of lb </li>\n</ol>",
      "rawMarkdown": "1. Yes you can use any way you want for training ! make it separate folders for image level  and study level or a single folder its al your choice \n2. There are total of 6 labels , 4 labels in study level i.e negative for pneumonia , atypical , intermediate and typical so a classic multiclass classification or you can make it a binary classification by using 4 models for each class  and for image level there are 2 labels opacity or none and you have to predict the bounding boxes in it [ object detection and images can have multiple bbox ] \n3. yes it should contains our predictions in the prediction string column, we have to do classifcation on the  id which has _study in it and object detection in the  id which has _image in it \n4. you can use mutiple models for each task it depends upon you because ulitmately you have to be at the top of lb",
      "votes": null
    },
    {
      "id": "1338690",
      "postDate": "06/06/2021 16:35:41",
      "content": "<p>Hi Shubham Thapa,<br>\nMany thanks for the reply. <br>\nI have a follow up question on question 3:<br>\nI checked on the test folder, it does not contain any CSV file indication \"study\" and \"image\" folder. Even the images in the test folders are not named as \"xxxx_image.dcm\" and \"xxxx_study.dcm\".<br>\nDoes it mean I have to perform classification and detection for the same image twice?<br>\nHowever, when I check the sample submission file, they haven't perform classification and detection for the same image.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Shubham Thapa,\nMany thanks for the reply. \nI have a follow up question on question 3:\nI checked on the test folder, it does not contain any CSV file indication \"study\" and \"image\" folder. Even the images in the test folders are not named as \"xxxx_image.dcm\" and \"xxxx_study.dcm\".\nDoes it mean I have to perform classification and detection for the same image twice?\nHowever, when I check the sample submission file, they haven't perform classification and detection for the same image.\n\n\nThanks!",
      "votes": null
    },
    {
      "id": "1338738",
      "postDate": "06/06/2021 17:17:19",
      "content": "<p>the test folder has all the images for testing and the sample submission file is your main test csv file here you can see the id column which is named as 00188a671292_study or a29c5a68b07b_image now from this you can distinguish which one to classify and which one for object detection , your predictions string is something like this for study level Class_name1 Confidence_score  0 0 1 1 Class_name2 Confidence_score  0 0 1 1  Class_name3 Confidence_score  0 0 1 1  Class_name4 Confidence_score  0 0 1 1   confidence score can be reffered to as the probability and for image level it is like this class_name  Confidence_score BBox_coordinates and so on if there are more objects in that single image  as an image can have multiple objects which we need to find a bbox or none 1 0 0 1 1 if there is no objects to be detected . <br>\nYou can find a detailed discussion here <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240329\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240329</a></p>",
      "rawMarkdown": "the test folder has all the images for testing and the sample submission file is your main test csv file here you can see the id column which is named as 00188a671292_study or a29c5a68b07b_image now from this you can distinguish which one to classify and which one for object detection , your predictions string is something like this for study level Class_name1 Confidence_score  0 0 1 1 Class_name2 Confidence_score  0 0 1 1  Class_name3 Confidence_score  0 0 1 1  Class_name4 Confidence_score  0 0 1 1   confidence score can be reffered to as the probability and for image level it is like this class_name  Confidence_score BBox_coordinates and so on if there are more objects in that single image  as an image can have multiple objects which we need to find a bbox or none 1 0 0 1 1 if there is no objects to be detected . \nYou can find a detailed discussion here https://www.kaggle.com/c/siim-covid19-detection/discussion/240329",
      "votes": null
    },
    {
      "id": "1338849",
      "postDate": "06/06/2021 19:10:14",
      "content": "<p>Hi Shubham Thapa,<br>\nThanks for the information. </p>",
      "rawMarkdown": "Hi Shubham Thapa,\nThanks for the information.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1338661,
      "author_name": "trooperog",
      "author_url": "",
      "post_date": "06/06/2021 16:08:37",
      "content": "<ol>\n<li>Yes you can use any way you want for training ! make it separate folders for image level  and study level or a single folder its al your choice </li>\n<li>There are total of 6 labels , 4 labels in study level i.e negative for pneumonia , atypical , intermediate and typical so a classic multiclass classification or you can make it a binary classification by using 4 models for each class  and for image level there are 2 labels opacity or none and you have to predict the bounding boxes in it [ object detection and images can have multiple bbox ] </li>\n<li>yes it should contains our predictions in the prediction string column, we have to do classifcation on the  id which has _study in it and object detection in the  id which has _image in it </li>\n<li>you can use mutiple models for each task it depends upon you because ulitmately you have to be at the top of lb </li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1338690,
      "author_name": "radike",
      "author_url": "",
      "post_date": "06/06/2021 16:35:41",
      "content": "<p>Hi Shubham Thapa,<br>\nMany thanks for the reply. <br>\nI have a follow up question on question 3:<br>\nI checked on the test folder, it does not contain any CSV file indication \"study\" and \"image\" folder. Even the images in the test folders are not named as \"xxxx_image.dcm\" and \"xxxx_study.dcm\".<br>\nDoes it mean I have to perform classification and detection for the same image twice?<br>\nHowever, when I check the sample submission file, they haven't perform classification and detection for the same image.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1338738,
          "author_name": "trooperog",
          "author_url": "",
          "post_date": "06/06/2021 17:17:19",
          "content": "<p>the test folder has all the images for testing and the sample submission file is your main test csv file here you can see the id column which is named as 00188a671292_study or a29c5a68b07b_image now from this you can distinguish which one to classify and which one for object detection , your predictions string is something like this for study level Class_name1 Confidence_score  0 0 1 1 Class_name2 Confidence_score  0 0 1 1  Class_name3 Confidence_score  0 0 1 1  Class_name4 Confidence_score  0 0 1 1   confidence score can be reffered to as the probability and for image level it is like this class_name  Confidence_score BBox_coordinates and so on if there are more objects in that single image  as an image can have multiple objects which we need to find a bbox or none 1 0 0 1 1 if there is no objects to be detected . <br>\nYou can find a detailed discussion here <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240329\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240329</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1338849,
      "author_name": "radike",
      "author_url": "",
      "post_date": "06/06/2021 19:10:14",
      "content": "<p>Hi Shubham Thapa,<br>\nThanks for the information. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1338620": "Hi All,\n\nI am new to Kaggle and this is my first real competition. I have few things to clarify about this competition.\nBased on my understating, this competition consists of two parts,\nimage classification and object detection. \n\nThe CSV file \"train_image_level.csv\" contains the folder names we have to read for the object detection task and the CSV file \"train_study_level.csv\" contains the folder names require for the classification task. \nAll the images for training are stored in the path :\n/kaggle/input/siim-covid19-detection/train\nand for testing:\n/kaggle/input/siim-covid19-detection/test\n\nMy questions are:\n1. Currently all the images are in different folders. Can we use the same folder structure for the model training or can I create my own folders for \"study\" and \"image\" in the kaggle folder structure and copy the relevant images and perform the training?\n2. Are all the images available for detection are also a classification problem? \n3. Should our submission file contains the results for the images in the \"/kaggle/input/siim-covid19-detection/test\" folder? But we do not know which images to perform classification and which images to perform detection. \n4. Do we have to perform classification and detection using one ML model or can we have two different models for the each task?\n\nThanks a lot!!",
    "1338661": "1. Yes you can use any way you want for training ! make it separate folders for image level  and study level or a single folder its al your choice \n2. There are total of 6 labels , 4 labels in study level i.e negative for pneumonia , atypical , intermediate and typical so a classic multiclass classification or you can make it a binary classification by using 4 models for each class  and for image level there are 2 labels opacity or none and you have to predict the bounding boxes in it [ object detection and images can have multiple bbox ] \n3. yes it should contains our predictions in the prediction string column, we have to do classifcation on the  id which has _study in it and object detection in the  id which has _image in it \n4. you can use mutiple models for each task it depends upon you because ulitmately you have to be at the top of lb",
    "1338690": "Hi Shubham Thapa,\nMany thanks for the reply. \nI have a follow up question on question 3:\nI checked on the test folder, it does not contain any CSV file indication \"study\" and \"image\" folder. Even the images in the test folders are not named as \"xxxx_image.dcm\" and \"xxxx_study.dcm\".\nDoes it mean I have to perform classification and detection for the same image twice?\nHowever, when I check the sample submission file, they haven't perform classification and detection for the same image.\n\n\nThanks!",
    "1338738": "the test folder has all the images for testing and the sample submission file is your main test csv file here you can see the id column which is named as 00188a671292_study or a29c5a68b07b_image now from this you can distinguish which one to classify and which one for object detection , your predictions string is something like this for study level Class_name1 Confidence_score  0 0 1 1 Class_name2 Confidence_score  0 0 1 1  Class_name3 Confidence_score  0 0 1 1  Class_name4 Confidence_score  0 0 1 1   confidence score can be reffered to as the probability and for image level it is like this class_name  Confidence_score BBox_coordinates and so on if there are more objects in that single image  as an image can have multiple objects which we need to find a bbox or none 1 0 0 1 1 if there is no objects to be detected . \nYou can find a detailed discussion here https://www.kaggle.com/c/siim-covid19-detection/discussion/240329",
    "1338849": "Hi Shubham Thapa,\nThanks for the information."
  },
  "source": "meta"
}