{
  "id": 265255,
  "title": "Cases of Submission failure and success",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/265255",
  "author_name": "",
  "post_date": "2021-08-15T07:53:43.336382Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>When I used DICOM image by myself without using any external dataset, my submission was successful. </p>\n<p>In other notebooks, I tried submitting several times using an external dataset of PNG-converted image, but all submissions resulted in failure. </p>\n<p>Even if the submission file is perfect, it seem that the submission system would not accept submissions from the notebook quoting some external data.</p>",
  "messages": [
    {
      "id": "1472910",
      "postDate": "08/15/2021 07:53:43",
      "content": "<p>When I used DICOM image by myself without using any external dataset, my submission was successful. </p>\n<p>In other notebooks, I tried submitting several times using an external dataset of PNG-converted image, but all submissions resulted in failure. </p>\n<p>Even if the submission file is perfect, it seem that the submission system would not accept submissions from the notebook quoting some external data.</p>",
      "rawMarkdown": "When I used DICOM image by myself without using any external dataset, my submission was successful. \n\nIn other notebooks, I tried submitting several times using an external dataset of PNG-converted image, but all submissions resulted in failure. \n\nEven if the submission file is perfect, it seem that the submission system would not accept submissions from the notebook quoting some external data.",
      "votes": null
    },
    {
      "id": "1473889",
      "postDate": "08/15/2021 18:43:35",
      "content": "<p>i am a beginner , is any way to convert PNG image to DICOM image</p>",
      "rawMarkdown": "i am a beginner , is any way to convert PNG image to DICOM image",
      "votes": null
    },
    {
      "id": "1473904",
      "postDate": "08/15/2021 19:05:09",
      "content": "<p>Hi or try to convert PNG image to DICOM image </p>",
      "rawMarkdown": "Hi or try to convert PNG image to DICOM image",
      "votes": null
    },
    {
      "id": "1477130",
      "postDate": "08/17/2021 11:11:09",
      "content": "<p>Hi there , to solve this problem i need to tell you about how scoring occurs in kaggle</p>\n<p>Given :-</p>\n<p>Training data :-</p>\n<p>you are given a training data folder with labelled images on which you train this model however this is just 80 percent of entire data </p>\n<p>80 % of data is  train folder</p>\n<p>and the rest 20 is coming next</p>\n<p>Test folder :-</p>\n<p>This folder contains maybe 1% of the 20 % part of the entire dataset . Then you write code to generate prediction in csv format from the test folder</p>\n<p>Data left :- 19 %</p>\n<p>Public Score :-</p>\n<p>The score you are getting until now is called Public score which basically is the area under roc curve metric ( this is selected by the competiton hosters , it might also have been accuracy metric however and is written in the evaluation page ) . And what is this metric calculated on maybe 3 % of 19 % data left . Basically the test folder has bunch more images and your code is rerun  with the updated folder . But if you use any external test dataset with png images that folders images will not be increased ( Rather the dicom images test folder) . And you will use the old one only , thus you will predict 87 samples in the submission but actually there  are a ton of images , thus it results in error . Thus you need to use dicom images</p>\n<p>Data left :- 16 %</p>\n<p>Private Score :-<br>\nYou are told to select two submissions which you think is the best . And is rerun on 16 % data and a new scores comes out known as private score ( Not given yet) . And on that your final rank is selected .</p>",
      "rawMarkdown": "Hi there , to solve this problem i need to tell you about how scoring occurs in kaggle\n\nGiven :-\n\nTraining data :-\n\nyou are given a training data folder with labelled images on which you train this model however this is just 80 percent of entire data \n\n80 % of data is  train folder\n\nand the rest 20 is coming next\n\nTest folder :-\n\nThis folder contains maybe 1% of the 20 % part of the entire dataset . Then you write code to generate prediction in csv format from the test folder\n\nData left :- 19 %\n\n Public Score :-\n\nThe score you are getting until now is called Public score which basically is the area under roc curve metric ( this is selected by the competiton hosters , it might also have been accuracy metric however and is written in the evaluation page ) . And what is this metric calculated on maybe 3 % of 19 % data left . Basically the test folder has bunch more images and your code is rerun  with the updated folder . But if you use any external test dataset with png images that folders images will not be increased ( Rather the dicom images test folder) . And you will use the old one only , thus you will predict 87 samples in the submission but actually there  are a ton of images , thus it results in error . Thus you need to use dicom images\n\nData left :- 16 %\n\nPrivate Score :-\nYou are told to select two submissions which you think is the best . And is rerun on 16 % data and a new scores comes out known as private score ( Not given yet) . And on that your final rank is selected .",
      "votes": null
    },
    {
      "id": "1477432",
      "postDate": "08/17/2021 13:29:04",
      "content": "<p>Hi, I assume you mean the other way around? I realised myself that my model built on png training data would be able to accept the test data since that was in dicom format, unless I built in a conversion from dicom to png as part of the inference notebook. I copied a lot from <a href=\"https://www.kaggle.com/anasshnn/fast-dicom-png-full-data-download-data\" target=\"_blank\">https://www.kaggle.com/anasshnn/fast-dicom-png-full-data-download-data</a> . </p>\n<p>However despite it successfully scoring the small test set initially provided, when you submit predictions I found it timed out. As Sayantan says above, there is more data to score in this part of the process. Therefore it seems like I cannot use png at all since it doesn't give enough time left to run predictions after converting from dicom to png. But if you want to try the above notebook contains a way to convert.</p>",
      "rawMarkdown": "Hi, I assume you mean the other way around? I realised myself that my model built on png training data would be able to accept the test data since that was in dicom format, unless I built in a conversion from dicom to png as part of the inference notebook. I copied a lot from https://www.kaggle.com/anasshnn/fast-dicom-png-full-data-download-data . \n\nHowever despite it successfully scoring the small test set initially provided, when you submit predictions I found it timed out. As Sayantan says above, there is more data to score in this part of the process. Therefore it seems like I cannot use png at all since it doesn't give enough time left to run predictions after converting from dicom to png. But if you want to try the above notebook contains a way to convert.",
      "votes": null
    },
    {
      "id": "1479384",
      "postDate": "08/18/2021 12:56:20",
      "content": "<p>even if you are not using the competition data for coding your submission notebook, add it to your submission notebook and try to submit</p>",
      "rawMarkdown": "even if you are not using the competition data for coding your submission notebook, add it to your submission notebook and try to submit",
      "votes": null
    },
    {
      "id": "1531071",
      "postDate": "10/01/2021 15:46:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/drhekalkhatib\" target=\"_blank\">@drhekalkhatib</a> , DICOM images have metadata that you have to include it, simple kit is a tool that you can use for it, follow the next link to understand better and see an example: </p>\n<h4><a href=\"https://simpleitk.readthedocs.io/en/next/Examples/DicomSeriesFromArray/Documentation.html\" target=\"_blank\">https://simpleitk.readthedocs.io/en/next/Examples/DicomSeriesFromArray/Documentation.html</a></h4>",
      "rawMarkdown": "Hi @drhekalkhatib , DICOM images have metadata that you have to include it, simple kit is a tool that you can use for it, follow the next link to understand better and see an example: \n#### https://simpleitk.readthedocs.io/en/next/Examples/DicomSeriesFromArray/Documentation.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1473889,
      "author_name": "drhekalkhatib",
      "author_url": "",
      "post_date": "08/15/2021 18:43:35",
      "content": "<p>i am a beginner , is any way to convert PNG image to DICOM image</p>",
      "votes": null,
      "replies": [
        {
          "id": 1477432,
          "author_name": "moonshots",
          "author_url": "",
          "post_date": "08/17/2021 13:29:04",
          "content": "<p>Hi, I assume you mean the other way around? I realised myself that my model built on png training data would be able to accept the test data since that was in dicom format, unless I built in a conversion from dicom to png as part of the inference notebook. I copied a lot from <a href=\"https://www.kaggle.com/anasshnn/fast-dicom-png-full-data-download-data\" target=\"_blank\">https://www.kaggle.com/anasshnn/fast-dicom-png-full-data-download-data</a> . </p>\n<p>However despite it successfully scoring the small test set initially provided, when you submit predictions I found it timed out. As Sayantan says above, there is more data to score in this part of the process. Therefore it seems like I cannot use png at all since it doesn't give enough time left to run predictions after converting from dicom to png. But if you want to try the above notebook contains a way to convert.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1531071,
          "author_name": "victorfernandezalbor",
          "author_url": "",
          "post_date": "10/01/2021 15:46:30",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/drhekalkhatib\" target=\"_blank\">@drhekalkhatib</a> , DICOM images have metadata that you have to include it, simple kit is a tool that you can use for it, follow the next link to understand better and see an example: </p>\n<h4><a href=\"https://simpleitk.readthedocs.io/en/next/Examples/DicomSeriesFromArray/Documentation.html\" target=\"_blank\">https://simpleitk.readthedocs.io/en/next/Examples/DicomSeriesFromArray/Documentation.html</a></h4>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1473904,
      "author_name": "drhekalkhatib",
      "author_url": "",
      "post_date": "08/15/2021 19:05:09",
      "content": "<p>Hi or try to convert PNG image to DICOM image </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1477130,
      "author_name": "swaralipibose",
      "author_url": "",
      "post_date": "08/17/2021 11:11:09",
      "content": "<p>Hi there , to solve this problem i need to tell you about how scoring occurs in kaggle</p>\n<p>Given :-</p>\n<p>Training data :-</p>\n<p>you are given a training data folder with labelled images on which you train this model however this is just 80 percent of entire data </p>\n<p>80 % of data is  train folder</p>\n<p>and the rest 20 is coming next</p>\n<p>Test folder :-</p>\n<p>This folder contains maybe 1% of the 20 % part of the entire dataset . Then you write code to generate prediction in csv format from the test folder</p>\n<p>Data left :- 19 %</p>\n<p>Public Score :-</p>\n<p>The score you are getting until now is called Public score which basically is the area under roc curve metric ( this is selected by the competiton hosters , it might also have been accuracy metric however and is written in the evaluation page ) . And what is this metric calculated on maybe 3 % of 19 % data left . Basically the test folder has bunch more images and your code is rerun  with the updated folder . But if you use any external test dataset with png images that folders images will not be increased ( Rather the dicom images test folder) . And you will use the old one only , thus you will predict 87 samples in the submission but actually there  are a ton of images , thus it results in error . Thus you need to use dicom images</p>\n<p>Data left :- 16 %</p>\n<p>Private Score :-<br>\nYou are told to select two submissions which you think is the best . And is rerun on 16 % data and a new scores comes out known as private score ( Not given yet) . And on that your final rank is selected .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1479384,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "08/18/2021 12:56:20",
      "content": "<p>even if you are not using the competition data for coding your submission notebook, add it to your submission notebook and try to submit</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1472910": "When I used DICOM image by myself without using any external dataset, my submission was successful. \n\nIn other notebooks, I tried submitting several times using an external dataset of PNG-converted image, but all submissions resulted in failure. \n\nEven if the submission file is perfect, it seem that the submission system would not accept submissions from the notebook quoting some external data.",
    "1473889": "i am a beginner , is any way to convert PNG image to DICOM image",
    "1473904": "Hi or try to convert PNG image to DICOM image",
    "1477130": "Hi there , to solve this problem i need to tell you about how scoring occurs in kaggle\n\nGiven :-\n\nTraining data :-\n\nyou are given a training data folder with labelled images on which you train this model however this is just 80 percent of entire data \n\n80 % of data is  train folder\n\nand the rest 20 is coming next\n\nTest folder :-\n\nThis folder contains maybe 1% of the 20 % part of the entire dataset . Then you write code to generate prediction in csv format from the test folder\n\nData left :- 19 %\n\n Public Score :-\n\nThe score you are getting until now is called Public score which basically is the area under roc curve metric ( this is selected by the competiton hosters , it might also have been accuracy metric however and is written in the evaluation page ) . And what is this metric calculated on maybe 3 % of 19 % data left . Basically the test folder has bunch more images and your code is rerun  with the updated folder . But if you use any external test dataset with png images that folders images will not be increased ( Rather the dicom images test folder) . And you will use the old one only , thus you will predict 87 samples in the submission but actually there  are a ton of images , thus it results in error . Thus you need to use dicom images\n\nData left :- 16 %\n\nPrivate Score :-\nYou are told to select two submissions which you think is the best . And is rerun on 16 % data and a new scores comes out known as private score ( Not given yet) . And on that your final rank is selected .",
    "1477432": "Hi, I assume you mean the other way around? I realised myself that my model built on png training data would be able to accept the test data since that was in dicom format, unless I built in a conversion from dicom to png as part of the inference notebook. I copied a lot from https://www.kaggle.com/anasshnn/fast-dicom-png-full-data-download-data . \n\nHowever despite it successfully scoring the small test set initially provided, when you submit predictions I found it timed out. As Sayantan says above, there is more data to score in this part of the process. Therefore it seems like I cannot use png at all since it doesn't give enough time left to run predictions after converting from dicom to png. But if you want to try the above notebook contains a way to convert.",
    "1479384": "even if you are not using the competition data for coding your submission notebook, add it to your submission notebook and try to submit",
    "1531071": "Hi @drhekalkhatib , DICOM images have metadata that you have to include it, simple kit is a tool that you can use for it, follow the next link to understand better and see an example: \n#### https://simpleitk.readthedocs.io/en/next/Examples/DicomSeriesFromArray/Documentation.html"
  },
  "source": "meta"
}