{
  "id": 160136,
  "title": "DICOM or jpeg?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160136",
  "author_name": "",
  "post_date": "2020-06-20T01:19:35.395558900Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>What are benefits of using DICOM instead JPEG? Which do you use?</p>",
  "messages": [
    {
      "id": "893832",
      "postDate": "06/20/2020 01:19:35",
      "content": "<p>What are benefits of using DICOM instead JPEG? Which do you use?</p>",
      "rawMarkdown": "What are benefits of using DICOM instead JPEG? Which do you use?",
      "votes": null
    },
    {
      "id": "893881",
      "postDate": "06/20/2020 04:00:05",
      "content": "<p>Hi Jacek, good to see you here. I have only been using TFRecords which contain both jpeg and meta data. I posted my datasets 768x768 <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512 <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a>, 384x384 <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, 256x256 <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>, and external data 512x512 <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a>.</p>\n\n<p>I think jpegs are just images without meta data whereas DICOM contain both image and meta data. Also TFRecords can contain both image and meta data. I have found that the jpegs (and TFRecords) have good enough quality to score high in CV and LB but perhaps the DICOM have more quality that could help, i'm not sure.</p>",
      "rawMarkdown": "Hi Jacek, good to see you here. I have only been using TFRecords which contain both jpeg and meta data. I posted my datasets 768x768 [here][1], 512x512 [here][2], 384x384 [here][3], 256x256 [here][4], and external data 512x512 [here][5].\n\nI think jpegs are just images without meta data whereas DICOM contain both image and meta data. Also TFRecords can contain both image and meta data. I have found that the jpegs (and TFRecords) have good enough quality to score high in CV and LB but perhaps the DICOM have more quality that could help, i'm not sure.\n  \n[1]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[2]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[3]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[5]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images",
      "votes": null
    },
    {
      "id": "893991",
      "postDate": "06/20/2020 05:23:11",
      "content": "<p>The background for what was produced by what with respect to the DICOM and JPEG images has been explained by Phil Culliton from Kaggle in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155415\">this discussion.</a></p>",
      "rawMarkdown": "The background for what was produced by what with respect to the DICOM and JPEG images has been explained by Phil Culliton from Kaggle in [this discussion.](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155415)",
      "votes": null
    },
    {
      "id": "894323",
      "postDate": "06/20/2020 10:36:58",
      "content": "<p>how does the interpolation algorithms, like BICUBIC vs LANCZOS vs others, during resizing affect model performance?</p>",
      "rawMarkdown": "how does the interpolation algorithms, like BICUBIC vs LANCZOS vs others, during resizing affect model performance?",
      "votes": null
    },
    {
      "id": "894418",
      "postDate": "06/20/2020 12:11:44",
      "content": "<p>Thanks - \"The source JPEGs' pixel information was extracted and written into the Pixel Data of the DICOM files.\" - I was thinking JPEG can be worse because lossy compression, now I know that's not the case.</p>",
      "rawMarkdown": "Thanks - \"The source JPEGs' pixel information was extracted and written into the Pixel Data of the DICOM files.\" - I was thinking JPEG can be worse because lossy compression, now I know that's not the case.",
      "votes": null
    },
    {
      "id": "894422",
      "postDate": "06/20/2020 12:13:15",
      "content": "<p>Hello <a href=\"/cdeotte\">@cdeotte</a> :)\nso we are using scaled-down images in this competion? I have seen size of the data so I wonder how long is the training, but in the kernels I was reading it takes few minutes per epoch for some reason.</p>",
      "rawMarkdown": "Hello @cdeotte :)\nso we are using scaled-down images in this competion? I have seen size of the data so I wonder how long is the training, but in the kernels I was reading it takes few minutes per epoch for some reason.",
      "votes": null
    },
    {
      "id": "898792",
      "postDate": "06/23/2020 18:40:11",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for sharing your data.\nDo you know if it is possible (and optimize) to use your tfrecord If I want to experiment the data on pytorch instead of Tensorflow ? Or is it better to extract the jpeg and metadata first and then use them with pytorch ?</p>",
      "rawMarkdown": "Thanks @cdeotte  for sharing your data.\nDo you know if it is possible (and optimize) to use your tfrecord If I want to experiment the data on pytorch instead of Tensorflow ? Or is it better to extract the jpeg and metadata first and then use them with pytorch ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 893881,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/20/2020 04:00:05",
      "content": "<p>Hi Jacek, good to see you here. I have only been using TFRecords which contain both jpeg and meta data. I posted my datasets 768x768 <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512 <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a>, 384x384 <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, 256x256 <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>, and external data 512x512 <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a>.</p>\n\n<p>I think jpegs are just images without meta data whereas DICOM contain both image and meta data. Also TFRecords can contain both image and meta data. I have found that the jpegs (and TFRecords) have good enough quality to score high in CV and LB but perhaps the DICOM have more quality that could help, i'm not sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 894323,
          "author_name": "shayekh",
          "author_url": "",
          "post_date": "06/20/2020 10:36:58",
          "content": "<p>how does the interpolation algorithms, like BICUBIC vs LANCZOS vs others, during resizing affect model performance?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 894422,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "06/20/2020 12:13:15",
          "content": "<p>Hello <a href=\"/cdeotte\">@cdeotte</a> :)\nso we are using scaled-down images in this competion? I have seen size of the data so I wonder how long is the training, but in the kernels I was reading it takes few minutes per epoch for some reason.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898792,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "06/23/2020 18:40:11",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for sharing your data.\nDo you know if it is possible (and optimize) to use your tfrecord If I want to experiment the data on pytorch instead of Tensorflow ? Or is it better to extract the jpeg and metadata first and then use them with pytorch ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893991,
      "author_name": "mutantspore",
      "author_url": "",
      "post_date": "06/20/2020 05:23:11",
      "content": "<p>The background for what was produced by what with respect to the DICOM and JPEG images has been explained by Phil Culliton from Kaggle in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155415\">this discussion.</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 894418,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "06/20/2020 12:11:44",
          "content": "<p>Thanks - \"The source JPEGs' pixel information was extracted and written into the Pixel Data of the DICOM files.\" - I was thinking JPEG can be worse because lossy compression, now I know that's not the case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "893832": "What are benefits of using DICOM instead JPEG? Which do you use?",
    "893881": "Hi Jacek, good to see you here. I have only been using TFRecords which contain both jpeg and meta data. I posted my datasets 768x768 [here][1], 512x512 [here][2], 384x384 [here][3], 256x256 [here][4], and external data 512x512 [here][5].\n\nI think jpegs are just images without meta data whereas DICOM contain both image and meta data. Also TFRecords can contain both image and meta data. I have found that the jpegs (and TFRecords) have good enough quality to score high in CV and LB but perhaps the DICOM have more quality that could help, i'm not sure.\n  \n[1]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[2]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[3]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[5]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images",
    "893991": "The background for what was produced by what with respect to the DICOM and JPEG images has been explained by Phil Culliton from Kaggle in [this discussion.](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155415)",
    "894323": "how does the interpolation algorithms, like BICUBIC vs LANCZOS vs others, during resizing affect model performance?",
    "894418": "Thanks - \"The source JPEGs' pixel information was extracted and written into the Pixel Data of the DICOM files.\" - I was thinking JPEG can be worse because lossy compression, now I know that's not the case.",
    "894422": "Hello @cdeotte :)\nso we are using scaled-down images in this competion? I have seen size of the data so I wonder how long is the training, but in the kernels I was reading it takes few minutes per epoch for some reason.",
    "898792": "Thanks @cdeotte  for sharing your data.\nDo you know if it is possible (and optimize) to use your tfrecord If I want to experiment the data on pytorch instead of Tensorflow ? Or is it better to extract the jpeg and metadata first and then use them with pytorch ?"
  },
  "source": "meta"
}