{
  "id": 155859,
  "title": "[Merge External Data]",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155859",
  "author_name": "Alex Shonenkov",
  "post_date": "2020-06-03T09:26:06.353000",
  "votes": 74,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>I have created <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">kernel with merge external data</a> (including metadata):</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection\">Melanoma Detection Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">Skin Lesion Images for Melanoma Classification</a></li>\n<li><a href=\"https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000\">Skin Cancer MNIST: HAM10000</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/data\">SIIM-ISIC Melanoma Classification</a></li>\n</ul>\n\n<p>I have got 53351 samples for class 0 and 5106 for class 1. </p>\n\n<p>Merged dataset you can find <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">here</a>\n- format: jpeg, rgb\n- size: 512x512</p>\n\n<p>if you need other format of data (tf-records or other image size) you can use <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">my kernel</a></p>\n\n<p>I hope it helps you!</p>\n\n<p>Welcome!</p>",
  "messages": [
    {
      "id": 872528,
      "postDate": "2020-06-03T09:26:06.353Z",
      "content": "<p>Hi everyone!</p>\n\n<p>I have created <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">kernel with merge external data</a> (including metadata):</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection\">Melanoma Detection Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">Skin Lesion Images for Melanoma Classification</a></li>\n<li><a href=\"https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000\">Skin Cancer MNIST: HAM10000</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/data\">SIIM-ISIC Melanoma Classification</a></li>\n</ul>\n\n<p>I have got 53351 samples for class 0 and 5106 for class 1. </p>\n\n<p>Merged dataset you can find <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">here</a>\n- format: jpeg, rgb\n- size: 512x512</p>\n\n<p>if you need other format of data (tf-records or other image size) you can use <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">my kernel</a></p>\n\n<p>I hope it helps you!</p>\n\n<p>Welcome!</p>",
      "rawMarkdown": "Hi everyone!\n\nI have created [kernel with merge external data](https://www.kaggle.com/shonenkov/merge-external-data) (including metadata):\n\n- [Melanoma Detection Dataset](https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection)\n- [Skin Lesion Images for Melanoma Classification](https://www.kaggle.com/andrewmvd/isic-2019)\n- [Skin Cancer MNIST: HAM10000](https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000)\n- [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification/data)\n\nI have got 53351 samples for class 0 and 5106 for class 1. \n\nMerged dataset you can find [here](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg)\n- format: jpeg, rgb\n- size: 512x512\n\nif you need other format of data (tf-records or other image size) you can use [my kernel](https://www.kaggle.com/shonenkov/merge-external-data)\n\nI hope it helps you!\n\nWelcome!\n",
      "votes": 74
    },
    {
      "id": 878163,
      "postDate": "2020-06-08T10:06:50.290Z",
      "content": "<p>Please be aware of some duplicated images in the marged dataset that may has negative effect on the CV. This is mainly due to the naming of the ISIC 2019 images which some has \"_downsampled\" term in the name. \nExample of duplicate images:\nISIC_0000036.jpg (fold 0)\nISIC_0000036_downsampled.jpg (fold 1)</p>",
      "rawMarkdown": "Please be aware of some duplicated images in the marged dataset that may has negative effect on the CV. This is mainly due to the naming of the ISIC 2019 images which some has \"_downsampled\" term in the name. \nExample of duplicate images:\nISIC_0000036.jpg (fold 0)\nISIC_0000036_downsampled.jpg (fold 1)",
      "votes": 4,
      "replies": [
        {
          "id": 878178,
          "postDate": "2020-06-08T10:22:54.760Z",
          "content": "<p>thank you! Sure, It needs some corrections :)</p>",
          "rawMarkdown": "thank you! Sure, It needs some corrections :)",
          "votes": 1
        },
        {
          "id": 878183,
          "postDate": "2020-06-08T10:35:50.563Z",
          "content": "<p>I just had a closer look at the datasets. I think it is enough to use the ISIC 2019 dataset (and not use the former datasets (i.e. ISIC 2018, ISIC 2017, ISIC 2016) as there is almost no new images in the former datasets. \nThe only benefit of using former datasets is that they have the images in the original resolutions and not in the downsampled format. Anyway as we may be forced to resize the images for training then even this benefit will not be useful. </p>",
          "rawMarkdown": "I just had a closer look at the datasets. I think it is enough to use the ISIC 2019 dataset (and not use the former datasets (i.e. ISIC 2018, ISIC 2017, ISIC 2016) as there is almost no new images in the former datasets. \nThe only benefit of using former datasets is that they have the images in the original resolutions and not in the downsampled format. Anyway as we may be forced to resize the images for training then even this benefit will not be useful. ",
          "votes": 3,
          "replies": [
            {
              "id": 969479,
              "postDate": "2020-08-13T17:55:49.870Z",
              "content": "<p>Hello ,<br>\n2019 competition winner have used multiple dataset ( <a href=\"https://arxiv.org/pdf/1910.03910.pdf\" target=\"_blank\">https://arxiv.org/pdf/1910.03910.pdf</a> ) ,</p>\n<ul>\n<li>HAM dataset</li>\n<li>7 point dataset</li>\n<li>MSK Dataset</li>\n<li>BCN 20000 Dataset</li>\n</ul>\n<p>Why these dataset not considered as external dataset ?</p>\n<p>From Paper,<br>\n A part of the training dataset is the HAM10000 dataset which contains images of<br>\nsize 600 × 450 that were centered and cropped around the lesion. The dataset<br>\ncurators applied histogram corrections to some images [16]. Another dataset,<br>\nBCN 20000, contains images of size 1024 × 1024. This dataset is particularly<br>\nchallenging as many images are uncropped and lesions in difficult and uncommon<br>\nlocations are present [5]. Last, the MSK dataset contains images with various<br>\nsizes.<br>\nThe dataset also contains meta information about the patient’s age group<br>\n(in steps of five years), the anatomical site (eight possible sites) and the sex<br>\n(male/female). The meta data is partially incomplete, i.e., there are missing<br>\nvalues for some images.<br>\nIn addition, we make use of external data. We use the 995 dermoscopic images<br>\nfrom the 7-point dataset [2]. </p>\n<p>Thanks,<br>\nAnthony</p>",
              "rawMarkdown": "Hello ,\n2019 competition winner have used multiple dataset ( https://arxiv.org/pdf/1910.03910.pdf ) ,\n- HAM dataset\n- 7 point dataset\n- MSK Dataset\n- BCN 20000 Dataset\n\nWhy these dataset not considered as external dataset ?\n\nFrom Paper,\n A part of the training dataset is the HAM10000 dataset which contains images of\nsize 600 × 450 that were centered and cropped around the lesion. The dataset\ncurators applied histogram corrections to some images [16]. Another dataset,\nBCN 20000, contains images of size 1024 × 1024. This dataset is particularly\nchallenging as many images are uncropped and lesions in difficult and uncommon\nlocations are present [5]. Last, the MSK dataset contains images with various\nsizes.\nThe dataset also contains meta information about the patient’s age group\n(in steps of five years), the anatomical site (eight possible sites) and the sex\n(male/female). The meta data is partially incomplete, i.e., there are missing\nvalues for some images.\nIn addition, we make use of external data. We use the 995 dermoscopic images\nfrom the 7-point dataset [2]. \n \nThanks,\nAnthony"
            }
          ]
        },
        {
          "id": 878250,
          "postDate": "2020-06-08T11:38:03.460Z",
          "content": "<p>I have added fixes, thank you! </p>",
          "rawMarkdown": "I have added fixes, thank you! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 872665,
      "postDate": "2020-06-03T12:09:40.317Z",
      "content": "<p>Nice.\nCould you please clarify how you performed resizing? Was it simple resize with a loss of original aspect ratio or was it crop -&gt; resize, keeping an original aspect ratio?</p>",
      "rawMarkdown": "Nice.\nCould you please clarify how you performed resizing? Was it simple resize with a loss of original aspect ratio or was it crop -&gt; resize, keeping an original aspect ratio?",
      "votes": 1,
      "replies": [
        {
          "id": 872672,
          "postDate": "2020-06-03T12:16:19.643Z",
          "content": "<p>Thanks a lot!</p>\n\n<p>I have used cv2.INTER_AREA without crop and saving ratio, you can see full code <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">here</a> ;) if need you can change resizing method. </p>\n\n<p>Welcome!</p>",
          "rawMarkdown": "Thanks a lot!\n\nI have used cv2.INTER_AREA without crop and saving ratio, you can see full code [here](https://www.kaggle.com/shonenkov/merge-external-data) ;) if need you can change resizing method. \n\nWelcome!",
          "votes": 1
        }
      ]
    },
    {
      "id": 872784,
      "postDate": "2020-06-03T13:59:55.517Z",
      "content": "<p>great work. </p>\n\n<p><a href=\"/shonenkov\">@shonenkov</a> hi, can you please ensure the data leakage issue. I think it's really important. I've also faced some issues (like other reporters). Thank you.</p>",
      "rawMarkdown": "great work. \n\n@shonenkov hi, can you please ensure the data leakage issue. I think it's really important. I've also faced some issues (like other reporters). Thank you.",
      "votes": 2
    },
    {
      "id": 1608565,
      "postDate": "2021-12-06T12:57:56.477Z",
      "content": "<p>Can you explain each file？I don't know what these .csv file represents？</p>",
      "rawMarkdown": "Can you explain each file？I don't know what these .csv file represents？"
    },
    {
      "id": 918796,
      "postDate": "2020-07-07T13:56:53.147Z",
      "content": "<p>This should be useful for some experimentation. Thank you!</p>",
      "rawMarkdown": "This should be useful for some experimentation. Thank you!\n"
    },
    {
      "id": 878898,
      "postDate": "2020-06-09T02:08:53.330Z",
      "content": "<p>I only use  ISIC2019 &amp; ISIC2020,but got overfitting,val 0.9269  LB0.887,do you know the reason?</p>",
      "rawMarkdown": "I only use  ISIC2019 &amp; ISIC2020,but got overfitting,val 0.9269  LB0.887,do you know the reason?",
      "replies": [
        {
          "id": 878945,
          "postDate": "2020-06-09T03:58:28.657Z",
          "content": "<ol>\n<li>data leakage</li>\n<li>val data is different from test data\nI guess....</li>\n</ol>",
          "rawMarkdown": "1. data leakage\n2. val data is different from test data\nI guess...."
        },
        {
          "id": 879438,
          "postDate": "2020-06-09T13:24:06.137Z",
          "content": "<p>now my val score 0.9168  lb score 0.917 ,I only fix the channel bgr -&gt;rgb </p>",
          "rawMarkdown": "now my val score 0.9168  lb score 0.917 ,I only fix the channel bgr -&gt;rgb "
        },
        {
          "id": 879448,
          "postDate": "2020-06-09T13:31:31.210Z",
          "content": "<p>For all the images?? </p>",
          "rawMarkdown": "For all the images?? "
        },
        {
          "id": 879453,
          "postDate": "2020-06-09T13:37:56.740Z",
          "content": "<p>Images are already in RGB according to dataset kernel</p>",
          "rawMarkdown": "Images are already in RGB according to dataset kernel"
        },
        {
          "id": 879485,
          "postDate": "2020-06-09T14:04:47Z",
          "content": "<p>ISIC2019 &amp; ISIC2020 the left img is RGB format,but overfit maybe is not the reason.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3132358%2F8b039fc9631dce853f8915576c39951c%2F2020-06-09%2022-01-49.png?generation=1591711413752119&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": " ISIC2019 &amp; ISIC2020 the left img is RGB format,but overfit maybe is not the reason.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3132358%2F8b039fc9631dce853f8915576c39951c%2F2020-06-09%2022-01-49.png?generation=1591711413752119&amp;alt=media)\n",
          "votes": 2
        },
        {
          "id": 879536,
          "postDate": "2020-06-09T14:38:30.510Z",
          "content": "<p>I'm still having trouble getting it. Could you please clarify it...</p>",
          "rawMarkdown": "I'm still having trouble getting it. Could you please clarify it..."
        },
        {
          "id": 880314,
          "postDate": "2020-06-10T07:11:23.460Z",
          "content": "<p>when you open the images using <code>cv2.imread</code>, by default, it has the BGR format. So if you need RGB format you need to switch the blue and red channels. However, if you have all the training and test images in BGR format, I assume it will not affect your final score. </p>",
          "rawMarkdown": "when you open the images using `cv2.imread`, by default, it has the BGR format. So if you need RGB format you need to switch the blue and red channels. However, if you have all the training and test images in BGR format, I assume it will not affect your final score. ",
          "votes": 2
        },
        {
          "id": 923334,
          "postDate": "2020-07-10T18:19:20.583Z",
          "content": "<p>I have the same overfit problem, I could try this…but like others said this  change of channel should not matter.…</p>",
          "rawMarkdown": "I have the same overfit problem, I could try this…but like others said this  change of channel should not matter.…"
        }
      ]
    },
    {
      "id": 876168,
      "postDate": "2020-06-06T13:51:40.040Z",
      "content": "<p>Please could you share the tf-records for the final dataset? Couldn't find it in the notebook..</p>\n\n<p>Thanks and brilliant job!</p>",
      "rawMarkdown": "Please could you share the tf-records for the final dataset? Couldn't find it in the notebook..\n\nThanks and brilliant job!",
      "replies": [
        {
          "id": 876173,
          "postDate": "2020-06-06T13:55:02.467Z",
          "content": "<p>Looks like I found it.</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">https://www.kaggle.com/cdeotte/how-to-create-tfrecords</a></p>\n\n<p>Anyways thanks!</p>",
          "rawMarkdown": "Looks like I found it.\n\nhttps://www.kaggle.com/cdeotte/how-to-create-tfrecords\n\nAnyways thanks!"
        }
      ]
    },
    {
      "id": 878304,
      "postDate": "2020-06-08T12:29:04.373Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 878386,
          "postDate": "2020-06-08T13:44:10.290Z",
          "content": "<p><a href=\"/awsaf49ieee\">@awsaf49ieee</a> I have added <code>folds_08062020.csv</code> with ISIC2019 &amp; ISIC2020. I did't change images and other files to avoid conflicts of versions</p>",
          "rawMarkdown": "@awsaf49ieee I have added `folds_08062020.csv` with ISIC2019 &amp; ISIC2020. I did't change images and other files to avoid conflicts of versions"
        }
      ]
    },
    {
      "id": 876670,
      "postDate": "2020-06-06T22:17:13.593Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 918401,
      "postDate": "2020-07-07T08:30:48.680Z",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!"
    }
  ],
  "comments": [
    {
      "id": 878163,
      "author_name": "Amirreza Mahbod",
      "author_url": "",
      "post_date": "2020-06-08T10:06:50.290000",
      "content": "<p>Please be aware of some duplicated images in the marged dataset that may has negative effect on the CV. This is mainly due to the naming of the ISIC 2019 images which some has \"_downsampled\" term in the name. \nExample of duplicate images:\nISIC_0000036.jpg (fold 0)\nISIC_0000036_downsampled.jpg (fold 1)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 878178,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-08T10:22:54.760000",
          "content": "<p>thank you! Sure, It needs some corrections :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 878183,
          "author_name": "Amirreza Mahbod",
          "author_url": "",
          "post_date": "2020-06-08T10:35:50.563000",
          "content": "<p>I just had a closer look at the datasets. I think it is enough to use the ISIC 2019 dataset (and not use the former datasets (i.e. ISIC 2018, ISIC 2017, ISIC 2016) as there is almost no new images in the former datasets. \nThe only benefit of using former datasets is that they have the images in the original resolutions and not in the downsampled format. Anyway as we may be forced to resize the images for training then even this benefit will not be useful. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 969479,
              "author_name": "Anthony Leo",
              "author_url": "",
              "post_date": "2020-08-13T17:55:49.870000",
              "content": "<p>Hello ,<br>\n2019 competition winner have used multiple dataset ( <a href=\"https://arxiv.org/pdf/1910.03910.pdf\" target=\"_blank\">https://arxiv.org/pdf/1910.03910.pdf</a> ) ,</p>\n<ul>\n<li>HAM dataset</li>\n<li>7 point dataset</li>\n<li>MSK Dataset</li>\n<li>BCN 20000 Dataset</li>\n</ul>\n<p>Why these dataset not considered as external dataset ?</p>\n<p>From Paper,<br>\n A part of the training dataset is the HAM10000 dataset which contains images of<br>\nsize 600 × 450 that were centered and cropped around the lesion. The dataset<br>\ncurators applied histogram corrections to some images [16]. Another dataset,<br>\nBCN 20000, contains images of size 1024 × 1024. This dataset is particularly<br>\nchallenging as many images are uncropped and lesions in difficult and uncommon<br>\nlocations are present [5]. Last, the MSK dataset contains images with various<br>\nsizes.<br>\nThe dataset also contains meta information about the patient’s age group<br>\n(in steps of five years), the anatomical site (eight possible sites) and the sex<br>\n(male/female). The meta data is partially incomplete, i.e., there are missing<br>\nvalues for some images.<br>\nIn addition, we make use of external data. We use the 995 dermoscopic images<br>\nfrom the 7-point dataset [2]. </p>\n<p>Thanks,<br>\nAnthony</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 878250,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-08T11:38:03.460000",
          "content": "<p>I have added fixes, thank you! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 872665,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-06-03T12:09:40.317000",
      "content": "<p>Nice.\nCould you please clarify how you performed resizing? Was it simple resize with a loss of original aspect ratio or was it crop -&gt; resize, keeping an original aspect ratio?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 872672,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-03T12:16:19.643000",
          "content": "<p>Thanks a lot!</p>\n\n<p>I have used cv2.INTER_AREA without crop and saving ratio, you can see full code <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">here</a> ;) if need you can change resizing method. </p>\n\n<p>Welcome!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 872784,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-06-03T13:59:55.517000",
      "content": "<p>great work. </p>\n\n<p><a href=\"/shonenkov\">@shonenkov</a> hi, can you please ensure the data leakage issue. I think it's really important. I've also faced some issues (like other reporters). Thank you.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1608565,
      "author_name": "zhujjie",
      "author_url": "",
      "post_date": "2021-12-06T12:57:56.477000",
      "content": "<p>Can you explain each file？I don't know what these .csv file represents？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 918796,
      "author_name": "Harsh Nagouda",
      "author_url": "",
      "post_date": "2020-07-07T13:56:53.147000",
      "content": "<p>This should be useful for some experimentation. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 878898,
      "author_name": "xiaopeng",
      "author_url": "",
      "post_date": "2020-06-09T02:08:53.330000",
      "content": "<p>I only use  ISIC2019 &amp; ISIC2020,but got overfitting,val 0.9269  LB0.887,do you know the reason?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 878945,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-06-09T03:58:28.657000",
          "content": "<ol>\n<li>data leakage</li>\n<li>val data is different from test data\nI guess....</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 879438,
          "author_name": "xiaopeng",
          "author_url": "",
          "post_date": "2020-06-09T13:24:06.137000",
          "content": "<p>now my val score 0.9168  lb score 0.917 ,I only fix the channel bgr -&gt;rgb </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 879448,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-06-09T13:31:31.210000",
          "content": "<p>For all the images?? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 879453,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-06-09T13:37:56.740000",
          "content": "<p>Images are already in RGB according to dataset kernel</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 879485,
          "author_name": "xiaopeng",
          "author_url": "",
          "post_date": "2020-06-09T14:04:47",
          "content": "<p>ISIC2019 &amp; ISIC2020 the left img is RGB format,but overfit maybe is not the reason.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3132358%2F8b039fc9631dce853f8915576c39951c%2F2020-06-09%2022-01-49.png?generation=1591711413752119&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 879536,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-06-09T14:38:30.510000",
          "content": "<p>I'm still having trouble getting it. Could you please clarify it...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880314,
          "author_name": "Amirreza Mahbod",
          "author_url": "",
          "post_date": "2020-06-10T07:11:23.460000",
          "content": "<p>when you open the images using <code>cv2.imread</code>, by default, it has the BGR format. So if you need RGB format you need to switch the blue and red channels. However, if you have all the training and test images in BGR format, I assume it will not affect your final score. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 923334,
          "author_name": "Licheng Zhang",
          "author_url": "",
          "post_date": "2020-07-10T18:19:20.583000",
          "content": "<p>I have the same overfit problem, I could try this…but like others said this  change of channel should not matter.…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 876168,
      "author_name": "DeepWilson",
      "author_url": "",
      "post_date": "2020-06-06T13:51:40.040000",
      "content": "<p>Please could you share the tf-records for the final dataset? Couldn't find it in the notebook..</p>\n\n<p>Thanks and brilliant job!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 876173,
          "author_name": "DeepWilson",
          "author_url": "",
          "post_date": "2020-06-06T13:55:02.467000",
          "content": "<p>Looks like I found it.</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">https://www.kaggle.com/cdeotte/how-to-create-tfrecords</a></p>\n\n<p>Anyways thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 878304,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-08T12:29:04.373000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 878386,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-08T13:44:10.290000",
          "content": "<p><a href=\"/awsaf49ieee\">@awsaf49ieee</a> I have added <code>folds_08062020.csv</code> with ISIC2019 &amp; ISIC2020. I did't change images and other files to avoid conflicts of versions</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 876670,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-06T22:17:13.593000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 918401,
      "author_name": "Sohag Kumar Mondal",
      "author_url": "",
      "post_date": "2020-07-07T08:30:48.680000",
      "content": "<p>Thank you very much!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "872528": "Hi everyone!\n\nI have created [kernel with merge external data](https://www.kaggle.com/shonenkov/merge-external-data) (including metadata):\n\n- [Melanoma Detection Dataset](https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection)\n- [Skin Lesion Images for Melanoma Classification](https://www.kaggle.com/andrewmvd/isic-2019)\n- [Skin Cancer MNIST: HAM10000](https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000)\n- [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification/data)\n\nI have got 53351 samples for class 0 and 5106 for class 1. \n\nMerged dataset you can find [here](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg)\n- format: jpeg, rgb\n- size: 512x512\n\nif you need other format of data (tf-records or other image size) you can use [my kernel](https://www.kaggle.com/shonenkov/merge-external-data)\n\nI hope it helps you!\n\nWelcome!\n",
    "878163": "Please be aware of some duplicated images in the marged dataset that may has negative effect on the CV. This is mainly due to the naming of the ISIC 2019 images which some has \"_downsampled\" term in the name. \nExample of duplicate images:\nISIC_0000036.jpg (fold 0)\nISIC_0000036_downsampled.jpg (fold 1)",
    "872665": "Nice.\nCould you please clarify how you performed resizing? Was it simple resize with a loss of original aspect ratio or was it crop -&gt; resize, keeping an original aspect ratio?",
    "872784": "great work. \n\n@shonenkov hi, can you please ensure the data leakage issue. I think it's really important. I've also faced some issues (like other reporters). Thank you.",
    "1608565": "Can you explain each file？I don't know what these .csv file represents？",
    "918796": "This should be useful for some experimentation. Thank you!\n",
    "878898": "I only use  ISIC2019 &amp; ISIC2020,but got overfitting,val 0.9269  LB0.887,do you know the reason?",
    "876168": "Please could you share the tf-records for the final dataset? Couldn't find it in the notebook..\n\nThanks and brilliant job!",
    "878304": "",
    "876670": "",
    "918401": "Thank you very much!"
  }
}