{
  "id": 168028,
  "title": "Mystery Images - RAPIDS cuML TSNE",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168028",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-18T23:25:39.581000",
  "votes": 121,
  "comment_count": 22,
  "views": 0,
  "content": "<h1>RAPIDS cuML TSNE</h1>\n\n<p>Using RAPIDS cuML t-SNE, we can project CNN image embeddings into the <code>x-y-plane</code>. (Detailed explanation <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\">here</a>, notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>). Each dot represents one image in the competition train or test dataset. Orange are train benign, blue are train malignant, and green are test. When 2 points are close then these images are similar.</p>\n\n<pre><code> import cuml\n model = cuml.TSNE()\n embed_all = np.concatenate([embed_train,embed_test,embed_ext])\n embed2D = model.fit_transform(embed_all)\n # TRAIN\n train['x'] = embed2D[:len(embed_train),0]\n train['y'] = embed2D[:len(embed_train),1]\n # TEST\n test['x'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),0]\n test['y'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),1]\n # EXTERNAL\n train_ext['x'] = embed2D[len(embed_train)+len(embed_test):,0]\n train_ext['y'] = embed2D[len(embed_train)+len(embed_test):,1]\n</code></pre>\n\n<h1>Mystery Images</h1>\n\n<h2>Train vs. Test</h2>\n\n<p>In the below plot, i have marked 3 rectangular regions. These regions contain test images but no train images. That means there are test images with no train images that look like them. Will our models generalize to these <strong>mystery</strong> images?? (These mystery images are 15% of all test images !!)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fe83d2ffd3ca9be6c7993a24fc3cc9b69%2FScreen%20Shot%202020-07-18%20at%203.55.21%20PM.png?generation=1595113174837597&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 1 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fddd42a3286a7d72b9f3e536c007fbf02%2F1.png?generation=1595113206584644&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 2 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F19dbcdbb71ecae22d1b355a77aa7ea73%2F2.png?generation=1595113218190890&amp;alt=media\" alt=\"\"></p>\n\n<h2>Train vs. 2019 Comp</h2>\n\n<p>In the below plot, i have marked 3 rectangular regions. Green is 2019 comp benign and red is 2019 comp malignant. These regions contain 2019 comp images (external data) but no train or test images (current comp data). That means that this external data is different than our comps data. Will using these <strong>mystery</strong> images as additional training data help our models?? (Note this plot is the new half of 2019. The old half of 2019 is the 2018 2017 comp data).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fdcd8cc7347a533ba5879f353ebb36b53%2FScreen%20Shot%202020-07-18%20at%204.01.14%20PM.png?generation=1595113338072729&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 4 - 10.6% malignant - DOES NOT EXIST IN TRAIN OR TEST</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F767f9ecb46655f52079af4cf5ba7e041%2F1b.png?generation=1595113354036056&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 5 - 38.7% malignant - DOES NOT EXIST IN TRAIN OR TEST</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9055ccd50dacee186c887b6a8ad75e1a%2F2b.png?generation=1595113390464894&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 6 - 24.8% malignant - DOES NOT EXIST IN TRAIN OR TEST</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F64f7fd9c52ccf38cea768b6bdaf801e2%2F3b.png?generation=1595113409505076&amp;alt=media\" alt=\"\"></p>\n\n<h2>Confirmation</h2>\n\n<p>We can double check our RAPIDS cuML TNSE discovery, by plotting a simple histogram of image average color. (I got this idea from Ertugrul's great EDA notebook <a href=\"https://www.kaggle.com/datafan07/starter-analysis-of-melanoma-metadata-and-images/\">here</a>). Each image is <code>256x256x3</code> numbers where each number is between <code>0 and 255</code>. For each train and test image, we take the average value and plot the histogram of all train and all test. We do this <strong>with</strong> and <strong>without</strong> the mystery test images\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc70d572a27aaa41028143de1433650c3%2FScreen%20Shot%202020-07-18%20at%206.01.44%20PM.png?generation=1595120523032662&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 934932,
      "postDate": "2020-07-18T23:25:39.580Z",
      "content": "<h1>RAPIDS cuML TSNE</h1>\n\n<p>Using RAPIDS cuML t-SNE, we can project CNN image embeddings into the <code>x-y-plane</code>. (Detailed explanation <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\">here</a>, notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>). Each dot represents one image in the competition train or test dataset. Orange are train benign, blue are train malignant, and green are test. When 2 points are close then these images are similar.</p>\n\n<pre><code> import cuml\n model = cuml.TSNE()\n embed_all = np.concatenate([embed_train,embed_test,embed_ext])\n embed2D = model.fit_transform(embed_all)\n # TRAIN\n train['x'] = embed2D[:len(embed_train),0]\n train['y'] = embed2D[:len(embed_train),1]\n # TEST\n test['x'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),0]\n test['y'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),1]\n # EXTERNAL\n train_ext['x'] = embed2D[len(embed_train)+len(embed_test):,0]\n train_ext['y'] = embed2D[len(embed_train)+len(embed_test):,1]\n</code></pre>\n\n<h1>Mystery Images</h1>\n\n<h2>Train vs. Test</h2>\n\n<p>In the below plot, i have marked 3 rectangular regions. These regions contain test images but no train images. That means there are test images with no train images that look like them. Will our models generalize to these <strong>mystery</strong> images?? (These mystery images are 15% of all test images !!)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fe83d2ffd3ca9be6c7993a24fc3cc9b69%2FScreen%20Shot%202020-07-18%20at%203.55.21%20PM.png?generation=1595113174837597&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 1 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fddd42a3286a7d72b9f3e536c007fbf02%2F1.png?generation=1595113206584644&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 2 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F19dbcdbb71ecae22d1b355a77aa7ea73%2F2.png?generation=1595113218190890&amp;alt=media\" alt=\"\"></p>\n\n<h2>Train vs. 2019 Comp</h2>\n\n<p>In the below plot, i have marked 3 rectangular regions. Green is 2019 comp benign and red is 2019 comp malignant. These regions contain 2019 comp images (external data) but no train or test images (current comp data). That means that this external data is different than our comps data. Will using these <strong>mystery</strong> images as additional training data help our models?? (Note this plot is the new half of 2019. The old half of 2019 is the 2018 2017 comp data).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fdcd8cc7347a533ba5879f353ebb36b53%2FScreen%20Shot%202020-07-18%20at%204.01.14%20PM.png?generation=1595113338072729&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 4 - 10.6% malignant - DOES NOT EXIST IN TRAIN OR TEST</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F767f9ecb46655f52079af4cf5ba7e041%2F1b.png?generation=1595113354036056&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 5 - 38.7% malignant - DOES NOT EXIST IN TRAIN OR TEST</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9055ccd50dacee186c887b6a8ad75e1a%2F2b.png?generation=1595113390464894&amp;alt=media\" alt=\"\"></p>\n\n<h3>Region 6 - 24.8% malignant - DOES NOT EXIST IN TRAIN OR TEST</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F64f7fd9c52ccf38cea768b6bdaf801e2%2F3b.png?generation=1595113409505076&amp;alt=media\" alt=\"\"></p>\n\n<h2>Confirmation</h2>\n\n<p>We can double check our RAPIDS cuML TNSE discovery, by plotting a simple histogram of image average color. (I got this idea from Ertugrul's great EDA notebook <a href=\"https://www.kaggle.com/datafan07/starter-analysis-of-melanoma-metadata-and-images/\">here</a>). Each image is <code>256x256x3</code> numbers where each number is between <code>0 and 255</code>. For each train and test image, we take the average value and plot the histogram of all train and all test. We do this <strong>with</strong> and <strong>without</strong> the mystery test images\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc70d572a27aaa41028143de1433650c3%2FScreen%20Shot%202020-07-18%20at%206.01.44%20PM.png?generation=1595120523032662&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# RAPIDS cuML TSNE\nUsing RAPIDS cuML t-SNE, we can project CNN image embeddings into the `x-y-plane`. (Detailed explanation [here][1], notebook [here][2]). Each dot represents one image in the competition train or test dataset. Orange are train benign, blue are train malignant, and green are test. When 2 points are close then these images are similar.\n  \n\n     import cuml\n     model = cuml.TSNE()\n     embed_all = np.concatenate([embed_train,embed_test,embed_ext])\n     embed2D = model.fit_transform(embed_all)\n     # TRAIN\n     train['x'] = embed2D[:len(embed_train),0]\n     train['y'] = embed2D[:len(embed_train),1]\n     # TEST\n     test['x'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),0]\n     test['y'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),1]\n     # EXTERNAL\n     train_ext['x'] = embed2D[len(embed_train)+len(embed_test):,0]\n     train_ext['y'] = embed2D[len(embed_train)+len(embed_test):,1]\n\n# Mystery Images\n\n## Train vs. Test\nIn the below plot, i have marked 3 rectangular regions. These regions contain test images but no train images. That means there are test images with no train images that look like them. Will our models generalize to these **mystery** images?? (These mystery images are 15% of all test images !!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fe83d2ffd3ca9be6c7993a24fc3cc9b69%2FScreen%20Shot%202020-07-18%20at%203.55.21%20PM.png?generation=1595113174837597&amp;alt=media)\n### Region 1 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fddd42a3286a7d72b9f3e536c007fbf02%2F1.png?generation=1595113206584644&amp;alt=media)\n### Region 2 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F19dbcdbb71ecae22d1b355a77aa7ea73%2F2.png?generation=1595113218190890&amp;alt=media)\n\n## Train vs. 2019 Comp\nIn the below plot, i have marked 3 rectangular regions. Green is 2019 comp benign and red is 2019 comp malignant. These regions contain 2019 comp images (external data) but no train or test images (current comp data). That means that this external data is different than our comps data. Will using these **mystery** images as additional training data help our models?? (Note this plot is the new half of 2019. The old half of 2019 is the 2018 2017 comp data).\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fdcd8cc7347a533ba5879f353ebb36b53%2FScreen%20Shot%202020-07-18%20at%204.01.14%20PM.png?generation=1595113338072729&amp;alt=media)\n### Region 4 - 10.6% malignant - DOES NOT EXIST IN TRAIN OR TEST\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F767f9ecb46655f52079af4cf5ba7e041%2F1b.png?generation=1595113354036056&amp;alt=media)\n### Region 5 - 38.7% malignant - DOES NOT EXIST IN TRAIN OR TEST\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9055ccd50dacee186c887b6a8ad75e1a%2F2b.png?generation=1595113390464894&amp;alt=media)\n### Region 6 - 24.8% malignant - DOES NOT EXIST IN TRAIN OR TEST\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F64f7fd9c52ccf38cea768b6bdaf801e2%2F3b.png?generation=1595113409505076&amp;alt=media)\n\n## Confirmation\nWe can double check our RAPIDS cuML TNSE discovery, by plotting a simple histogram of image average color. (I got this idea from Ertugrul's great EDA notebook [here][3]). Each image is `256x256x3` numbers where each number is between `0 and 255`. For each train and test image, we take the average value and plot the histogram of all train and all test. We do this **with** and **without** the mystery test images\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc70d572a27aaa41028143de1433650c3%2FScreen%20Shot%202020-07-18%20at%206.01.44%20PM.png?generation=1595120523032662&amp;alt=media)\n\n\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\n[2]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[3]: https://www.kaggle.com/datafan07/starter-analysis-of-melanoma-metadata-and-images/",
      "votes": 120
    },
    {
      "id": 945537,
      "postDate": "2020-07-25T23:40:48.977Z",
      "content": "<p>Judging by the t-sne plots and combining it with my limited understanding, I'd say there is something that makes the new portion of the 2019 data different from the rest.</p>\n\n<p>Looking at the average color plots, I'd say some extra preprocessing or data augmentation steps might help.\nThe first thing that comes to my mind is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876\">color constancy</a> as a preprocessing step, or maybe random hsv, brightness, contrast augmentations, which might reduce the distance of these samples.</p>",
      "rawMarkdown": "Judging by the t-sne plots and combining it with my limited understanding, I'd say there is something that makes the new portion of the 2019 data different from the rest.\n\nLooking at the average color plots, I'd say some extra preprocessing or data augmentation steps might help.\nThe first thing that comes to my mind is [color constancy](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876) as a preprocessing step, or maybe random hsv, brightness, contrast augmentations, which might reduce the distance of these samples.",
      "votes": 1
    },
    {
      "id": 938717,
      "postDate": "2020-07-21T17:42:01.513Z",
      "content": "<p>Amazing work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> especially finding out the mystery images.  I had a question about model generalization since these test images are completely new wouldn't a metric learning-based approach be preferred, essentially mapping similar images closer to \"prototypes\" of the labels in a high dimensional plane?. What are your thoughts? Seems like interesting data to apply \"manifold learning\" !! I have been experimenting with metric learning but I am not able to get good results 😓, but I'm still gonna try pushing and keep experimenting. </p>",
      "rawMarkdown": "Amazing work @cdeotte especially finding out the mystery images.  I had a question about model generalization since these test images are completely new wouldn't a metric learning-based approach be preferred, essentially mapping similar images closer to \"prototypes\" of the labels in a high dimensional plane?. What are your thoughts? Seems like interesting data to apply \"manifold learning\" !! I have been experimenting with metric learning but I am not able to get good results 😓, but I'm still gonna try pushing and keep experimenting. ",
      "votes": 1
    },
    {
      "id": 938649,
      "postDate": "2020-07-21T16:45:20.273Z",
      "content": "<p>This is an amazing work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , looking at that distribution, it seems to be a good idea to only sample older data that has a similar distribution, it may be the reason why using the complete data was not working well for me</p>",
      "rawMarkdown": "This is an amazing work @cdeotte , looking at that distribution, it seems to be a good idea to only sample older data that has a similar distribution, it may be the reason why using the complete data was not working well for me",
      "votes": 1,
      "replies": [
        {
          "id": 938675,
          "postDate": "2020-07-21T17:06:44.727Z",
          "content": "<p>Yes. When i use all of 2019 comp data it decreases my CV LB. If you wish to only use 2018 2017, those are the even numbered TFRecords in my external Kaggle dataset.</p>\n<p>I can explain why most people are fooled. The 2019 data is half \"new 2019\" and half \"old 2019\" (which is 2018 2017). The \"new 2019\" data is easy to classify, so if you add the \"new 2019\" data to your validation folds, then your CV score goes up. But that isn't really your CV score.</p>\n<p>All experiments should use the same validation. So we should not add external data to our validation folds. We should always use 2020 comp data for validation. When we do this, we see that using \"new 2019\" (for training) decreases CV.</p>",
          "rawMarkdown": "Yes. When i use all of 2019 comp data it decreases my CV LB. If you wish to only use 2018 2017, those are the even numbered TFRecords in my external Kaggle dataset.\n\nI can explain why most people are fooled. The 2019 data is half \"new 2019\" and half \"old 2019\" (which is 2018 2017). The \"new 2019\" data is easy to classify, so if you add the \"new 2019\" data to your validation folds, then your CV score goes up. But that isn't really your CV score.\n\nAll experiments should use the same validation. So we should not add external data to our validation folds. We should always use 2020 comp data for validation. When we do this, we see that using \"new 2019\" (for training) decreases CV.",
          "votes": 2
        },
        {
          "id": 938731,
          "postDate": "2020-07-21T17:57:44.530Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , I was going to search on how to get only old data from your TFRecs.</p>\n<p>I did some experiments using part of 2020 data as validation set and all 2019 + the rest of 2020 data as train set, and saw no improvements, this was probably because as you said I also added the new and easy 2019 data to the train set, maybe I should try to remove those samples.</p>",
          "rawMarkdown": "Thanks @cdeotte , I was going to search on how to get only old data from your TFRecs.\n\nI did some experiments using part of 2020 data as validation set and all 2019 + the rest of 2020 data as train set, and saw no improvements, this was probably because as you said I also added the new and easy 2019 data to the train set, maybe I should try to remove those samples."
        },
        {
          "id": 938747,
          "postDate": "2020-07-21T18:15:55.867Z",
          "content": "<p>When i train with 2020 plus \"old 2019\", it improves my CV LB</p>",
          "rawMarkdown": "When i train with 2020 plus \"old 2019\", it improves my CV LB",
          "votes": 3
        },
        {
          "id": 939409,
          "postDate": "2020-07-22T08:17:52.113Z",
          "content": "<p>How interesting! Do you estimate this deterioration due to quality (\"<a href=\"https://arxiv.org/abs/1908.02288\" target=\"_blank\">Lesions in the <em>wild</em></a>\"), or \"easy-ness\" (to overfit?) of \"new ISIC2019\"? </p>\n<p>Basically, as far as I see, \"old 2019\" is the <a href=\"https://doi.org/10.1038/sdata.2018.161\" target=\"_blank\">HAM10000</a> dataset, and <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\" target=\"_blank\">MSK- &amp; UDA-dataset(s) from the ISIC-archive</a>. I guess things got complicated with the overlap of datasets and challenge-datasets in 2019, and the non-availability of \"new ISIC2019\" through the official ISIC-archive. My guess is that a lot of confusion would be omitted by using the original dataset terms.</p>",
          "rawMarkdown": "How interesting! Do you estimate this deterioration due to quality (\"[Lesions in the *wild*](https://arxiv.org/abs/1908.02288)\"), or \"easy-ness\" (to overfit?) of \"new ISIC2019\"? \n\nBasically, as far as I see, \"old 2019\" is the [HAM10000](https://doi.org/10.1038/sdata.2018.161) dataset, and [MSK- &amp; UDA-dataset(s) from the ISIC-archive](https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery). I guess things got complicated with the overlap of datasets and challenge-datasets in 2019, and the non-availability of \"new ISIC2019\" through the official ISIC-archive. My guess is that a lot of confusion would be omitted by using the original dataset terms."
        }
      ]
    },
    {
      "id": 936663,
      "postDate": "2020-07-20T12:21:19.573Z",
      "content": "<p>It is some weeks that I'm wondering why it is so common to get very differently AUC from CV and LB. I have suspected that there are differences between test and train distribution, and your analysis, especially the histogram, support it. I don't think I'm a big expert, but I would say that this is one of the reasons why it is easier to get better results with many trials with ensembles. Then, we will see what happens with private LB (that could have a significantly different distribution).  </p>",
      "rawMarkdown": "It is some weeks that I'm wondering why it is so common to get very differently AUC from CV and LB. I have suspected that there are differences between test and train distribution, and your analysis, especially the histogram, support it. I don't think I'm a big expert, but I would say that this is one of the reasons why it is easier to get better results with many trials with ensembles. Then, we will see what happens with private LB (that could have a significantly different distribution).  ",
      "votes": 1
    },
    {
      "id": 936108,
      "postDate": "2020-07-20T01:31:04.450Z",
      "content": "<p>Nice Work!!!</p>",
      "rawMarkdown": "Nice Work!!!",
      "votes": 1
    },
    {
      "id": 935397,
      "postDate": "2020-07-19T11:03:58.527Z",
      "content": "<ol>\n<li>What about only<code>2017-2018</code> Dataset?</li>\n<li>As <code>2019 (2017-18)</code> Dataset is different from <code>2020</code> then why result improves using it?</li>\n</ol>",
      "rawMarkdown": "1. What about only` 2017-2018` Dataset?\n2. As `2019 (2017-18)` Dataset is different from `2020` then why result improves using it?",
      "votes": 1,
      "replies": [
        {
          "id": 935835,
          "postDate": "2020-07-19T17:29:17.543Z",
          "content": "<p>The 2019 comp data is 25,000 images. Half of those images have original size <code>1024x1024</code> and those are the \"new 2019\", the other half of 2019 comp data is the 2018 2017 comp data. (In my notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> you can experiment just adding the \"new 2019\" or the \"old 2019 = 2018 + 2017\"). Have you seen an improvement using the \"new 2019\"? When i add \"new 2019\", it lowers my CV and LB. (In the plots above, I'm plotting the \"new 2019\").</p>\n\n<p>If you plot the \"old 2019 = 2018 + 2017\" data together with this years 2020 comp using RAPIDS cuML TNSE, you will see that they are similar. (The \"old 2019\" coincides with about half of this years 2020 data).</p>\n\n<p>I would add the plot to this discussion post, but every time you run TNSE, you get a different plot, so i would have to redo everything. I plotted the \"old 2019\" compared with \"this year 2020\" data in a comment to my discussion post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551#932485\">here</a>. Scroll the comments and look for the plots.</p>",
          "rawMarkdown": "The 2019 comp data is 25,000 images. Half of those images have original size `1024x1024` and those are the \"new 2019\", the other half of 2019 comp data is the 2018 2017 comp data. (In my notebook [here][1] you can experiment just adding the \"new 2019\" or the \"old 2019 = 2018 + 2017\"). Have you seen an improvement using the \"new 2019\"? When i add \"new 2019\", it lowers my CV and LB. (In the plots above, I'm plotting the \"new 2019\").\n\nIf you plot the \"old 2019 = 2018 + 2017\" data together with this years 2020 comp using RAPIDS cuML TNSE, you will see that they are similar. (The \"old 2019\" coincides with about half of this years 2020 data).\n\nI would add the plot to this discussion post, but every time you run TNSE, you get a different plot, so i would have to redo everything. I plotted the \"old 2019\" compared with \"this year 2020\" data in a comment to my discussion post [here][2]. Scroll the comments and look for the plots.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551#932485",
          "votes": 4
        },
        {
          "id": 936948,
          "postDate": "2020-07-20T16:10:53.870Z",
          "content": "<p>hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for sharing this.</p>\n<p>Why is that that we can't reproduce the exact same plot? is this due to t-sne? or cuML implementation?</p>\n<p>I think UMAP is more used nowadays than t-SNE and it seems that there is an cuML implementation of UMAP, is there a specific reason you chose to work with t-SNE instead of UMAP?</p>\n<p>Thanks!</p>",
          "rawMarkdown": "hello @cdeotte, thanks for sharing this.\n\nWhy is that that we can't reproduce the exact same plot? is this due to t-sne? or cuML implementation?\n\nI think UMAP is more used nowadays than t-SNE and it seems that there is an cuML implementation of UMAP, is there a specific reason you chose to work with t-SNE instead of UMAP?\n\nThanks!",
          "votes": 1
        },
        {
          "id": 937000,
          "postDate": "2020-07-20T16:45:34.477Z",
          "content": "<p>Both t-sne and umap produce different plots each time you run them. I think this is because they use random sampling in the algorithms. </p>\n\n<p>I didn't choose t-sne for any specific reason, it is just the one i have used more often (and the pictures look prettier). Below is a UMAP plot of 2020 train data, 2019 new portion data, 2018 2017 data. We see that 2019 new portion adds a whole new island to the picture</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F12368ea4b1ce87a2f878eac4249993e8%2FScreen%20Shot%202020-07-20%20at%209.42.17%20AM.png?generation=1595263779145169&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Both t-sne and umap produce different plots each time you run them. I think this is because they use random sampling in the algorithms. \n\nI didn't choose t-sne for any specific reason, it is just the one i have used more often (and the pictures look prettier). Below is a UMAP plot of 2020 train data, 2019 new portion data, 2018 2017 data. We see that 2019 new portion adds a whole new island to the picture\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F12368ea4b1ce87a2f878eac4249993e8%2FScreen%20Shot%202020-07-20%20at%209.42.17%20AM.png?generation=1595263779145169&amp;alt=media)",
          "votes": 2
        },
        {
          "id": 937023,
          "postDate": "2020-07-20T17:10:25.733Z",
          "content": "<p>thanks! Just wanted to be sure I was not missing something important!</p>",
          "rawMarkdown": "thanks! Just wanted to be sure I was not missing something important!",
          "votes": 1
        }
      ]
    },
    {
      "id": 934935,
      "postDate": "2020-07-18T23:38:46.137Z",
      "content": "<p>Amazing work! <a href=\"/cdeotte\">@cdeotte</a>  </p>",
      "rawMarkdown": "Amazing work! @cdeotte  ",
      "votes": 1
    },
    {
      "id": 938821,
      "postDate": "2020-07-21T19:19:59.007Z",
      "content": "<p>I think it will be cool to make an adversarial test on ROC-AUC metric with 2020 versus new 2019, 2020 versus old 2019 and maybe some other setups. \nBut your investigations give a good intuition about the results of such checks</p>",
      "rawMarkdown": "I think it will be cool to make an adversarial test on ROC-AUC metric with 2020 versus new 2019, 2020 versus old 2019 and maybe some other setups. \nBut your investigations give a good intuition about the results of such checks",
      "votes": 2
    },
    {
      "id": 934971,
      "postDate": "2020-07-19T01:04:01.870Z",
      "content": "<p>UPDATE: I posted a confirmation histogram using image average color.</p>",
      "rawMarkdown": "UPDATE: I posted a confirmation histogram using image average color.",
      "votes": 2
    },
    {
      "id": 968813,
      "postDate": "2020-08-13T09:00:18.987Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> how did you extract the samples that are present in test but not in train? So I mean after embedding the images and then visualizing the regions how did you obtain the images from the embedding regions with no training samples present? (eg the ones that you showed above)</p>",
      "rawMarkdown": "@cdeotte how did you extract the samples that are present in test but not in train? So I mean after embedding the images and then visualizing the regions how did you obtain the images from the embedding regions with no training samples present? (eg the ones that you showed above)"
    },
    {
      "id": 934945,
      "postDate": "2020-07-19T00:04:28.650Z",
      "content": "<p>what if we use GAN to produce more examples ??</p>",
      "rawMarkdown": "what if we use GAN to produce more examples ??",
      "replies": [
        {
          "id": 934956,
          "postDate": "2020-07-19T00:32:12.163Z",
          "content": "<p>That's a cool idea. You could use a GAN to fill in the white space and get fuller coverage of the image training space.</p>",
          "rawMarkdown": "That's a cool idea. You could use a GAN to fill in the white space and get fuller coverage of the image training space.",
          "votes": 6
        }
      ]
    },
    {
      "id": 943677,
      "postDate": "2020-07-24T14:10:51.010Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 960576,
      "postDate": "2020-08-06T14:12:07.353Z",
      "content": "<p>Thanks. Nice work.</p>",
      "rawMarkdown": "Thanks. Nice work.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 945537,
      "author_name": "Stephan",
      "author_url": "",
      "post_date": "2020-07-25T23:40:48.977000",
      "content": "<p>Judging by the t-sne plots and combining it with my limited understanding, I'd say there is something that makes the new portion of the 2019 data different from the rest.</p>\n\n<p>Looking at the average color plots, I'd say some extra preprocessing or data augmentation steps might help.\nThe first thing that comes to my mind is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876\">color constancy</a> as a preprocessing step, or maybe random hsv, brightness, contrast augmentations, which might reduce the distance of these samples.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 938717,
      "author_name": "Atharva Phatak",
      "author_url": "",
      "post_date": "2020-07-21T17:42:01.513000",
      "content": "<p>Amazing work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> especially finding out the mystery images.  I had a question about model generalization since these test images are completely new wouldn't a metric learning-based approach be preferred, essentially mapping similar images closer to \"prototypes\" of the labels in a high dimensional plane?. What are your thoughts? Seems like interesting data to apply \"manifold learning\" !! I have been experimenting with metric learning but I am not able to get good results 😓, but I'm still gonna try pushing and keep experimenting. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 938649,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-07-21T16:45:20.273000",
      "content": "<p>This is an amazing work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , looking at that distribution, it seems to be a good idea to only sample older data that has a similar distribution, it may be the reason why using the complete data was not working well for me</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938675,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-21T17:06:44.727000",
          "content": "<p>Yes. When i use all of 2019 comp data it decreases my CV LB. If you wish to only use 2018 2017, those are the even numbered TFRecords in my external Kaggle dataset.</p>\n<p>I can explain why most people are fooled. The 2019 data is half \"new 2019\" and half \"old 2019\" (which is 2018 2017). The \"new 2019\" data is easy to classify, so if you add the \"new 2019\" data to your validation folds, then your CV score goes up. But that isn't really your CV score.</p>\n<p>All experiments should use the same validation. So we should not add external data to our validation folds. We should always use 2020 comp data for validation. When we do this, we see that using \"new 2019\" (for training) decreases CV.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938731,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-07-21T17:57:44.530000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , I was going to search on how to get only old data from your TFRecs.</p>\n<p>I did some experiments using part of 2020 data as validation set and all 2019 + the rest of 2020 data as train set, and saw no improvements, this was probably because as you said I also added the new and easy 2019 data to the train set, maybe I should try to remove those samples.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938747,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-21T18:15:55.867000",
          "content": "<p>When i train with 2020 plus \"old 2019\", it improves my CV LB</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 939409,
          "author_name": "chdlr",
          "author_url": "",
          "post_date": "2020-07-22T08:17:52.113000",
          "content": "<p>How interesting! Do you estimate this deterioration due to quality (\"<a href=\"https://arxiv.org/abs/1908.02288\" target=\"_blank\">Lesions in the <em>wild</em></a>\"), or \"easy-ness\" (to overfit?) of \"new ISIC2019\"? </p>\n<p>Basically, as far as I see, \"old 2019\" is the <a href=\"https://doi.org/10.1038/sdata.2018.161\" target=\"_blank\">HAM10000</a> dataset, and <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\" target=\"_blank\">MSK- &amp; UDA-dataset(s) from the ISIC-archive</a>. I guess things got complicated with the overlap of datasets and challenge-datasets in 2019, and the non-availability of \"new ISIC2019\" through the official ISIC-archive. My guess is that a lot of confusion would be omitted by using the original dataset terms.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 936663,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2020-07-20T12:21:19.573000",
      "content": "<p>It is some weeks that I'm wondering why it is so common to get very differently AUC from CV and LB. I have suspected that there are differences between test and train distribution, and your analysis, especially the histogram, support it. I don't think I'm a big expert, but I would say that this is one of the reasons why it is easier to get better results with many trials with ensembles. Then, we will see what happens with private LB (that could have a significantly different distribution).  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936108,
      "author_name": "VarVelVol",
      "author_url": "",
      "post_date": "2020-07-20T01:31:04.450000",
      "content": "<p>Nice Work!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 935397,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2020-07-19T11:03:58.527000",
      "content": "<ol>\n<li>What about only<code>2017-2018</code> Dataset?</li>\n<li>As <code>2019 (2017-18)</code> Dataset is different from <code>2020</code> then why result improves using it?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 935835,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-19T17:29:17.543000",
          "content": "<p>The 2019 comp data is 25,000 images. Half of those images have original size <code>1024x1024</code> and those are the \"new 2019\", the other half of 2019 comp data is the 2018 2017 comp data. (In my notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> you can experiment just adding the \"new 2019\" or the \"old 2019 = 2018 + 2017\"). Have you seen an improvement using the \"new 2019\"? When i add \"new 2019\", it lowers my CV and LB. (In the plots above, I'm plotting the \"new 2019\").</p>\n\n<p>If you plot the \"old 2019 = 2018 + 2017\" data together with this years 2020 comp using RAPIDS cuML TNSE, you will see that they are similar. (The \"old 2019\" coincides with about half of this years 2020 data).</p>\n\n<p>I would add the plot to this discussion post, but every time you run TNSE, you get a different plot, so i would have to redo everything. I plotted the \"old 2019\" compared with \"this year 2020\" data in a comment to my discussion post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551#932485\">here</a>. Scroll the comments and look for the plots.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 936948,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-20T16:10:53.870000",
          "content": "<p>hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for sharing this.</p>\n<p>Why is that that we can't reproduce the exact same plot? is this due to t-sne? or cuML implementation?</p>\n<p>I think UMAP is more used nowadays than t-SNE and it seems that there is an cuML implementation of UMAP, is there a specific reason you chose to work with t-SNE instead of UMAP?</p>\n<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 937000,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-20T16:45:34.477000",
          "content": "<p>Both t-sne and umap produce different plots each time you run them. I think this is because they use random sampling in the algorithms. </p>\n\n<p>I didn't choose t-sne for any specific reason, it is just the one i have used more often (and the pictures look prettier). Below is a UMAP plot of 2020 train data, 2019 new portion data, 2018 2017 data. We see that 2019 new portion adds a whole new island to the picture</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F12368ea4b1ce87a2f878eac4249993e8%2FScreen%20Shot%202020-07-20%20at%209.42.17%20AM.png?generation=1595263779145169&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 937023,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-20T17:10:25.733000",
          "content": "<p>thanks! Just wanted to be sure I was not missing something important!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 934935,
      "author_name": "Aravind Ram Nathan",
      "author_url": "",
      "post_date": "2020-07-18T23:38:46.137000",
      "content": "<p>Amazing work! <a href=\"/cdeotte\">@cdeotte</a>  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 938821,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2020-07-21T19:19:59.007000",
      "content": "<p>I think it will be cool to make an adversarial test on ROC-AUC metric with 2020 versus new 2019, 2020 versus old 2019 and maybe some other setups. \nBut your investigations give a good intuition about the results of such checks</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 934971,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-19T01:04:01.870000",
      "content": "<p>UPDATE: I posted a confirmation histogram using image average color.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 968813,
      "author_name": "hrunic",
      "author_url": "",
      "post_date": "2020-08-13T09:00:18.987000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> how did you extract the samples that are present in test but not in train? So I mean after embedding the images and then visualizing the regions how did you obtain the images from the embedding regions with no training samples present? (eg the ones that you showed above)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 934945,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2020-07-19T00:04:28.650000",
      "content": "<p>what if we use GAN to produce more examples ??</p>",
      "votes": 0,
      "replies": [
        {
          "id": 934956,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-19T00:32:12.163000",
          "content": "<p>That's a cool idea. You could use a GAN to fill in the white space and get fuller coverage of the image training space.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 943677,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T14:10:51.010000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 960576,
      "author_name": "Anders Ericsson Gnosco",
      "author_url": "",
      "post_date": "2020-08-06T14:12:07.353000",
      "content": "<p>Thanks. Nice work.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "934932": "# RAPIDS cuML TSNE\nUsing RAPIDS cuML t-SNE, we can project CNN image embeddings into the `x-y-plane`. (Detailed explanation [here][1], notebook [here][2]). Each dot represents one image in the competition train or test dataset. Orange are train benign, blue are train malignant, and green are test. When 2 points are close then these images are similar.\n  \n\n     import cuml\n     model = cuml.TSNE()\n     embed_all = np.concatenate([embed_train,embed_test,embed_ext])\n     embed2D = model.fit_transform(embed_all)\n     # TRAIN\n     train['x'] = embed2D[:len(embed_train),0]\n     train['y'] = embed2D[:len(embed_train),1]\n     # TEST\n     test['x'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),0]\n     test['y'] = embed2D[len(embed_train):len(embed_train)+len(embed_test),1]\n     # EXTERNAL\n     train_ext['x'] = embed2D[len(embed_train)+len(embed_test):,0]\n     train_ext['y'] = embed2D[len(embed_train)+len(embed_test):,1]\n\n# Mystery Images\n\n## Train vs. Test\nIn the below plot, i have marked 3 rectangular regions. These regions contain test images but no train images. That means there are test images with no train images that look like them. Will our models generalize to these **mystery** images?? (These mystery images are 15% of all test images !!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fe83d2ffd3ca9be6c7993a24fc3cc9b69%2FScreen%20Shot%202020-07-18%20at%203.55.21%20PM.png?generation=1595113174837597&amp;alt=media)\n### Region 1 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fddd42a3286a7d72b9f3e536c007fbf02%2F1.png?generation=1595113206584644&amp;alt=media)\n### Region 2 - DOES NOT EXIST IN TRAIN BUT DOES EXIST IN TEST!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F19dbcdbb71ecae22d1b355a77aa7ea73%2F2.png?generation=1595113218190890&amp;alt=media)\n\n## Train vs. 2019 Comp\nIn the below plot, i have marked 3 rectangular regions. Green is 2019 comp benign and red is 2019 comp malignant. These regions contain 2019 comp images (external data) but no train or test images (current comp data). That means that this external data is different than our comps data. Will using these **mystery** images as additional training data help our models?? (Note this plot is the new half of 2019. The old half of 2019 is the 2018 2017 comp data).\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fdcd8cc7347a533ba5879f353ebb36b53%2FScreen%20Shot%202020-07-18%20at%204.01.14%20PM.png?generation=1595113338072729&amp;alt=media)\n### Region 4 - 10.6% malignant - DOES NOT EXIST IN TRAIN OR TEST\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F767f9ecb46655f52079af4cf5ba7e041%2F1b.png?generation=1595113354036056&amp;alt=media)\n### Region 5 - 38.7% malignant - DOES NOT EXIST IN TRAIN OR TEST\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9055ccd50dacee186c887b6a8ad75e1a%2F2b.png?generation=1595113390464894&amp;alt=media)\n### Region 6 - 24.8% malignant - DOES NOT EXIST IN TRAIN OR TEST\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F64f7fd9c52ccf38cea768b6bdaf801e2%2F3b.png?generation=1595113409505076&amp;alt=media)\n\n## Confirmation\nWe can double check our RAPIDS cuML TNSE discovery, by plotting a simple histogram of image average color. (I got this idea from Ertugrul's great EDA notebook [here][3]). Each image is `256x256x3` numbers where each number is between `0 and 255`. For each train and test image, we take the average value and plot the histogram of all train and all test. We do this **with** and **without** the mystery test images\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc70d572a27aaa41028143de1433650c3%2FScreen%20Shot%202020-07-18%20at%206.01.44%20PM.png?generation=1595120523032662&amp;alt=media)\n\n\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\n[2]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[3]: https://www.kaggle.com/datafan07/starter-analysis-of-melanoma-metadata-and-images/",
    "945537": "Judging by the t-sne plots and combining it with my limited understanding, I'd say there is something that makes the new portion of the 2019 data different from the rest.\n\nLooking at the average color plots, I'd say some extra preprocessing or data augmentation steps might help.\nThe first thing that comes to my mind is [color constancy](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876) as a preprocessing step, or maybe random hsv, brightness, contrast augmentations, which might reduce the distance of these samples.",
    "938717": "Amazing work @cdeotte especially finding out the mystery images.  I had a question about model generalization since these test images are completely new wouldn't a metric learning-based approach be preferred, essentially mapping similar images closer to \"prototypes\" of the labels in a high dimensional plane?. What are your thoughts? Seems like interesting data to apply \"manifold learning\" !! I have been experimenting with metric learning but I am not able to get good results 😓, but I'm still gonna try pushing and keep experimenting. ",
    "938649": "This is an amazing work @cdeotte , looking at that distribution, it seems to be a good idea to only sample older data that has a similar distribution, it may be the reason why using the complete data was not working well for me",
    "936663": "It is some weeks that I'm wondering why it is so common to get very differently AUC from CV and LB. I have suspected that there are differences between test and train distribution, and your analysis, especially the histogram, support it. I don't think I'm a big expert, but I would say that this is one of the reasons why it is easier to get better results with many trials with ensembles. Then, we will see what happens with private LB (that could have a significantly different distribution).  ",
    "936108": "Nice Work!!!",
    "935397": "1. What about only` 2017-2018` Dataset?\n2. As `2019 (2017-18)` Dataset is different from `2020` then why result improves using it?",
    "934935": "Amazing work! @cdeotte  ",
    "938821": "I think it will be cool to make an adversarial test on ROC-AUC metric with 2020 versus new 2019, 2020 versus old 2019 and maybe some other setups. \nBut your investigations give a good intuition about the results of such checks",
    "934971": "UPDATE: I posted a confirmation histogram using image average color.",
    "968813": "@cdeotte how did you extract the samples that are present in test but not in train? So I mean after embedding the images and then visualizing the regions how did you obtain the images from the embedding regions with no training samples present? (eg the ones that you showed above)",
    "934945": "what if we use GAN to produce more examples ??",
    "943677": "",
    "960576": "Thanks. Nice work."
  }
}