{
  "id": 167551,
  "title": "View Images with RAPIDS cuML TSNE",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/167551",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-17T04:02:12.848000",
  "votes": 136,
  "comment_count": 45,
  "views": 0,
  "content": "<h1>How CNN Works</h1>\n\n<p>A CNN is basically two parts. The bottom layers convert images into features and the top layers convert features into classification</p>\n\n<h1>Transfer Learning</h1>\n\n<p>When we download a pretrained CNN from the internet, we download the bottom layers for feature generation. Then we build our own top layers (custom head) and train them to convert features into classification for our specific problem (such as Melanoma Comp).</p>\n\n<h1>Feature Extraction</h1>\n\n<p>Instead of using a CNN for classification, we can just extract the features. Then for every image, we have a vector of numbers that represents the image. These features (also called image embeddings) have a <strong>magical property</strong>. When the distance between two image vectors is small then the images are similar. (Cosine similarity or Euclidean similarity). </p>\n\n<h1>RAPIDS cuML TSNE</h1>\n\n<p>If we extract features from EfficientNetB0, then we get vectors of length 1280. So each image is represented by a point in 1280 dimensional space. It's hard to visualize this but RAPIDS cuML TSNE can project this into 2 dimensions. Then we can plot each image as a dot in the <code>x-y-plane</code> below.</p>\n\n<h1>View Images</h1>\n\n<p>All dots that are close together are similar images. We can randomly select 10 regions and look at the images in those regions. We can also calculate the proportion of images in those regions that are malignant.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F6c9be2831918aae56af077e2bb604f57%2FScreen%20Shot%202020-07-16%20at%208.39.15%20PM.png?generation=1594957184583179&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 10 - 0.0% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F661d7ff283f66b9b035eda0cd93a904a%2FScreen%20Shot%202020-07-16%20at%208.45.30%20PM.png?generation=1594957547795576&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 9 - 1.5% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0a8694f40f085acb1de5fb6dfb99b6a0%2FScreen%20Shot%202020-07-16%20at%208.46.43%20PM.png?generation=1594957614965043&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 8 - 2.1% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F095cd7fa1d4ca3d2e6bb30652fdc8ab8%2FScreen%20Shot%202020-07-16%20at%208.47.34%20PM.png?generation=1594957665483816&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 7 - 4.0% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F523a1f26ec7a3f159b23c4d27b715262%2FScreen%20Shot%202020-07-16%20at%208.48.19%20PM.png?generation=1594957709694231&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 6 - 4.9% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffeae6e2c67abebc9cdb3fa1d8b9cd20a%2FScreen%20Shot%202020-07-16%20at%208.49.29%20PM.png?generation=1594957778977807&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 5 - 5.7% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F13cb7eaddef1599c089cc4150bbf3f52%2FScreen%20Shot%202020-07-16%20at%208.50.39%20PM.png?generation=1594957852045805&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 4 - 7.4% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcc393e67d21e132bc4eaa8d644d30b78%2FScreen%20Shot%202020-07-16%20at%208.51.23%20PM.png?generation=1594957893670098&amp;alt=media\" alt=\"\"></p>\n\n<h1>Region 3 - 9.4% malignant</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fa3b174517759d56220f215f748989182%2FScreen%20Shot%202020-07-16%20at%208.52.36%20PM.png?generation=1594957969518207&amp;alt=media\" alt=\"\"></p>\n\n<h1>Region 2 - 10.8% malignant</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5549e7e55ea222368a3c705328d1ece9%2FScreen%20Shot%202020-07-16%20at%208.52.07%20PM.png?generation=1594958036239102&amp;alt=media\" alt=\"\"></p>\n\n<h1>Region 1 - 17.6% malignant</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7870c7996fb01709afe81007d505361a%2FScreen%20Shot%202020-07-16%20at%208.52.25%20PM.png?generation=1594957993975170&amp;alt=media\" alt=\"\"></p>\n\n<h1>Starter Notebook</h1>\n\n<p>If you want to play with image embeddings, I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>. We can use image embeddings to find the duplicate images in the train set and the duplicate images between this year's comp and last year's comp. We can also use image embeddings as input into another ML model like XGB, or kNN, or Logistic Regression. An advantage of using another ML model is that we can then easily input meta features also.</p>\n\n<p>Additionally you can use RAPIDS cuML KMeans and cuML kNN to visualize similar images. This is also demonstrated in the starter notebook.</p>\n\n<h1>Train Test Comparison</h1>\n\n<p>I did a train data test data comparison using RAPIDS cuML TNSE and posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">here</a>. The results are surprising. The test data contains <strong>mystery</strong> images!</p>",
  "messages": [
    {
      "id": 932425,
      "postDate": "2020-07-17T04:02:12.850Z",
      "content": "<h1>How CNN Works</h1>\n\n<p>A CNN is basically two parts. The bottom layers convert images into features and the top layers convert features into classification</p>\n\n<h1>Transfer Learning</h1>\n\n<p>When we download a pretrained CNN from the internet, we download the bottom layers for feature generation. Then we build our own top layers (custom head) and train them to convert features into classification for our specific problem (such as Melanoma Comp).</p>\n\n<h1>Feature Extraction</h1>\n\n<p>Instead of using a CNN for classification, we can just extract the features. Then for every image, we have a vector of numbers that represents the image. These features (also called image embeddings) have a <strong>magical property</strong>. When the distance between two image vectors is small then the images are similar. (Cosine similarity or Euclidean similarity). </p>\n\n<h1>RAPIDS cuML TSNE</h1>\n\n<p>If we extract features from EfficientNetB0, then we get vectors of length 1280. So each image is represented by a point in 1280 dimensional space. It's hard to visualize this but RAPIDS cuML TSNE can project this into 2 dimensions. Then we can plot each image as a dot in the <code>x-y-plane</code> below.</p>\n\n<h1>View Images</h1>\n\n<p>All dots that are close together are similar images. We can randomly select 10 regions and look at the images in those regions. We can also calculate the proportion of images in those regions that are malignant.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F6c9be2831918aae56af077e2bb604f57%2FScreen%20Shot%202020-07-16%20at%208.39.15%20PM.png?generation=1594957184583179&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 10 - 0.0% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F661d7ff283f66b9b035eda0cd93a904a%2FScreen%20Shot%202020-07-16%20at%208.45.30%20PM.png?generation=1594957547795576&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 9 - 1.5% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0a8694f40f085acb1de5fb6dfb99b6a0%2FScreen%20Shot%202020-07-16%20at%208.46.43%20PM.png?generation=1594957614965043&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 8 - 2.1% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F095cd7fa1d4ca3d2e6bb30652fdc8ab8%2FScreen%20Shot%202020-07-16%20at%208.47.34%20PM.png?generation=1594957665483816&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 7 - 4.0% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F523a1f26ec7a3f159b23c4d27b715262%2FScreen%20Shot%202020-07-16%20at%208.48.19%20PM.png?generation=1594957709694231&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 6 - 4.9% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffeae6e2c67abebc9cdb3fa1d8b9cd20a%2FScreen%20Shot%202020-07-16%20at%208.49.29%20PM.png?generation=1594957778977807&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 5 - 5.7% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F13cb7eaddef1599c089cc4150bbf3f52%2FScreen%20Shot%202020-07-16%20at%208.50.39%20PM.png?generation=1594957852045805&amp;alt=media\" alt=\"\"></p>\n\n<h2>Region 4 - 7.4% malignant</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcc393e67d21e132bc4eaa8d644d30b78%2FScreen%20Shot%202020-07-16%20at%208.51.23%20PM.png?generation=1594957893670098&amp;alt=media\" alt=\"\"></p>\n\n<h1>Region 3 - 9.4% malignant</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fa3b174517759d56220f215f748989182%2FScreen%20Shot%202020-07-16%20at%208.52.36%20PM.png?generation=1594957969518207&amp;alt=media\" alt=\"\"></p>\n\n<h1>Region 2 - 10.8% malignant</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5549e7e55ea222368a3c705328d1ece9%2FScreen%20Shot%202020-07-16%20at%208.52.07%20PM.png?generation=1594958036239102&amp;alt=media\" alt=\"\"></p>\n\n<h1>Region 1 - 17.6% malignant</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7870c7996fb01709afe81007d505361a%2FScreen%20Shot%202020-07-16%20at%208.52.25%20PM.png?generation=1594957993975170&amp;alt=media\" alt=\"\"></p>\n\n<h1>Starter Notebook</h1>\n\n<p>If you want to play with image embeddings, I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>. We can use image embeddings to find the duplicate images in the train set and the duplicate images between this year's comp and last year's comp. We can also use image embeddings as input into another ML model like XGB, or kNN, or Logistic Regression. An advantage of using another ML model is that we can then easily input meta features also.</p>\n\n<p>Additionally you can use RAPIDS cuML KMeans and cuML kNN to visualize similar images. This is also demonstrated in the starter notebook.</p>\n\n<h1>Train Test Comparison</h1>\n\n<p>I did a train data test data comparison using RAPIDS cuML TNSE and posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">here</a>. The results are surprising. The test data contains <strong>mystery</strong> images!</p>",
      "rawMarkdown": "# How CNN Works\nA CNN is basically two parts. The bottom layers convert images into features and the top layers convert features into classification\n\n# Transfer Learning\nWhen we download a pretrained CNN from the internet, we download the bottom layers for feature generation. Then we build our own top layers (custom head) and train them to convert features into classification for our specific problem (such as Melanoma Comp).\n\n# Feature Extraction\nInstead of using a CNN for classification, we can just extract the features. Then for every image, we have a vector of numbers that represents the image. These features (also called image embeddings) have a **magical property**. When the distance between two image vectors is small then the images are similar. (Cosine similarity or Euclidean similarity). \n\n# RAPIDS cuML TSNE\nIf we extract features from EfficientNetB0, then we get vectors of length 1280. So each image is represented by a point in 1280 dimensional space. It's hard to visualize this but RAPIDS cuML TSNE can project this into 2 dimensions. Then we can plot each image as a dot in the `x-y-plane` below.\n\n# View Images\nAll dots that are close together are similar images. We can randomly select 10 regions and look at the images in those regions. We can also calculate the proportion of images in those regions that are malignant.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F6c9be2831918aae56af077e2bb604f57%2FScreen%20Shot%202020-07-16%20at%208.39.15%20PM.png?generation=1594957184583179&amp;alt=media)\n\n## Region 10 - 0.0% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F661d7ff283f66b9b035eda0cd93a904a%2FScreen%20Shot%202020-07-16%20at%208.45.30%20PM.png?generation=1594957547795576&amp;alt=media)\n\n## Region 9 - 1.5% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0a8694f40f085acb1de5fb6dfb99b6a0%2FScreen%20Shot%202020-07-16%20at%208.46.43%20PM.png?generation=1594957614965043&amp;alt=media)\n\n## Region 8 - 2.1% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F095cd7fa1d4ca3d2e6bb30652fdc8ab8%2FScreen%20Shot%202020-07-16%20at%208.47.34%20PM.png?generation=1594957665483816&amp;alt=media)\n\n## Region 7 - 4.0% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F523a1f26ec7a3f159b23c4d27b715262%2FScreen%20Shot%202020-07-16%20at%208.48.19%20PM.png?generation=1594957709694231&amp;alt=media)\n\n## Region 6 - 4.9% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffeae6e2c67abebc9cdb3fa1d8b9cd20a%2FScreen%20Shot%202020-07-16%20at%208.49.29%20PM.png?generation=1594957778977807&amp;alt=media)\n\n## Region 5 - 5.7% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F13cb7eaddef1599c089cc4150bbf3f52%2FScreen%20Shot%202020-07-16%20at%208.50.39%20PM.png?generation=1594957852045805&amp;alt=media)\n\n## Region 4 - 7.4% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcc393e67d21e132bc4eaa8d644d30b78%2FScreen%20Shot%202020-07-16%20at%208.51.23%20PM.png?generation=1594957893670098&amp;alt=media)\n\n# Region 3 - 9.4% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fa3b174517759d56220f215f748989182%2FScreen%20Shot%202020-07-16%20at%208.52.36%20PM.png?generation=1594957969518207&amp;alt=media)\n\n# Region 2 - 10.8% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5549e7e55ea222368a3c705328d1ece9%2FScreen%20Shot%202020-07-16%20at%208.52.07%20PM.png?generation=1594958036239102&amp;alt=media)\n\n\n# Region 1 - 17.6% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7870c7996fb01709afe81007d505361a%2FScreen%20Shot%202020-07-16%20at%208.52.25%20PM.png?generation=1594957993975170&amp;alt=media)\n\n# Starter Notebook\nIf you want to play with image embeddings, I posted a starter notebook [here][1]. We can use image embeddings to find the duplicate images in the train set and the duplicate images between this year's comp and last year's comp. We can also use image embeddings as input into another ML model like XGB, or kNN, or Logistic Regression. An advantage of using another ML model is that we can then easily input meta features also.\n\nAdditionally you can use RAPIDS cuML KMeans and cuML kNN to visualize similar images. This is also demonstrated in the starter notebook.\n\n# Train Test Comparison\nI did a train data test data comparison using RAPIDS cuML TNSE and posted [here][2]. The results are surprising. The test data contains **mystery** images!\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n",
      "votes": 134
    },
    {
      "id": 932485,
      "postDate": "2020-07-17T05:00:49.060Z",
      "content": "<p>The following plots are quite enlightening! They show many secrets!\n* The 2020 test data mostly matches the 2020 train data\n* There is a section of 2020 test data that is not in the 2020 train data\n* The 2018+2017 comp data is similar to a portion of the 2020 train and test data\n* The new portion of 2019 comp data is different than 2020, 2018, 2017 data</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0ab11f71c0a48dd722dbf4fa3644718e%2FScreen%20Shot%202020-07-16%20at%2010.15.23%20PM.png?generation=1594962953082990&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F40c982354929bbe5bb2fdfc337ee95d7%2FScreen%20Shot%202020-07-16%20at%2010.15.34%20PM.png?generation=1594962964835771&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The following plots are quite enlightening! They show many secrets!\n* The 2020 test data mostly matches the 2020 train data\n* There is a section of 2020 test data that is not in the 2020 train data\n* The 2018+2017 comp data is similar to a portion of the 2020 train and test data\n* The new portion of 2019 comp data is different than 2020, 2018, 2017 data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0ab11f71c0a48dd722dbf4fa3644718e%2FScreen%20Shot%202020-07-16%20at%2010.15.23%20PM.png?generation=1594962953082990&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F40c982354929bbe5bb2fdfc337ee95d7%2FScreen%20Shot%202020-07-16%20at%2010.15.34%20PM.png?generation=1594962964835771&amp;alt=media)\n",
      "votes": 24,
      "replies": [
        {
          "id": 932578,
          "postDate": "2020-07-17T06:44:34.340Z",
          "content": "<p>Very nice work. I guess you could use this to 'hand' select (Using some heuristic threshold on the cosine similarity of the images in the embedded space) images from the 2019 dataset which best represent the 2020 dataset.</p>\n\n<p>The 2019 image embeddings could be different due to scale / color, I guess you could play with transformations and look at how the embedding changes.</p>",
          "rawMarkdown": "Very nice work. I guess you could use this to 'hand' select (Using some heuristic threshold on the cosine similarity of the images in the embedded space) images from the 2019 dataset which best represent the 2020 dataset.\n\nThe 2019 image embeddings could be different due to scale / color, I guess you could play with transformations and look at how the embedding changes.",
          "votes": 4
        },
        {
          "id": 933041,
          "postDate": "2020-07-17T12:51:38.383Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 933101,
          "postDate": "2020-07-17T13:41:20.963Z",
          "content": "<p>Did you use t-SNE for your plots? or another method?</p>",
          "rawMarkdown": "Did you use t-SNE for your plots? or another method?",
          "votes": 1
        },
        {
          "id": 933202,
          "postDate": "2020-07-17T15:02:10.480Z",
          "content": "<p>Yes these plots are RAPIDS cuML t-SNE. I use the notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a> and modify the code to:</p>\n\n<pre><code>model = cuml.TSNE()\nembed2D = model.fit_transform(np.concatenate([embed,embed_test,embed_ext]))\n# TRAIN\ntrain['x'] = embed2D[:len(embed),0]\ntrain['y'] = embed2D[:len(embed),1]\n# TEST\ntest['x'] = embed2D[len(embed):len(embed)+len(embed_test),0]\ntest['y'] = embed2D[len(embed):len(embed)+len(embed_test),1]\n# EXTERNAL\ntrain_ext['x'] = embed2D[len(embed)+len(embed_test):,0]\ntrain_ext['y'] = embed2D[len(embed)+len(embed_test):,1]\n</code></pre>\n\n<p>And then you plot with something like</p>\n\n<pre><code>plt.figure(figsize=(20,10))\n\nplt.subplot(1,2,1)\ndf1 = train.loc[train.target==0]\nplt.scatter(df1.x,df1.y,color='orange',s=10,label='2020 Train Benign')\ndf2 = train.loc[train.target==1]\nplt.scatter(df2.x,df2.y,color='blue',s=10,label='2020 Train Malignant')\nplt.legend()\n\nplt.subplot(1,2,2)\ndf1 = train.loc[train.target==0]\nplt.scatter(df1.x,df1.y,color='orange',s=10,label='2020 Train Benign')\ndf2 = train.loc[train.target==1]\nplt.scatter(df2.x,df2.y,color='blue',s=10,label='2020 Train Malignant')\ndf3 = test\nplt.scatter(df3.x,df3.y,color='green',s=10,label='2020 Test')\nplt.legend()\n\nplt.show()\n</code></pre>",
          "rawMarkdown": "Yes these plots are RAPIDS cuML t-SNE. I use the notebook [here][1] and modify the code to:\n\n    model = cuml.TSNE()\n    embed2D = model.fit_transform(np.concatenate([embed,embed_test,embed_ext]))\n    # TRAIN\n    train['x'] = embed2D[:len(embed),0]\n    train['y'] = embed2D[:len(embed),1]\n    # TEST\n    test['x'] = embed2D[len(embed):len(embed)+len(embed_test),0]\n    test['y'] = embed2D[len(embed):len(embed)+len(embed_test),1]\n    # EXTERNAL\n    train_ext['x'] = embed2D[len(embed)+len(embed_test):,0]\n    train_ext['y'] = embed2D[len(embed)+len(embed_test):,1]\n\nAnd then you plot with something like\n\n    plt.figure(figsize=(20,10))\n\n    plt.subplot(1,2,1)\n    df1 = train.loc[train.target==0]\n    plt.scatter(df1.x,df1.y,color='orange',s=10,label='2020 Train Benign')\n    df2 = train.loc[train.target==1]\n    plt.scatter(df2.x,df2.y,color='blue',s=10,label='2020 Train Malignant')\n    plt.legend()\n\n    plt.subplot(1,2,2)\n    df1 = train.loc[train.target==0]\n    plt.scatter(df1.x,df1.y,color='orange',s=10,label='2020 Train Benign')\n    df2 = train.loc[train.target==1]\n    plt.scatter(df2.x,df2.y,color='blue',s=10,label='2020 Train Malignant')\n    df3 = test\n    plt.scatter(df3.x,df3.y,color='green',s=10,label='2020 Test')\n    plt.legend()\n\n    plt.show()\n    \n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates",
          "votes": 6
        }
      ]
    },
    {
      "id": 933418,
      "postDate": "2020-07-17T17:32:46.277Z",
      "content": "<p>RAPIDS has UMAP too. If you wish to explore UMAP just change <code>model = cuml.TSNE()</code> to <code>model = cuml.UMAP()</code> in the notebook. Both take only seconds to execute.</p>",
      "rawMarkdown": "RAPIDS has UMAP too. If you wish to explore UMAP just change `model = cuml.TSNE()` to `model = cuml.UMAP()` in the notebook. Both take only seconds to execute.",
      "votes": 5
    },
    {
      "id": 977072,
      "postDate": "2020-08-19T09:34:48.560Z",
      "content": "<p>Wow its so cool, feel myself like a magician, thanks<br>\n(EfficientNet-B0 number of features: 20480)<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2919655%2F912d013c51835f60ad6d414e6a8ea346%2Ffull_tsne.png?generation=1597829550882459&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Wow its so cool, feel myself like a magician, thanks\n(EfficientNet-B0 number of features: 20480)\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2919655%2F912d013c51835f60ad6d414e6a8ea346%2Ffull_tsne.png?generation=1597829550882459&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 977556,
          "postDate": "2020-08-19T14:59:37.950Z",
          "content": "<p>Awesome plot</p>",
          "rawMarkdown": "Awesome plot",
          "votes": 1
        }
      ]
    },
    {
      "id": 933330,
      "postDate": "2020-07-17T16:50:41.987Z",
      "content": "<p>Your second point is interesting: </p>\n\n<p>One could implement an Ugly Duckling method using UMAP. Unlike t-SNE UMAP provides a mapping to the embedded space so any instance can be projected into the lower dimensional space - it can actually be used in a classification pipeline. One could transform all of a patients images into the lower dimensional space at test time and identify the ugly ducklings in this way.</p>",
      "rawMarkdown": "Your second point is interesting: \n\nOne could implement an Ugly Duckling method using UMAP. Unlike t-SNE UMAP provides a mapping to the embedded space so any instance can be projected into the lower dimensional space - it can actually be used in a classification pipeline. One could transform all of a patients images into the lower dimensional space at test time and identify the ugly ducklings in this way.",
      "votes": 3,
      "replies": [
        {
          "id": 933527,
          "postDate": "2020-07-17T18:47:07.987Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 933305,
      "postDate": "2020-07-17T16:28:51.930Z",
      "content": "<p>With only a few tweaks, this notebook can easily be used to evaluate predictions of our models. I took the OOF outcome of my best-scoring model (CV .945), and visualized the most confident false negatives (200) and false positives (500). Colors are blending a bit, but few things can be observed nevertheless.</p>\n<ul>\n<li>The most resilient cases to predict are indeed tricky. False positives very closely copy true positives and on the contrary, false negatives frequently deviate.</li>\n<li>For this reason, fine-tuning models and ensembling can only achieve as much, and increasing AUC beyond .95-96 will most probably need more sophisticated work with patient-level info, such as the Ugly Duckling method. Looking forward to implementing it!</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3919031%2Fb00b31ca741a79c1bd3100a558b2153d%2FOOF_plot.png\" alt=\"\"></p>\n<p>PS: If you're not seeing the image, see the attachment. I'm having troubles displaying it :-(</p>",
      "rawMarkdown": "With only a few tweaks, this notebook can easily be used to evaluate predictions of our models. I took the OOF outcome of my best-scoring model (CV .945), and visualized the most confident false negatives (200) and false positives (500). Colors are blending a bit, but few things can be observed nevertheless.\n- The most resilient cases to predict are indeed tricky. False positives very closely copy true positives and on the contrary, false negatives frequently deviate.\n- For this reason, fine-tuning models and ensembling can only achieve as much, and increasing AUC beyond .95-96 will most probably need more sophisticated work with patient-level info, such as the Ugly Duckling method. Looking forward to implementing it!\n\n![](https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3919031%2Fb00b31ca741a79c1bd3100a558b2153d%2FOOF_plot.png)\n\nPS: If you're not seeing the image, see the attachment. I'm having troubles displaying it :-(",
      "votes": 3
    },
    {
      "id": 964738,
      "postDate": "2020-08-10T06:13:49.777Z",
      "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> for the awsome detail explanation 👍 !</p>",
      "rawMarkdown": "Thank you @cdeotte for the awsome detail explanation 👍 !",
      "votes": 1
    },
    {
      "id": 936937,
      "postDate": "2020-07-20T15:59:58.700Z",
      "content": "<p>Great !!</p>",
      "rawMarkdown": "Great !!",
      "votes": 1
    },
    {
      "id": 935860,
      "postDate": "2020-07-19T17:59:24.983Z",
      "content": "<p>Wow very nice and clear!</p>",
      "rawMarkdown": "Wow very nice and clear!",
      "votes": 1
    },
    {
      "id": 934390,
      "postDate": "2020-07-18T12:06:07.987Z",
      "content": "<p>Nice idea ! Thanks for that! really insightful</p>",
      "rawMarkdown": "Nice idea ! Thanks for that! really insightful",
      "votes": 1,
      "replies": [
        {
          "id": 934938,
          "postDate": "2020-07-18T23:47:05.440Z",
          "content": "<p>Thanks Manh. I also used RAPIDS cuML TNSE to compare our test dataset with train dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">here</a>. And I compared the new portion of the 2019 comp data with our train and test too. The results are enlightening.</p>",
          "rawMarkdown": "Thanks Manh. I also used RAPIDS cuML TNSE to compare our test dataset with train dataset [here][1]. And I compared the new portion of the 2019 comp data with our train and test too. The results are enlightening.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028"
        }
      ]
    },
    {
      "id": 934014,
      "postDate": "2020-07-18T07:34:57.257Z",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!",
      "votes": 1
    },
    {
      "id": 933858,
      "postDate": "2020-07-18T04:52:52.393Z",
      "content": "<p>Thanks for the very clear and detailed analysis!</p>",
      "rawMarkdown": "Thanks for the very clear and detailed analysis!",
      "votes": 1
    },
    {
      "id": 933835,
      "postDate": "2020-07-18T04:26:12.057Z",
      "content": "<p>Really insightful kernel! (upvoted)</p>",
      "rawMarkdown": "Really insightful kernel! (upvoted)",
      "votes": 1,
      "replies": [
        {
          "id": 933837,
          "postDate": "2020-07-18T04:30:21.997Z",
          "content": "<p>Thanks Souhardya</p>",
          "rawMarkdown": "Thanks Souhardya",
          "votes": 1
        }
      ]
    },
    {
      "id": 933639,
      "postDate": "2020-07-17T21:03:08.553Z",
      "content": "<p>WELL EXPLAINED!!!</p>",
      "rawMarkdown": "WELL EXPLAINED!!!",
      "votes": 1,
      "replies": [
        {
          "id": 933640,
          "postDate": "2020-07-17T21:05:15.663Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 933118,
      "postDate": "2020-07-17T14:00:12.303Z",
      "content": "<p>I'm trying to put this comment in the right thread. <a href=\"/cdeotte\">@cdeotte</a>  please attend as many kaggle competitions as you can, I've learnt the most from your sharing in this competition</p>",
      "rawMarkdown": "I'm trying to put this comment in the right thread. @cdeotte  please attend as many kaggle competitions as you can, I've learnt the most from your sharing in this competition",
      "votes": 1
    },
    {
      "id": 932967,
      "postDate": "2020-07-17T11:59:51.977Z",
      "content": "<p>Wow <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you keep surprising me. Though I had a similar idea lately, I had little to no idea how to execute it. I sincerely thank you for all your contributions, this is my first comp and I keep learning so much from you.</p>",
      "rawMarkdown": "Wow @cdeotte, you keep surprising me. Though I had a similar idea lately, I had little to no idea how to execute it. I sincerely thank you for all your contributions, this is my first comp and I keep learning so much from you.",
      "votes": 1
    },
    {
      "id": 940383,
      "postDate": "2020-07-22T22:39:35.777Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 935411,
      "postDate": "2020-07-19T11:14:42.817Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 933516,
      "postDate": "2020-07-17T18:35:14.783Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 933514,
      "postDate": "2020-07-17T18:34:46.207Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 933425,
      "postDate": "2020-07-17T17:35:49.180Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 933421,
      "postDate": "2020-07-17T17:33:20.353Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 960569,
      "postDate": "2020-08-06T14:09:08.430Z",
      "content": "<p>This was really interesting. Thanks.</p>",
      "rawMarkdown": "This was really interesting. Thanks.",
      "votes": 1
    },
    {
      "id": 955206,
      "postDate": "2020-08-02T12:52:44.730Z",
      "content": "<p>very informative thanks sir.</p>",
      "rawMarkdown": "very informative thanks sir.",
      "votes": 1
    },
    {
      "id": 953841,
      "postDate": "2020-08-01T07:06:22.397Z",
      "content": "<p>Really useful. Thanks for sharing!</p>",
      "rawMarkdown": "Really useful. Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 944299,
      "postDate": "2020-07-25T02:11:16.703Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 937735,
      "postDate": "2020-07-21T06:26:56.193Z",
      "content": "<p>Thanks!\nAwesome!</p>",
      "rawMarkdown": "Thanks!\nAwesome!",
      "votes": 1
    },
    {
      "id": 937500,
      "postDate": "2020-07-21T03:58:03.030Z",
      "content": "<p>Thanks! Awesome work!</p>",
      "rawMarkdown": "Thanks! Awesome work!",
      "votes": 1
    },
    {
      "id": 937356,
      "postDate": "2020-07-21T00:47:46Z",
      "content": "<p>Thanks!\nreally Awesome work</p>",
      "rawMarkdown": "Thanks!\nreally Awesome work",
      "votes": 1
    },
    {
      "id": 937032,
      "postDate": "2020-07-20T17:17:16.680Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 936196,
      "postDate": "2020-07-20T04:09:56.977Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 936180,
      "postDate": "2020-07-20T03:42:46.440Z",
      "content": "<p>Nice Idea Thanks for Sharing</p>",
      "rawMarkdown": "Nice Idea Thanks for Sharing",
      "votes": 1
    },
    {
      "id": 935131,
      "postDate": "2020-07-19T05:18:59.430Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 934943,
      "postDate": "2020-07-19T00:01:42.347Z",
      "content": "<p>Thanks for sharing a wonderful idea! </p>",
      "rawMarkdown": "Thanks for sharing a wonderful idea! ",
      "votes": 1
    },
    {
      "id": 933015,
      "postDate": "2020-07-17T12:27:53.610Z",
      "content": "<p>Thanks Mr Chris for the awesome help.</p>",
      "rawMarkdown": "Thanks Mr Chris for the awesome help.",
      "votes": 1
    },
    {
      "id": 932440,
      "postDate": "2020-07-17T04:28:18.393Z",
      "content": "<p>Thanks for sharing:)</p>",
      "rawMarkdown": "Thanks for sharing:)",
      "votes": 1
    },
    {
      "id": 936610,
      "postDate": "2020-07-20T11:01:20.023Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!\n"
    }
  ],
  "comments": [
    {
      "id": 932485,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-17T05:00:49.060000",
      "content": "<p>The following plots are quite enlightening! They show many secrets!\n* The 2020 test data mostly matches the 2020 train data\n* There is a section of 2020 test data that is not in the 2020 train data\n* The 2018+2017 comp data is similar to a portion of the 2020 train and test data\n* The new portion of 2019 comp data is different than 2020, 2018, 2017 data</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0ab11f71c0a48dd722dbf4fa3644718e%2FScreen%20Shot%202020-07-16%20at%2010.15.23%20PM.png?generation=1594962953082990&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F40c982354929bbe5bb2fdfc337ee95d7%2FScreen%20Shot%202020-07-16%20at%2010.15.34%20PM.png?generation=1594962964835771&amp;alt=media\" alt=\"\"></p>",
      "votes": 24,
      "replies": [
        {
          "id": 932578,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-07-17T06:44:34.340000",
          "content": "<p>Very nice work. I guess you could use this to 'hand' select (Using some heuristic threshold on the cosine similarity of the images in the embedded space) images from the 2019 dataset which best represent the 2020 dataset.</p>\n\n<p>The 2019 image embeddings could be different due to scale / color, I guess you could play with transformations and look at how the embedding changes.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 933041,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-17T12:51:38.383000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 933101,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-07-17T13:41:20.963000",
          "content": "<p>Did you use t-SNE for your plots? or another method?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 933202,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-17T15:02:10.480000",
          "content": "<p>Yes these plots are RAPIDS cuML t-SNE. I use the notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a> and modify the code to:</p>\n\n<pre><code>model = cuml.TSNE()\nembed2D = model.fit_transform(np.concatenate([embed,embed_test,embed_ext]))\n# TRAIN\ntrain['x'] = embed2D[:len(embed),0]\ntrain['y'] = embed2D[:len(embed),1]\n# TEST\ntest['x'] = embed2D[len(embed):len(embed)+len(embed_test),0]\ntest['y'] = embed2D[len(embed):len(embed)+len(embed_test),1]\n# EXTERNAL\ntrain_ext['x'] = embed2D[len(embed)+len(embed_test):,0]\ntrain_ext['y'] = embed2D[len(embed)+len(embed_test):,1]\n</code></pre>\n\n<p>And then you plot with something like</p>\n\n<pre><code>plt.figure(figsize=(20,10))\n\nplt.subplot(1,2,1)\ndf1 = train.loc[train.target==0]\nplt.scatter(df1.x,df1.y,color='orange',s=10,label='2020 Train Benign')\ndf2 = train.loc[train.target==1]\nplt.scatter(df2.x,df2.y,color='blue',s=10,label='2020 Train Malignant')\nplt.legend()\n\nplt.subplot(1,2,2)\ndf1 = train.loc[train.target==0]\nplt.scatter(df1.x,df1.y,color='orange',s=10,label='2020 Train Benign')\ndf2 = train.loc[train.target==1]\nplt.scatter(df2.x,df2.y,color='blue',s=10,label='2020 Train Malignant')\ndf3 = test\nplt.scatter(df3.x,df3.y,color='green',s=10,label='2020 Test')\nplt.legend()\n\nplt.show()\n</code></pre>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 933418,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-17T17:32:46.277000",
      "content": "<p>RAPIDS has UMAP too. If you wish to explore UMAP just change <code>model = cuml.TSNE()</code> to <code>model = cuml.UMAP()</code> in the notebook. Both take only seconds to execute.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 977072,
      "author_name": "Ivan Void",
      "author_url": "",
      "post_date": "2020-08-19T09:34:48.560000",
      "content": "<p>Wow its so cool, feel myself like a magician, thanks<br>\n(EfficientNet-B0 number of features: 20480)<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2919655%2F912d013c51835f60ad6d414e6a8ea346%2Ffull_tsne.png?generation=1597829550882459&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 977556,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T14:59:37.950000",
          "content": "<p>Awesome plot</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 933330,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2020-07-17T16:50:41.987000",
      "content": "<p>Your second point is interesting: </p>\n\n<p>One could implement an Ugly Duckling method using UMAP. Unlike t-SNE UMAP provides a mapping to the embedded space so any instance can be projected into the lower dimensional space - it can actually be used in a classification pipeline. One could transform all of a patients images into the lower dimensional space at test time and identify the ugly ducklings in this way.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 933527,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-17T18:47:07.987000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 933305,
      "author_name": "Andrej Rybár",
      "author_url": "",
      "post_date": "2020-07-17T16:28:51.930000",
      "content": "<p>With only a few tweaks, this notebook can easily be used to evaluate predictions of our models. I took the OOF outcome of my best-scoring model (CV .945), and visualized the most confident false negatives (200) and false positives (500). Colors are blending a bit, but few things can be observed nevertheless.</p>\n<ul>\n<li>The most resilient cases to predict are indeed tricky. False positives very closely copy true positives and on the contrary, false negatives frequently deviate.</li>\n<li>For this reason, fine-tuning models and ensembling can only achieve as much, and increasing AUC beyond .95-96 will most probably need more sophisticated work with patient-level info, such as the Ugly Duckling method. Looking forward to implementing it!</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3919031%2Fb00b31ca741a79c1bd3100a558b2153d%2FOOF_plot.png\" alt=\"\"></p>\n<p>PS: If you're not seeing the image, see the attachment. I'm having troubles displaying it :-(</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 964738,
      "author_name": "Beans",
      "author_url": "",
      "post_date": "2020-08-10T06:13:49.777000",
      "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> for the awsome detail explanation 👍 !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936937,
      "author_name": "yatok36789",
      "author_url": "",
      "post_date": "2020-07-20T15:59:58.700000",
      "content": "<p>Great !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 935860,
      "author_name": "Karthikeyanmnec",
      "author_url": "",
      "post_date": "2020-07-19T17:59:24.983000",
      "content": "<p>Wow very nice and clear!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934390,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-07-18T12:06:07.987000",
      "content": "<p>Nice idea ! Thanks for that! really insightful</p>",
      "votes": 1,
      "replies": [
        {
          "id": 934938,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-18T23:47:05.440000",
          "content": "<p>Thanks Manh. I also used RAPIDS cuML TNSE to compare our test dataset with train dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">here</a>. And I compared the new portion of the 2019 comp data with our train and test too. The results are enlightening.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 934014,
      "author_name": "Souradeep Bera",
      "author_url": "",
      "post_date": "2020-07-18T07:34:57.257000",
      "content": "<p>Nice!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 933858,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-07-18T04:52:52.393000",
      "content": "<p>Thanks for the very clear and detailed analysis!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 933835,
      "author_name": "Souhardya Ganguly",
      "author_url": "",
      "post_date": "2020-07-18T04:26:12.057000",
      "content": "<p>Really insightful kernel! (upvoted)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 933837,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-18T04:30:21.997000",
          "content": "<p>Thanks Souhardya</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 933639,
      "author_name": "SAI AKHIL K",
      "author_url": "",
      "post_date": "2020-07-17T21:03:08.553000",
      "content": "<p>WELL EXPLAINED!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 933640,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-17T21:05:15.663000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 933118,
      "author_name": "ZHU CHAO",
      "author_url": "",
      "post_date": "2020-07-17T14:00:12.303000",
      "content": "<p>I'm trying to put this comment in the right thread. <a href=\"/cdeotte\">@cdeotte</a>  please attend as many kaggle competitions as you can, I've learnt the most from your sharing in this competition</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 932967,
      "author_name": "Andrej Rybár",
      "author_url": "",
      "post_date": "2020-07-17T11:59:51.977000",
      "content": "<p>Wow <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you keep surprising me. Though I had a similar idea lately, I had little to no idea how to execute it. I sincerely thank you for all your contributions, this is my first comp and I keep learning so much from you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940383,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-22T22:39:35.777000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 935411,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-19T11:14:42.817000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 933516,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T18:35:14.783000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 933514,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T18:34:46.207000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 933425,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T17:35:49.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 933421,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T17:33:20.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 960569,
      "author_name": "Anders Ericsson Gnosco",
      "author_url": "",
      "post_date": "2020-08-06T14:09:08.430000",
      "content": "<p>This was really interesting. Thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 955206,
      "author_name": "Sachin kumawat",
      "author_url": "",
      "post_date": "2020-08-02T12:52:44.730000",
      "content": "<p>very informative thanks sir.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 953841,
      "author_name": "Jaseem C K",
      "author_url": "",
      "post_date": "2020-08-01T07:06:22.397000",
      "content": "<p>Really useful. Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944299,
      "author_name": "Sanjay K",
      "author_url": "",
      "post_date": "2020-07-25T02:11:16.703000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937735,
      "author_name": "Kensei Oya",
      "author_url": "",
      "post_date": "2020-07-21T06:26:56.193000",
      "content": "<p>Thanks!\nAwesome!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937500,
      "author_name": "AegonDenver",
      "author_url": "",
      "post_date": "2020-07-21T03:58:03.030000",
      "content": "<p>Thanks! Awesome work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937356,
      "author_name": "Luke CHO",
      "author_url": "",
      "post_date": "2020-07-21T00:47:46",
      "content": "<p>Thanks!\nreally Awesome work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937032,
      "author_name": "Emirhan Ergin",
      "author_url": "",
      "post_date": "2020-07-20T17:17:16.680000",
      "content": "<p>Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936196,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-20T04:09:56.977000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936180,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-20T03:42:46.440000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 935131,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-19T05:18:59.430000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934943,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-19T00:01:42.347000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 933015,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T12:27:53.610000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 932440,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T04:28:18.393000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936610,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-20T11:01:20.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "932425": "# How CNN Works\nA CNN is basically two parts. The bottom layers convert images into features and the top layers convert features into classification\n\n# Transfer Learning\nWhen we download a pretrained CNN from the internet, we download the bottom layers for feature generation. Then we build our own top layers (custom head) and train them to convert features into classification for our specific problem (such as Melanoma Comp).\n\n# Feature Extraction\nInstead of using a CNN for classification, we can just extract the features. Then for every image, we have a vector of numbers that represents the image. These features (also called image embeddings) have a **magical property**. When the distance between two image vectors is small then the images are similar. (Cosine similarity or Euclidean similarity). \n\n# RAPIDS cuML TSNE\nIf we extract features from EfficientNetB0, then we get vectors of length 1280. So each image is represented by a point in 1280 dimensional space. It's hard to visualize this but RAPIDS cuML TSNE can project this into 2 dimensions. Then we can plot each image as a dot in the `x-y-plane` below.\n\n# View Images\nAll dots that are close together are similar images. We can randomly select 10 regions and look at the images in those regions. We can also calculate the proportion of images in those regions that are malignant.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F6c9be2831918aae56af077e2bb604f57%2FScreen%20Shot%202020-07-16%20at%208.39.15%20PM.png?generation=1594957184583179&amp;alt=media)\n\n## Region 10 - 0.0% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F661d7ff283f66b9b035eda0cd93a904a%2FScreen%20Shot%202020-07-16%20at%208.45.30%20PM.png?generation=1594957547795576&amp;alt=media)\n\n## Region 9 - 1.5% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0a8694f40f085acb1de5fb6dfb99b6a0%2FScreen%20Shot%202020-07-16%20at%208.46.43%20PM.png?generation=1594957614965043&amp;alt=media)\n\n## Region 8 - 2.1% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F095cd7fa1d4ca3d2e6bb30652fdc8ab8%2FScreen%20Shot%202020-07-16%20at%208.47.34%20PM.png?generation=1594957665483816&amp;alt=media)\n\n## Region 7 - 4.0% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F523a1f26ec7a3f159b23c4d27b715262%2FScreen%20Shot%202020-07-16%20at%208.48.19%20PM.png?generation=1594957709694231&amp;alt=media)\n\n## Region 6 - 4.9% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ffeae6e2c67abebc9cdb3fa1d8b9cd20a%2FScreen%20Shot%202020-07-16%20at%208.49.29%20PM.png?generation=1594957778977807&amp;alt=media)\n\n## Region 5 - 5.7% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F13cb7eaddef1599c089cc4150bbf3f52%2FScreen%20Shot%202020-07-16%20at%208.50.39%20PM.png?generation=1594957852045805&amp;alt=media)\n\n## Region 4 - 7.4% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcc393e67d21e132bc4eaa8d644d30b78%2FScreen%20Shot%202020-07-16%20at%208.51.23%20PM.png?generation=1594957893670098&amp;alt=media)\n\n# Region 3 - 9.4% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fa3b174517759d56220f215f748989182%2FScreen%20Shot%202020-07-16%20at%208.52.36%20PM.png?generation=1594957969518207&amp;alt=media)\n\n# Region 2 - 10.8% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5549e7e55ea222368a3c705328d1ece9%2FScreen%20Shot%202020-07-16%20at%208.52.07%20PM.png?generation=1594958036239102&amp;alt=media)\n\n\n# Region 1 - 17.6% malignant\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7870c7996fb01709afe81007d505361a%2FScreen%20Shot%202020-07-16%20at%208.52.25%20PM.png?generation=1594957993975170&amp;alt=media)\n\n# Starter Notebook\nIf you want to play with image embeddings, I posted a starter notebook [here][1]. We can use image embeddings to find the duplicate images in the train set and the duplicate images between this year's comp and last year's comp. We can also use image embeddings as input into another ML model like XGB, or kNN, or Logistic Regression. An advantage of using another ML model is that we can then easily input meta features also.\n\nAdditionally you can use RAPIDS cuML KMeans and cuML kNN to visualize similar images. This is also demonstrated in the starter notebook.\n\n# Train Test Comparison\nI did a train data test data comparison using RAPIDS cuML TNSE and posted [here][2]. The results are surprising. The test data contains **mystery** images!\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n",
    "932485": "The following plots are quite enlightening! They show many secrets!\n* The 2020 test data mostly matches the 2020 train data\n* There is a section of 2020 test data that is not in the 2020 train data\n* The 2018+2017 comp data is similar to a portion of the 2020 train and test data\n* The new portion of 2019 comp data is different than 2020, 2018, 2017 data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0ab11f71c0a48dd722dbf4fa3644718e%2FScreen%20Shot%202020-07-16%20at%2010.15.23%20PM.png?generation=1594962953082990&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F40c982354929bbe5bb2fdfc337ee95d7%2FScreen%20Shot%202020-07-16%20at%2010.15.34%20PM.png?generation=1594962964835771&amp;alt=media)\n",
    "933418": "RAPIDS has UMAP too. If you wish to explore UMAP just change `model = cuml.TSNE()` to `model = cuml.UMAP()` in the notebook. Both take only seconds to execute.",
    "977072": "Wow its so cool, feel myself like a magician, thanks\n(EfficientNet-B0 number of features: 20480)\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2919655%2F912d013c51835f60ad6d414e6a8ea346%2Ffull_tsne.png?generation=1597829550882459&alt=media)",
    "933330": "Your second point is interesting: \n\nOne could implement an Ugly Duckling method using UMAP. Unlike t-SNE UMAP provides a mapping to the embedded space so any instance can be projected into the lower dimensional space - it can actually be used in a classification pipeline. One could transform all of a patients images into the lower dimensional space at test time and identify the ugly ducklings in this way.",
    "933305": "With only a few tweaks, this notebook can easily be used to evaluate predictions of our models. I took the OOF outcome of my best-scoring model (CV .945), and visualized the most confident false negatives (200) and false positives (500). Colors are blending a bit, but few things can be observed nevertheless.\n- The most resilient cases to predict are indeed tricky. False positives very closely copy true positives and on the contrary, false negatives frequently deviate.\n- For this reason, fine-tuning models and ensembling can only achieve as much, and increasing AUC beyond .95-96 will most probably need more sophisticated work with patient-level info, such as the Ugly Duckling method. Looking forward to implementing it!\n\n![](https://www.googleapis.com/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3919031%2Fb00b31ca741a79c1bd3100a558b2153d%2FOOF_plot.png)\n\nPS: If you're not seeing the image, see the attachment. I'm having troubles displaying it :-(",
    "964738": "Thank you @cdeotte for the awsome detail explanation 👍 !",
    "936937": "Great !!",
    "935860": "Wow very nice and clear!",
    "934390": "Nice idea ! Thanks for that! really insightful",
    "934014": "Nice!",
    "933858": "Thanks for the very clear and detailed analysis!",
    "933835": "Really insightful kernel! (upvoted)",
    "933639": "WELL EXPLAINED!!!",
    "933118": "I'm trying to put this comment in the right thread. @cdeotte  please attend as many kaggle competitions as you can, I've learnt the most from your sharing in this competition",
    "932967": "Wow @cdeotte, you keep surprising me. Though I had a similar idea lately, I had little to no idea how to execute it. I sincerely thank you for all your contributions, this is my first comp and I keep learning so much from you.",
    "940383": "",
    "935411": "",
    "933516": "",
    "933514": "",
    "933425": "",
    "933421": "",
    "960569": "This was really interesting. Thanks.",
    "955206": "very informative thanks sir.",
    "953841": "Really useful. Thanks for sharing!",
    "944299": "Thanks for sharing!",
    "937735": "Thanks!\nAwesome!",
    "937500": "Thanks! Awesome work!",
    "937356": "Thanks!\nreally Awesome work",
    "937032": "Thanks!",
    "936196": "Thanks!",
    "936180": "Nice Idea Thanks for Sharing",
    "935131": "Thanks!",
    "934943": "Thanks for sharing a wonderful idea! ",
    "933015": "Thanks Mr Chris for the awesome help.",
    "932440": "Thanks for sharing:)",
    "936610": "Thanks!\n"
  }
}