{
  "id": 168456,
  "title": "Ugly duckling detection - UMAP and pairwise distance",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168456",
  "author_name": "",
  "post_date": "2020-07-20T17:38:48.480484400Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>@cdeotte  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">has shown us</a> how one can use t-SNE and UMAP to project CNN image embeddings into a 2D space. </p>\n\n<p>In the case of UMAP (unlike t-SNE) your model learns a direct transformation of the CNN embedding to the 2D space which means you can use it in a classifier. I've been investigating this and have shown that these 2D embeddings could display characteristics of the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\">ugly-duckling concept</a>.</p>\n\n<p>To do so simply follow Chris' excellent notebook (using UMAP) train and save a UMAP transformer and you can then project all of a patients images (either in the train or test) into the 2D space. For example consider these  images which shows four (hand selected) patients' images projected into the 2D space, where yellow points indicate images of malignant lesions.</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F5c9396bb0ba3b968f83b2d3110bb99c9%2Fpatient_embedding.png?generation=1595265660164754&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F5c9396bb0ba3b968f83b2d3110bb99c9%2Fpatient_embedding.png?generation=1595265660164754&amp;alt=media</a> =500x500)</p>\n\n<p>You can see that some malignant images are far away from what a 'normal' image is for a patient, although this is not always true. </p>\n\n<p>For every image (in training or test) one can then calculate the distance in this 2D space between its N nearest neighbours (for a given patient) and this distance shows some correlation to malignancy. For example, I've calculated it for N=3 and the below image shows the fraction of lesions which are malignant as a function of their mean distance to its three nearest neighbours - those further from there nearest-neighbours are more likely to be malignant.</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2Fd931eec054069c1fcef982b4377e8cb9%2Ffraction_malignant.png?generation=1595266011305628&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2Fd931eec054069c1fcef982b4377e8cb9%2Ffraction_malignant.png?generation=1595266011305628&amp;alt=media</a> =500x500)</p>\n\n<p>Of course, at some level this is to be expected: malignant lesions should not look like non-malignant lesions and should be far away in this 2D space. But this method could add additional information as it makes use of patient level information to characterise what is normal for a given patient.</p>\n\n<p>Using the mean distance of the image to its three nearest neighbours as a simple predictor yielded an OOF AUC of 0.7 but using it in my ensembles did not improve CV. However, the initial CNN (EFN B3) embedding is not trained on lesion images (image_weights='noisy-student'), so the method could improve if a better initial CNN embedding is used...</p>",
  "messages": [
    {
      "id": "937055",
      "postDate": "07/20/2020 17:38:48",
      "content": "<p>@cdeotte  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">has shown us</a> how one can use t-SNE and UMAP to project CNN image embeddings into a 2D space. </p>\n\n<p>In the case of UMAP (unlike t-SNE) your model learns a direct transformation of the CNN embedding to the 2D space which means you can use it in a classifier. I've been investigating this and have shown that these 2D embeddings could display characteristics of the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\">ugly-duckling concept</a>.</p>\n\n<p>To do so simply follow Chris' excellent notebook (using UMAP) train and save a UMAP transformer and you can then project all of a patients images (either in the train or test) into the 2D space. For example consider these  images which shows four (hand selected) patients' images projected into the 2D space, where yellow points indicate images of malignant lesions.</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F5c9396bb0ba3b968f83b2d3110bb99c9%2Fpatient_embedding.png?generation=1595265660164754&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F5c9396bb0ba3b968f83b2d3110bb99c9%2Fpatient_embedding.png?generation=1595265660164754&amp;alt=media</a> =500x500)</p>\n\n<p>You can see that some malignant images are far away from what a 'normal' image is for a patient, although this is not always true. </p>\n\n<p>For every image (in training or test) one can then calculate the distance in this 2D space between its N nearest neighbours (for a given patient) and this distance shows some correlation to malignancy. For example, I've calculated it for N=3 and the below image shows the fraction of lesions which are malignant as a function of their mean distance to its three nearest neighbours - those further from there nearest-neighbours are more likely to be malignant.</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2Fd931eec054069c1fcef982b4377e8cb9%2Ffraction_malignant.png?generation=1595266011305628&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2Fd931eec054069c1fcef982b4377e8cb9%2Ffraction_malignant.png?generation=1595266011305628&amp;alt=media</a> =500x500)</p>\n\n<p>Of course, at some level this is to be expected: malignant lesions should not look like non-malignant lesions and should be far away in this 2D space. But this method could add additional information as it makes use of patient level information to characterise what is normal for a given patient.</p>\n\n<p>Using the mean distance of the image to its three nearest neighbours as a simple predictor yielded an OOF AUC of 0.7 but using it in my ensembles did not improve CV. However, the initial CNN (EFN B3) embedding is not trained on lesion images (image_weights='noisy-student'), so the method could improve if a better initial CNN embedding is used...</p>",
      "rawMarkdown": "cdeotte  [has shown us](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028) how one can use t-SNE and UMAP to project CNN image embeddings into a 2D space. \n\nIn the case of UMAP (unlike t-SNE) your model learns a direct transformation of the CNN embedding to the 2D space which means you can use it in a classifier. I've been investigating this and have shown that these 2D embeddings could display characteristics of the [ugly-duckling concept](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348).\n\nTo do so simply follow Chris' excellent notebook (using UMAP) train and save a UMAP transformer and you can then project all of a patients images (either in the train or test) into the 2D space. For example consider these  images which shows four (hand selected) patients' images projected into the 2D space, where yellow points indicate images of malignant lesions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F5c9396bb0ba3b968f83b2d3110bb99c9%2Fpatient_embedding.png?generation=1595265660164754&amp;alt=media =500x500)\n\nYou can see that some malignant images are far away from what a 'normal' image is for a patient, although this is not always true. \n\nFor every image (in training or test) one can then calculate the distance in this 2D space between its N nearest neighbours (for a given patient) and this distance shows some correlation to malignancy. For example, I've calculated it for N=3 and the below image shows the fraction of lesions which are malignant as a function of their mean distance to its three nearest neighbours - those further from there nearest-neighbours are more likely to be malignant.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2Fd931eec054069c1fcef982b4377e8cb9%2Ffraction_malignant.png?generation=1595266011305628&amp;alt=media =500x500)\n\nOf course, at some level this is to be expected: malignant lesions should not look like non-malignant lesions and should be far away in this 2D space. But this method could add additional information as it makes use of patient level information to characterise what is normal for a given patient.\n\nUsing the mean distance of the image to its three nearest neighbours as a simple predictor yielded an OOF AUC of 0.7 but using it in my ensembles did not improve CV. However, the initial CNN (EFN B3) embedding is not trained on lesion images (image_weights='noisy-student'), so the method could improve if a better initial CNN embedding is used...",
      "votes": null
    },
    {
      "id": "937126",
      "postDate": "07/20/2020 18:34:13",
      "content": "<p>Instead of UMAP, why not just look for prediction values which are much higher for the same patient?</p>\n\n<p>i.e. first generate predictions, and then perform a post-processing step where you group predictions by patient_id, and look for images whose scores are very different from the rest of the same patient's images (say using simple <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284\">Dixon's Q-test</a>)</p>",
      "rawMarkdown": "Instead of UMAP, why not just look for prediction values which are much higher for the same patient?\n\ni.e. first generate predictions, and then perform a post-processing step where you group predictions by patient_id, and look for images whose scores are very different from the rest of the same patient's images (say using simple [Dixon's Q-test](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284))",
      "votes": null
    },
    {
      "id": "937171",
      "postDate": "07/20/2020 19:19:51",
      "content": "<p>Good question, what you suggest could produce similar results but not necessarily the same result - its because (as far as I'm aware) the CNN predictions has no concept (without explicit feature engineering) of patients other images, it treats each one as independent. Additionally, when using UMAP the model has not explicitly learned anything about malignancy,  it's never given a target, but just clusters on the image features (which are generic ones learned from transfer-learning).</p>\n\n<p>For the method you outline the patient may have two very different images but with similarly low predictions (i.e., they are in different regions of the 2D embedding but the classifier sees both of these regions as the same risk). However, the UMAP method would detect the two images are far about in the embedded space and flag this. This could be relevant if someone has a lot of lesions and then just one very different to the primary group, the different one may be low-risk to many patients but perhaps the presence of one, when you have different lesion characteristics, is relevant.</p>",
      "rawMarkdown": "Good question, what you suggest could produce similar results but not necessarily the same result - its because (as far as I'm aware) the CNN predictions has no concept (without explicit feature engineering) of patients other images, it treats each one as independent. Additionally, when using UMAP the model has not explicitly learned anything about malignancy,  it's never given a target, but just clusters on the image features (which are generic ones learned from transfer-learning).\n\nFor the method you outline the patient may have two very different images but with similarly low predictions (i.e., they are in different regions of the 2D embedding but the classifier sees both of these regions as the same risk). However, the UMAP method would detect the two images are far about in the embedded space and flag this. This could be relevant if someone has a lot of lesions and then just one very different to the primary group, the different one may be low-risk to many patients but perhaps the presence of one, when you have different lesion characteristics, is relevant.",
      "votes": null
    },
    {
      "id": "937218",
      "postDate": "07/20/2020 20:32:22",
      "content": "<p>Fantastic stuff! But: what about (\"simply\") using Cosine Distance of the feature vector between images of a patient?</p>",
      "rawMarkdown": "Fantastic stuff! But: what about (\"simply\") using Cosine Distance of the feature vector between images of a patient?",
      "votes": null
    },
    {
      "id": "937246",
      "postDate": "07/20/2020 21:05:50",
      "content": "<p>That may work better in a predictive model - definetly worth checking. My only thought is the feature vector is high dimensionality (1000 + dimensions), may make things difficult? Other than that I can't see why it would be worse than UMAP. UMAP has the disadvantage of losing information but has the advantage of allowing you to visualise the space, useful for investigating the ugly duckling concept for a given set of patients.</p>",
      "rawMarkdown": "That may work better in a predictive model - definetly worth checking. My only thought is the feature vector is high dimensionality (1000 + dimensions), may make things difficult? Other than that I can't see why it would be worse than UMAP. UMAP has the disadvantage of losing information but has the advantage of allowing you to visualise the space, useful for investigating the ugly duckling concept for a given set of patients.",
      "votes": null
    },
    {
      "id": "937611",
      "postDate": "07/21/2020 05:23:08",
      "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> Hi sirishks, Thanks for point out this, after Dixon Q Test, how can we merge UD result with model pred scores?</p>",
      "rawMarkdown": "sirishks Hi sirishks, Thanks for point out this, after Dixon Q Test, how can we merge UD result with model pred scores?",
      "votes": null
    },
    {
      "id": "939393",
      "postDate": "07/22/2020 08:04:38",
      "content": "<p>Hi - afaik Cosine Distance is quite beneficial for comparisons in high dimensionality as well in this case, and in my experience captures visual \"similarity\" quite well (<a href=\"https://ypsono.com/search/cbir\" target=\"_blank\">try-out possible here</a>), and could potentially be harnessed for <a href=\"https://onlinelibrary.wiley.com/doi/full/10.1111/bjd.17189\" target=\"_blank\">classification-by-retrieval</a>. </p>\n<p>Would be great if you let us know if it worked if you try it out :)</p>",
      "rawMarkdown": "Hi - afaik Cosine Distance is quite beneficial for comparisons in high dimensionality as well in this case, and in my experience captures visual \"similarity\" quite well ([try-out possible here](https://ypsono.com/search/cbir)), and could potentially be harnessed for [classification-by-retrieval](https://onlinelibrary.wiley.com/doi/full/10.1111/bjd.17189). \n\nWould be great if you let us know if it worked if you try it out :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937126,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "07/20/2020 18:34:13",
      "content": "<p>Instead of UMAP, why not just look for prediction values which are much higher for the same patient?</p>\n\n<p>i.e. first generate predictions, and then perform a post-processing step where you group predictions by patient_id, and look for images whose scores are very different from the rest of the same patient's images (say using simple <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284\">Dixon's Q-test</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 937171,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "07/20/2020 19:19:51",
          "content": "<p>Good question, what you suggest could produce similar results but not necessarily the same result - its because (as far as I'm aware) the CNN predictions has no concept (without explicit feature engineering) of patients other images, it treats each one as independent. Additionally, when using UMAP the model has not explicitly learned anything about malignancy,  it's never given a target, but just clusters on the image features (which are generic ones learned from transfer-learning).</p>\n\n<p>For the method you outline the patient may have two very different images but with similarly low predictions (i.e., they are in different regions of the 2D embedding but the classifier sees both of these regions as the same risk). However, the UMAP method would detect the two images are far about in the embedded space and flag this. This could be relevant if someone has a lot of lesions and then just one very different to the primary group, the different one may be low-risk to many patients but perhaps the presence of one, when you have different lesion characteristics, is relevant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937611,
          "author_name": "feiwofeifeixiaowo",
          "author_url": "",
          "post_date": "07/21/2020 05:23:08",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> Hi sirishks, Thanks for point out this, after Dixon Q Test, how can we merge UD result with model pred scores?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937218,
      "author_name": "tschandl",
      "author_url": "",
      "post_date": "07/20/2020 20:32:22",
      "content": "<p>Fantastic stuff! But: what about (\"simply\") using Cosine Distance of the feature vector between images of a patient?</p>",
      "votes": null,
      "replies": [
        {
          "id": 937246,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "07/20/2020 21:05:50",
          "content": "<p>That may work better in a predictive model - definetly worth checking. My only thought is the feature vector is high dimensionality (1000 + dimensions), may make things difficult? Other than that I can't see why it would be worse than UMAP. UMAP has the disadvantage of losing information but has the advantage of allowing you to visualise the space, useful for investigating the ugly duckling concept for a given set of patients.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939393,
          "author_name": "tschandl",
          "author_url": "",
          "post_date": "07/22/2020 08:04:38",
          "content": "<p>Hi - afaik Cosine Distance is quite beneficial for comparisons in high dimensionality as well in this case, and in my experience captures visual \"similarity\" quite well (<a href=\"https://ypsono.com/search/cbir\" target=\"_blank\">try-out possible here</a>), and could potentially be harnessed for <a href=\"https://onlinelibrary.wiley.com/doi/full/10.1111/bjd.17189\" target=\"_blank\">classification-by-retrieval</a>. </p>\n<p>Would be great if you let us know if it worked if you try it out :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "937055": "cdeotte  [has shown us](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028) how one can use t-SNE and UMAP to project CNN image embeddings into a 2D space. \n\nIn the case of UMAP (unlike t-SNE) your model learns a direct transformation of the CNN embedding to the 2D space which means you can use it in a classifier. I've been investigating this and have shown that these 2D embeddings could display characteristics of the [ugly-duckling concept](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348).\n\nTo do so simply follow Chris' excellent notebook (using UMAP) train and save a UMAP transformer and you can then project all of a patients images (either in the train or test) into the 2D space. For example consider these  images which shows four (hand selected) patients' images projected into the 2D space, where yellow points indicate images of malignant lesions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F5c9396bb0ba3b968f83b2d3110bb99c9%2Fpatient_embedding.png?generation=1595265660164754&amp;alt=media =500x500)\n\nYou can see that some malignant images are far away from what a 'normal' image is for a patient, although this is not always true. \n\nFor every image (in training or test) one can then calculate the distance in this 2D space between its N nearest neighbours (for a given patient) and this distance shows some correlation to malignancy. For example, I've calculated it for N=3 and the below image shows the fraction of lesions which are malignant as a function of their mean distance to its three nearest neighbours - those further from there nearest-neighbours are more likely to be malignant.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2Fd931eec054069c1fcef982b4377e8cb9%2Ffraction_malignant.png?generation=1595266011305628&amp;alt=media =500x500)\n\nOf course, at some level this is to be expected: malignant lesions should not look like non-malignant lesions and should be far away in this 2D space. But this method could add additional information as it makes use of patient level information to characterise what is normal for a given patient.\n\nUsing the mean distance of the image to its three nearest neighbours as a simple predictor yielded an OOF AUC of 0.7 but using it in my ensembles did not improve CV. However, the initial CNN (EFN B3) embedding is not trained on lesion images (image_weights='noisy-student'), so the method could improve if a better initial CNN embedding is used...",
    "937126": "Instead of UMAP, why not just look for prediction values which are much higher for the same patient?\n\ni.e. first generate predictions, and then perform a post-processing step where you group predictions by patient_id, and look for images whose scores are very different from the rest of the same patient's images (say using simple [Dixon's Q-test](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284))",
    "937171": "Good question, what you suggest could produce similar results but not necessarily the same result - its because (as far as I'm aware) the CNN predictions has no concept (without explicit feature engineering) of patients other images, it treats each one as independent. Additionally, when using UMAP the model has not explicitly learned anything about malignancy,  it's never given a target, but just clusters on the image features (which are generic ones learned from transfer-learning).\n\nFor the method you outline the patient may have two very different images but with similarly low predictions (i.e., they are in different regions of the 2D embedding but the classifier sees both of these regions as the same risk). However, the UMAP method would detect the two images are far about in the embedded space and flag this. This could be relevant if someone has a lot of lesions and then just one very different to the primary group, the different one may be low-risk to many patients but perhaps the presence of one, when you have different lesion characteristics, is relevant.",
    "937218": "Fantastic stuff! But: what about (\"simply\") using Cosine Distance of the feature vector between images of a patient?",
    "937246": "That may work better in a predictive model - definetly worth checking. My only thought is the feature vector is high dimensionality (1000 + dimensions), may make things difficult? Other than that I can't see why it would be worse than UMAP. UMAP has the disadvantage of losing information but has the advantage of allowing you to visualise the space, useful for investigating the ugly duckling concept for a given set of patients.",
    "937611": "sirishks Hi sirishks, Thanks for point out this, after Dixon Q Test, how can we merge UD result with model pred scores?",
    "939393": "Hi - afaik Cosine Distance is quite beneficial for comparisons in high dimensionality as well in this case, and in my experience captures visual \"similarity\" quite well ([try-out possible here](https://ypsono.com/search/cbir)), and could potentially be harnessed for [classification-by-retrieval](https://onlinelibrary.wiley.com/doi/full/10.1111/bjd.17189). \n\nWould be great if you let us know if it worked if you try it out :)"
  },
  "source": "meta"
}