{
  "id": 175500,
  "title": "Ugly Duckling: Did Anything Work?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175500",
  "author_name": "",
  "post_date": "2020-08-18T11:14:46.923369300Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Reading through the solution overviews, it appears that accounting for the \"ugly duckling\" concept was not very useful among the top teams. I was wondering if anyone had any luck with incorporating it into their solution and would like to share findings from our experiments on that matter. </p>\n<h1>\"Ugly Duckling\" Concept</h1>\n<p>At the beginning of the competition, there were multiple discussion topics on the \"ugly duckling\" concept:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\" target=\"_blank\">Understanding \"Ugly Duckling Concept\"</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284\" target=\"_blank\">Dixon's Q-test to identify \"Ugly duckling\" outliers</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456\" target=\"_blank\">Ugly duckling detection - UMAP and pairwise distance</a></li>\n</ul>\n<p>In a nutshell, the idea is that a lesion that is very different from the other lesions of the same patient is more likely to be malignant, even if the lesion itself demonstrates fewer signs of being malignant when considered as a single image. This seemed a very promising direction, so we looked at several options to account for it. </p>\n<h1>Our Attempts</h1>\n<p>We tried two approaches to account for the \"ugly duckling\". </p>\n<p>First, we used CNN to extract image embeddings similar to <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">this notebook from Chirs</a>. We then use KNN to compute mean distances to <code>[1, 2, ..., 5]</code> closest neighbors within the same patient if they have at least 5 images. Higher mean distance indicates outlier images that are more likely to be melanoma. Ranking images by mean distances achieved <code>0.7733</code> on CV but failed on LB with just <code>0.5346</code> public. Below is an example of how mean distance to three closest images separates the two classes on training data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554091%2Fd65477466cf1357478d51fc839606149%2Fugly_dist.jpg?generation=1597747035090498&amp;alt=media\" alt=\"\"></p>\n<p>The second approach was identifying patient-level outliers by inspecting CNN predictions as a part of a post-processing technique. For each patient, we identified images where the model outputs a prediction that exceeds the second-highest prediction per this patient by a gap <code>t</code> and multiplied such predictions by <code>k &gt; 1</code>. In theory, this should help to increase the rank of images that stand out within a patient even if the absolute value of the prediction is not very high. We tried to optimize <code>t</code> and <code>k</code> within the CV and achieved a boost of <code>+0.0004</code> OOF AUC by using <code>t = 0.5</code> and <code>k = 1.25</code>. However, using this post-processing technique with our best model resulted in the LB drop of <code>0.0005</code>. </p>\n<p>I have also experimented with <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284\" target=\"_blank\">Dixon's Q-test</a> suggested by <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> but did not have any luck.  Adding any of the distance measures, prediction gaps and Q-test statistic values as features to the meta-model also did not help to boost the LB performance despite improving the CV. In the end, we did not account for the patient-level image differences in our final submissions.</p>\n<p>Did any of you have luck with similar or different options to account for the \"ugly duckling\"?</p>\n<h1>Other Patient-Level Information</h1>\n<p>Patient-level information can be used beyond the \"ugly-duckling\" concept as well. <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175467\" target=\"_blank\">This discussion topic</a> covers a few ideas to use contextual information. Judging from the messages posted so far, there has not been much success either.</p>",
  "messages": [
    {
      "id": "975582",
      "postDate": "08/18/2020 11:14:46",
      "content": "<p>Reading through the solution overviews, it appears that accounting for the \"ugly duckling\" concept was not very useful among the top teams. I was wondering if anyone had any luck with incorporating it into their solution and would like to share findings from our experiments on that matter. </p>\n<h1>\"Ugly Duckling\" Concept</h1>\n<p>At the beginning of the competition, there were multiple discussion topics on the \"ugly duckling\" concept:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\" target=\"_blank\">Understanding \"Ugly Duckling Concept\"</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284\" target=\"_blank\">Dixon's Q-test to identify \"Ugly duckling\" outliers</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456\" target=\"_blank\">Ugly duckling detection - UMAP and pairwise distance</a></li>\n</ul>\n<p>In a nutshell, the idea is that a lesion that is very different from the other lesions of the same patient is more likely to be malignant, even if the lesion itself demonstrates fewer signs of being malignant when considered as a single image. This seemed a very promising direction, so we looked at several options to account for it. </p>\n<h1>Our Attempts</h1>\n<p>We tried two approaches to account for the \"ugly duckling\". </p>\n<p>First, we used CNN to extract image embeddings similar to <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">this notebook from Chirs</a>. We then use KNN to compute mean distances to <code>[1, 2, ..., 5]</code> closest neighbors within the same patient if they have at least 5 images. Higher mean distance indicates outlier images that are more likely to be melanoma. Ranking images by mean distances achieved <code>0.7733</code> on CV but failed on LB with just <code>0.5346</code> public. Below is an example of how mean distance to three closest images separates the two classes on training data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554091%2Fd65477466cf1357478d51fc839606149%2Fugly_dist.jpg?generation=1597747035090498&amp;alt=media\" alt=\"\"></p>\n<p>The second approach was identifying patient-level outliers by inspecting CNN predictions as a part of a post-processing technique. For each patient, we identified images where the model outputs a prediction that exceeds the second-highest prediction per this patient by a gap <code>t</code> and multiplied such predictions by <code>k &gt; 1</code>. In theory, this should help to increase the rank of images that stand out within a patient even if the absolute value of the prediction is not very high. We tried to optimize <code>t</code> and <code>k</code> within the CV and achieved a boost of <code>+0.0004</code> OOF AUC by using <code>t = 0.5</code> and <code>k = 1.25</code>. However, using this post-processing technique with our best model resulted in the LB drop of <code>0.0005</code>. </p>\n<p>I have also experimented with <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284\" target=\"_blank\">Dixon's Q-test</a> suggested by <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> but did not have any luck.  Adding any of the distance measures, prediction gaps and Q-test statistic values as features to the meta-model also did not help to boost the LB performance despite improving the CV. In the end, we did not account for the patient-level image differences in our final submissions.</p>\n<p>Did any of you have luck with similar or different options to account for the \"ugly duckling\"?</p>\n<h1>Other Patient-Level Information</h1>\n<p>Patient-level information can be used beyond the \"ugly-duckling\" concept as well. <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175467\" target=\"_blank\">This discussion topic</a> covers a few ideas to use contextual information. Judging from the messages posted so far, there has not been much success either.</p>",
      "rawMarkdown": "Reading through the solution overviews, it appears that accounting for the \"ugly duckling\" concept was not very useful among the top teams. I was wondering if anyone had any luck with incorporating it into their solution and would like to share findings from our experiments on that matter. \n\n\n# \"Ugly Duckling\" Concept\n\nAt the beginning of the competition, there were multiple discussion topics on the \"ugly duckling\" concept:\n- [Understanding \"Ugly Duckling Concept\"](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348)\n- [Dixon's Q-test to identify \"Ugly duckling\" outliers](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284)\n- [Ugly duckling detection - UMAP and pairwise distance](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456)\n\nIn a nutshell, the idea is that a lesion that is very different from the other lesions of the same patient is more likely to be malignant, even if the lesion itself demonstrates fewer signs of being malignant when considered as a single image. This seemed a very promising direction, so we looked at several options to account for it. \n\n\n# Our Attempts \n\nWe tried two approaches to account for the \"ugly duckling\". \n\nFirst, we used CNN to extract image embeddings similar to [this notebook from Chirs](https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates). We then use KNN to compute mean distances to `[1, 2, ..., 5]` closest neighbors within the same patient if they have at least 5 images. Higher mean distance indicates outlier images that are more likely to be melanoma. Ranking images by mean distances achieved `0.7733` on CV but failed on LB with just `0.5346` public. Below is an example of how mean distance to three closest images separates the two classes on training data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554091%2Fd65477466cf1357478d51fc839606149%2Fugly_dist.jpg?generation=1597747035090498&alt=media)\n\n\nThe second approach was identifying patient-level outliers by inspecting CNN predictions as a part of a post-processing technique. For each patient, we identified images where the model outputs a prediction that exceeds the second-highest prediction per this patient by a gap `t` and multiplied such predictions by `k > 1`. In theory, this should help to increase the rank of images that stand out within a patient even if the absolute value of the prediction is not very high. We tried to optimize `t` and `k` within the CV and achieved a boost of `+0.0004` OOF AUC by using `t = 0.5` and `k = 1.25`. However, using this post-processing technique with our best model resulted in the LB drop of `0.0005`. \n\nI have also experimented with [Dixon's Q-test](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284) suggested by @sirishks but did not have any luck.  Adding any of the distance measures, prediction gaps and Q-test statistic values as features to the meta-model also did not help to boost the LB performance despite improving the CV. In the end, we did not account for the patient-level image differences in our final submissions.\n\nDid any of you have luck with similar or different options to account for the \"ugly duckling\"?\n\n\n# Other Patient-Level Information\n\nPatient-level information can be used beyond the \"ugly-duckling\" concept as well. [This discussion topic](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175467) covers a few ideas to use contextual information. Judging from the messages posted so far, there has not been much success either.",
      "votes": null
    },
    {
      "id": "993643",
      "postDate": "09/01/2020 05:08:57",
      "content": "<p>My primary interest in this competition was looking for some kind of advantage using the \"Ugly Duckling\" concept since it might have interesting implications for a number of computer vision tasks and future research. I had trouble implementing some of my ideas since it was a little beyond my skill level and time investment, but it seems like overall my experience was similar to yours: </p>\n<ul>\n<li>Some trivial/unstable improvements using post-processing based on patient-level outliers.</li>\n<li>No improvements gained by adding \"context embeddings\" from other images from the patient. This just blew up my model complexity and training time.</li>\n</ul>\n<p>One of the challenges finding a meaningful \"Ugly Duckling\" effect in this data is the distribution of patient images. Many patients had only ~ 3 images (meaning that there was one image being classified with 2 context/baseline images), while some others had hundreds. So I wanted to develop some kind of aggregation over the \"context\" images, but aggregation over 2 context images was almost meaningless.</p>\n<p>One other area I would have liked to investigate but didn't have time, is patient specific pre-processing. Since some visual characteristics like skin-tone vary mostly at the patient level, applying patient level colour adjustments might help to improve generalizability over different types of skin.</p>",
      "rawMarkdown": "My primary interest in this competition was looking for some kind of advantage using the \"Ugly Duckling\" concept since it might have interesting implications for a number of computer vision tasks and future research. I had trouble implementing some of my ideas since it was a little beyond my skill level and time investment, but it seems like overall my experience was similar to yours: \n- Some trivial/unstable improvements using post-processing based on patient-level outliers.\n- No improvements gained by adding \"context embeddings\" from other images from the patient. This just blew up my model complexity and training time.\n\nOne of the challenges finding a meaningful \"Ugly Duckling\" effect in this data is the distribution of patient images. Many patients had only ~ 3 images (meaning that there was one image being classified with 2 context/baseline images), while some others had hundreds. So I wanted to develop some kind of aggregation over the \"context\" images, but aggregation over 2 context images was almost meaningless.\n\nOne other area I would have liked to investigate but didn't have time, is patient specific pre-processing. Since some visual characteristics like skin-tone vary mostly at the patient level, applying patient level colour adjustments might help to improve generalizability over different types of skin.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 993643,
      "author_name": "ravy101",
      "author_url": "",
      "post_date": "09/01/2020 05:08:57",
      "content": "<p>My primary interest in this competition was looking for some kind of advantage using the \"Ugly Duckling\" concept since it might have interesting implications for a number of computer vision tasks and future research. I had trouble implementing some of my ideas since it was a little beyond my skill level and time investment, but it seems like overall my experience was similar to yours: </p>\n<ul>\n<li>Some trivial/unstable improvements using post-processing based on patient-level outliers.</li>\n<li>No improvements gained by adding \"context embeddings\" from other images from the patient. This just blew up my model complexity and training time.</li>\n</ul>\n<p>One of the challenges finding a meaningful \"Ugly Duckling\" effect in this data is the distribution of patient images. Many patients had only ~ 3 images (meaning that there was one image being classified with 2 context/baseline images), while some others had hundreds. So I wanted to develop some kind of aggregation over the \"context\" images, but aggregation over 2 context images was almost meaningless.</p>\n<p>One other area I would have liked to investigate but didn't have time, is patient specific pre-processing. Since some visual characteristics like skin-tone vary mostly at the patient level, applying patient level colour adjustments might help to improve generalizability over different types of skin.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "975582": "Reading through the solution overviews, it appears that accounting for the \"ugly duckling\" concept was not very useful among the top teams. I was wondering if anyone had any luck with incorporating it into their solution and would like to share findings from our experiments on that matter. \n\n\n# \"Ugly Duckling\" Concept\n\nAt the beginning of the competition, there were multiple discussion topics on the \"ugly duckling\" concept:\n- [Understanding \"Ugly Duckling Concept\"](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348)\n- [Dixon's Q-test to identify \"Ugly duckling\" outliers](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284)\n- [Ugly duckling detection - UMAP and pairwise distance](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456)\n\nIn a nutshell, the idea is that a lesion that is very different from the other lesions of the same patient is more likely to be malignant, even if the lesion itself demonstrates fewer signs of being malignant when considered as a single image. This seemed a very promising direction, so we looked at several options to account for it. \n\n\n# Our Attempts \n\nWe tried two approaches to account for the \"ugly duckling\". \n\nFirst, we used CNN to extract image embeddings similar to [this notebook from Chirs](https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates). We then use KNN to compute mean distances to `[1, 2, ..., 5]` closest neighbors within the same patient if they have at least 5 images. Higher mean distance indicates outlier images that are more likely to be melanoma. Ranking images by mean distances achieved `0.7733` on CV but failed on LB with just `0.5346` public. Below is an example of how mean distance to three closest images separates the two classes on training data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554091%2Fd65477466cf1357478d51fc839606149%2Fugly_dist.jpg?generation=1597747035090498&alt=media)\n\n\nThe second approach was identifying patient-level outliers by inspecting CNN predictions as a part of a post-processing technique. For each patient, we identified images where the model outputs a prediction that exceeds the second-highest prediction per this patient by a gap `t` and multiplied such predictions by `k > 1`. In theory, this should help to increase the rank of images that stand out within a patient even if the absolute value of the prediction is not very high. We tried to optimize `t` and `k` within the CV and achieved a boost of `+0.0004` OOF AUC by using `t = 0.5` and `k = 1.25`. However, using this post-processing technique with our best model resulted in the LB drop of `0.0005`. \n\nI have also experimented with [Dixon's Q-test](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156284) suggested by @sirishks but did not have any luck.  Adding any of the distance measures, prediction gaps and Q-test statistic values as features to the meta-model also did not help to boost the LB performance despite improving the CV. In the end, we did not account for the patient-level image differences in our final submissions.\n\nDid any of you have luck with similar or different options to account for the \"ugly duckling\"?\n\n\n# Other Patient-Level Information\n\nPatient-level information can be used beyond the \"ugly-duckling\" concept as well. [This discussion topic](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175467) covers a few ideas to use contextual information. Judging from the messages posted so far, there has not been much success either.",
    "993643": "My primary interest in this competition was looking for some kind of advantage using the \"Ugly Duckling\" concept since it might have interesting implications for a number of computer vision tasks and future research. I had trouble implementing some of my ideas since it was a little beyond my skill level and time investment, but it seems like overall my experience was similar to yours: \n- Some trivial/unstable improvements using post-processing based on patient-level outliers.\n- No improvements gained by adding \"context embeddings\" from other images from the patient. This just blew up my model complexity and training time.\n\nOne of the challenges finding a meaningful \"Ugly Duckling\" effect in this data is the distribution of patient images. Many patients had only ~ 3 images (meaning that there was one image being classified with 2 context/baseline images), while some others had hundreds. So I wanted to develop some kind of aggregation over the \"context\" images, but aggregation over 2 context images was almost meaningless.\n\nOne other area I would have liked to investigate but didn't have time, is patient specific pre-processing. Since some visual characteristics like skin-tone vary mostly at the patient level, applying patient level colour adjustments might help to improve generalizability over different types of skin."
  },
  "source": "meta"
}