{
  "id": 158829,
  "title": "Interactive 3D data visualization with UMAP on RGB histograms with Fisher metric",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/158829",
  "author_name": "",
  "post_date": "2020-06-15T14:34:02.794164500Z",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F0e09c8550d8320bca4a362ba80ce82dc%2Fumap_with_histograms.png?generation=1592226072209541&amp;alt=media\" alt=\"\"></p>\n\n<p>I tried to understand the data a bit better and played with data visualizations. One informative visualization uses 3D UMAP embedding in the space of RGB histograms. You can see a summary on the image above. Here is the link to the <a href=\"https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms\">kernel</a></p>\n\n<h1>About the kernel:</h1>\n\n<ul>\n<li><p>it uses GPU (rapids UMAP is very fast)</p></li>\n<li><p>if you don't want to use GPU, you can browse for example this <a href=\"https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms?scriptVersionId=36429142\">version</a> -- the js plot is interactive, except clicking on points doesn't do anything.</p></li>\n<li><p>the 3D plots are interactive, hovering over a point shows the related tabular data </p></li>\n<li>clicking on a data point shows 10 most similar images + histograms. For efficiency, the similarity search is done in the space of UMAP embedding, but the results seem reasonable.</li>\n<li>I used slightly trimmed images, because some images have white/black vignettes (more on this below)</li>\n<li>there seems to be a bug in plotly: custom colors seems to screw up point selections</li>\n</ul>\n\n<h1>Some insights:</h1>\n\n<ul>\n<li><p>Initially I worked on (histograms of) original images. This yields two extra clusters containing images with white/black vignettes. The white vignette seems added in postprocessing, the black one may come from the imaging method. There are 109 images with white vignette, 22% of which are malignant -- much higher percentage than normally. </p></li>\n<li><p>The enlarged points correspond to target==1. As you can see, there is a big portion of data -- the  'fin' extending to the right of the UMAP plot -- where malignant cases are very rare. More precisely, there are 12345 (sic!) train images in this part, out of which only 13 are malignant. It seems this part corresponds to images with healthy-looking pinkish skin with not-too-dark changes.</p></li>\n<li><p>[Updated]: I also checked that the train and test data have similar structure without visible outliers. It'd be interesting to see if the available external data is similar.</p></li>\n<li><p>The presence of apparent loops in the UMAP embedding is quite interesting, and it's present for all (~50) random seeds I checked. Rapids fast GPU implementation of UMAP was definitely useful for that. I would speculate that these loops are present in the original space, since UMAP should be good at preserving such topological features.</p></li>\n</ul>\n\n<h1>Technicalities:</h1>\n\n<ul>\n<li>I use the Fisher information metric to compute the distance between two histograms. In terms of implementation it boils down to taking the np.sqrt of each histogram -- sqrt is an isometry between the Fisher space and the Euclidean space. So running UMAP on the transformed space is equivalent to running UMAP with Fisher metric as the source metric, except rapids UMAP doesn't seem to support custom metrics. This is quite an important point -- for example using the Euclidean distance to compare histograms fails to capture one of two extra components; also the loops are not as prominently separated.</li>\n</ul>\n\n<h1>Summary:</h1>\n\n<p>I hope someone finds this interesting. I will play with this more -- feel free to let me know if you'd like to see some particular experiment or need more details about the above.</p>",
  "messages": [
    {
      "id": "887203",
      "postDate": "06/15/2020 14:34:02",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F0e09c8550d8320bca4a362ba80ce82dc%2Fumap_with_histograms.png?generation=1592226072209541&amp;alt=media\" alt=\"\"></p>\n\n<p>I tried to understand the data a bit better and played with data visualizations. One informative visualization uses 3D UMAP embedding in the space of RGB histograms. You can see a summary on the image above. Here is the link to the <a href=\"https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms\">kernel</a></p>\n\n<h1>About the kernel:</h1>\n\n<ul>\n<li><p>it uses GPU (rapids UMAP is very fast)</p></li>\n<li><p>if you don't want to use GPU, you can browse for example this <a href=\"https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms?scriptVersionId=36429142\">version</a> -- the js plot is interactive, except clicking on points doesn't do anything.</p></li>\n<li><p>the 3D plots are interactive, hovering over a point shows the related tabular data </p></li>\n<li>clicking on a data point shows 10 most similar images + histograms. For efficiency, the similarity search is done in the space of UMAP embedding, but the results seem reasonable.</li>\n<li>I used slightly trimmed images, because some images have white/black vignettes (more on this below)</li>\n<li>there seems to be a bug in plotly: custom colors seems to screw up point selections</li>\n</ul>\n\n<h1>Some insights:</h1>\n\n<ul>\n<li><p>Initially I worked on (histograms of) original images. This yields two extra clusters containing images with white/black vignettes. The white vignette seems added in postprocessing, the black one may come from the imaging method. There are 109 images with white vignette, 22% of which are malignant -- much higher percentage than normally. </p></li>\n<li><p>The enlarged points correspond to target==1. As you can see, there is a big portion of data -- the  'fin' extending to the right of the UMAP plot -- where malignant cases are very rare. More precisely, there are 12345 (sic!) train images in this part, out of which only 13 are malignant. It seems this part corresponds to images with healthy-looking pinkish skin with not-too-dark changes.</p></li>\n<li><p>[Updated]: I also checked that the train and test data have similar structure without visible outliers. It'd be interesting to see if the available external data is similar.</p></li>\n<li><p>The presence of apparent loops in the UMAP embedding is quite interesting, and it's present for all (~50) random seeds I checked. Rapids fast GPU implementation of UMAP was definitely useful for that. I would speculate that these loops are present in the original space, since UMAP should be good at preserving such topological features.</p></li>\n</ul>\n\n<h1>Technicalities:</h1>\n\n<ul>\n<li>I use the Fisher information metric to compute the distance between two histograms. In terms of implementation it boils down to taking the np.sqrt of each histogram -- sqrt is an isometry between the Fisher space and the Euclidean space. So running UMAP on the transformed space is equivalent to running UMAP with Fisher metric as the source metric, except rapids UMAP doesn't seem to support custom metrics. This is quite an important point -- for example using the Euclidean distance to compare histograms fails to capture one of two extra components; also the loops are not as prominently separated.</li>\n</ul>\n\n<h1>Summary:</h1>\n\n<p>I hope someone finds this interesting. I will play with this more -- feel free to let me know if you'd like to see some particular experiment or need more details about the above.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F0e09c8550d8320bca4a362ba80ce82dc%2Fumap_with_histograms.png?generation=1592226072209541&amp;alt=media)\n\nI tried to understand the data a bit better and played with data visualizations. One informative visualization uses 3D UMAP embedding in the space of RGB histograms. You can see a summary on the image above. Here is the link to the [kernel](https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms)\n\n# About the kernel:\n- it uses GPU (rapids UMAP is very fast)\n\n- if you don't want to use GPU, you can browse for example this [version](https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms?scriptVersionId=36429142) -- the js plot is interactive, except clicking on points doesn't do anything.\n\n- the 3D plots are interactive, hovering over a point shows the related tabular data \n- clicking on a data point shows 10 most similar images + histograms. For efficiency, the similarity search is done in the space of UMAP embedding, but the results seem reasonable.\n- I used slightly trimmed images, because some images have white/black vignettes (more on this below)\n- there seems to be a bug in plotly: custom colors seems to screw up point selections\n\n# Some insights:\n- Initially I worked on (histograms of) original images. This yields two extra clusters containing images with white/black vignettes. The white vignette seems added in postprocessing, the black one may come from the imaging method. There are 109 images with white vignette, 22% of which are malignant -- much higher percentage than normally. \n\n- The enlarged points correspond to target==1. As you can see, there is a big portion of data -- the  'fin' extending to the right of the UMAP plot -- where malignant cases are very rare. More precisely, there are 12345 (sic!) train images in this part, out of which only 13 are malignant. It seems this part corresponds to images with healthy-looking pinkish skin with not-too-dark changes.\n\n- [Updated]: I also checked that the train and test data have similar structure without visible outliers. It'd be interesting to see if the available external data is similar.\n\n- The presence of apparent loops in the UMAP embedding is quite interesting, and it's present for all (~50) random seeds I checked. Rapids fast GPU implementation of UMAP was definitely useful for that. I would speculate that these loops are present in the original space, since UMAP should be good at preserving such topological features.\n\n# Technicalities:\n- I use the Fisher information metric to compute the distance between two histograms. In terms of implementation it boils down to taking the np.sqrt of each histogram -- sqrt is an isometry between the Fisher space and the Euclidean space. So running UMAP on the transformed space is equivalent to running UMAP with Fisher metric as the source metric, except rapids UMAP doesn't seem to support custom metrics. This is quite an important point -- for example using the Euclidean distance to compare histograms fails to capture one of two extra components; also the loops are not as prominently separated.\n\n# Summary:\nI hope someone finds this interesting. I will play with this more -- feel free to let me know if you'd like to see some particular experiment or need more details about the above.",
      "votes": null
    },
    {
      "id": "887343",
      "postDate": "06/15/2020 16:07:13",
      "content": "<p>Hi, super nice visualisation. I like it very much. I did not know it is possible to have callbacks to python in notebooks.</p>\n\n<p>You write that train and test have similar structure. This is correct, but if you compare test and train, there is a big cluster in test, with not so many corresponding images in train. Have a look at <a href=\"https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference\">https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference</a> </p>",
      "rawMarkdown": "Hi, super nice visualisation. I like it very much. I did not know it is possible to have callbacks to python in notebooks.\n\nYou write that train and test have similar structure. This is correct, but if you compare test and train, there is a big cluster in test, with not so many corresponding images in train. Have a look at [https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference](https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference)",
      "votes": null
    },
    {
      "id": "887662",
      "postDate": "06/15/2020 19:40:11",
      "content": "<p>Thanks!</p>\n\n<p>You're right. I had a closer look at my embedding and there is a similar hot-spot of test images. It's in the lower-left corner my summary image. Here is a sample from this region:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F7ddf79b3df7773d108350da7a7468585%2Fnew_type.png?generation=1592248984235844&amp;alt=media\" alt=\"\"></p>\n\n<p>~80% of the data in this region is test data. I'm curious how many of those images are in the public test set -- and can we really generalize from train data in this region... Do we actually know what's the public-private split? </p>",
      "rawMarkdown": "Thanks!\n\nYou're right. I had a closer look at my embedding and there is a similar hot-spot of test images. It's in the lower-left corner my summary image. Here is a sample from this region:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F7ddf79b3df7773d108350da7a7468585%2Fnew_type.png?generation=1592248984235844&amp;alt=media)\n\n~80% of the data in this region is test data. I'm curious how many of those images are in the public test set -- and can we really generalize from train data in this region... Do we actually know what's the public-private split?",
      "votes": null
    },
    {
      "id": "961870",
      "postDate": "08/07/2020 15:16:43",
      "content": "<p>I really like the visualization, Great work!</p>",
      "rawMarkdown": "I really like the visualization, Great work!",
      "votes": null
    },
    {
      "id": "2327681",
      "postDate": "07/03/2023 05:38:14",
      "content": "<p>Realtime RGB Histogram is here : <a href=\"https://youtu.be/k5rLn7VlAhI\" target=\"_blank\">https://youtu.be/k5rLn7VlAhI</a></p>",
      "rawMarkdown": "Realtime RGB Histogram is here : https://youtu.be/k5rLn7VlAhI",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961870,
      "author_name": "delllectron",
      "author_url": "",
      "post_date": "08/07/2020 15:16:43",
      "content": "<p>I really like the visualization, Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887343,
      "author_name": "agentauers",
      "author_url": "",
      "post_date": "06/15/2020 16:07:13",
      "content": "<p>Hi, super nice visualisation. I like it very much. I did not know it is possible to have callbacks to python in notebooks.</p>\n\n<p>You write that train and test have similar structure. This is correct, but if you compare test and train, there is a big cluster in test, with not so many corresponding images in train. Have a look at <a href=\"https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference\">https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 887662,
          "author_name": "hubwag",
          "author_url": "",
          "post_date": "06/15/2020 19:40:11",
          "content": "<p>Thanks!</p>\n\n<p>You're right. I had a closer look at my embedding and there is a similar hot-spot of test images. It's in the lower-left corner my summary image. Here is a sample from this region:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F7ddf79b3df7773d108350da7a7468585%2Fnew_type.png?generation=1592248984235844&amp;alt=media\" alt=\"\"></p>\n\n<p>~80% of the data in this region is test data. I'm curious how many of those images are in the public test set -- and can we really generalize from train data in this region... Do we actually know what's the public-private split? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2327681,
              "author_name": "bemorekgg",
              "author_url": "",
              "post_date": "07/03/2023 05:38:14",
              "content": "<p>Realtime RGB Histogram is here : <a href=\"https://youtu.be/k5rLn7VlAhI\" target=\"_blank\">https://youtu.be/k5rLn7VlAhI</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "887203": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F0e09c8550d8320bca4a362ba80ce82dc%2Fumap_with_histograms.png?generation=1592226072209541&amp;alt=media)\n\nI tried to understand the data a bit better and played with data visualizations. One informative visualization uses 3D UMAP embedding in the space of RGB histograms. You can see a summary on the image above. Here is the link to the [kernel](https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms)\n\n# About the kernel:\n- it uses GPU (rapids UMAP is very fast)\n\n- if you don't want to use GPU, you can browse for example this [version](https://www.kaggle.com/hubwag/rapids-umap-with-fisher-metric-on-rgb-histograms?scriptVersionId=36429142) -- the js plot is interactive, except clicking on points doesn't do anything.\n\n- the 3D plots are interactive, hovering over a point shows the related tabular data \n- clicking on a data point shows 10 most similar images + histograms. For efficiency, the similarity search is done in the space of UMAP embedding, but the results seem reasonable.\n- I used slightly trimmed images, because some images have white/black vignettes (more on this below)\n- there seems to be a bug in plotly: custom colors seems to screw up point selections\n\n# Some insights:\n- Initially I worked on (histograms of) original images. This yields two extra clusters containing images with white/black vignettes. The white vignette seems added in postprocessing, the black one may come from the imaging method. There are 109 images with white vignette, 22% of which are malignant -- much higher percentage than normally. \n\n- The enlarged points correspond to target==1. As you can see, there is a big portion of data -- the  'fin' extending to the right of the UMAP plot -- where malignant cases are very rare. More precisely, there are 12345 (sic!) train images in this part, out of which only 13 are malignant. It seems this part corresponds to images with healthy-looking pinkish skin with not-too-dark changes.\n\n- [Updated]: I also checked that the train and test data have similar structure without visible outliers. It'd be interesting to see if the available external data is similar.\n\n- The presence of apparent loops in the UMAP embedding is quite interesting, and it's present for all (~50) random seeds I checked. Rapids fast GPU implementation of UMAP was definitely useful for that. I would speculate that these loops are present in the original space, since UMAP should be good at preserving such topological features.\n\n# Technicalities:\n- I use the Fisher information metric to compute the distance between two histograms. In terms of implementation it boils down to taking the np.sqrt of each histogram -- sqrt is an isometry between the Fisher space and the Euclidean space. So running UMAP on the transformed space is equivalent to running UMAP with Fisher metric as the source metric, except rapids UMAP doesn't seem to support custom metrics. This is quite an important point -- for example using the Euclidean distance to compare histograms fails to capture one of two extra components; also the loops are not as prominently separated.\n\n# Summary:\nI hope someone finds this interesting. I will play with this more -- feel free to let me know if you'd like to see some particular experiment or need more details about the above.",
    "887343": "Hi, super nice visualisation. I like it very much. I did not know it is possible to have callbacks to python in notebooks.\n\nYou write that train and test have similar structure. This is correct, but if you compare test and train, there is a big cluster in test, with not so many corresponding images in train. Have a look at [https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference](https://www.kaggle.com/agentauers/more-rapids-umap-train-test-difference)",
    "887662": "Thanks!\n\nYou're right. I had a closer look at my embedding and there is a similar hot-spot of test images. It's in the lower-left corner my summary image. Here is a sample from this region:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F466630%2F7ddf79b3df7773d108350da7a7468585%2Fnew_type.png?generation=1592248984235844&amp;alt=media)\n\n~80% of the data in this region is test data. I'm curious how many of those images are in the public test set -- and can we really generalize from train data in this region... Do we actually know what's the public-private split?",
    "961870": "I really like the visualization, Great work!",
    "2327681": "Realtime RGB Histogram is here : https://youtu.be/k5rLn7VlAhI"
  },
  "source": "meta"
}