{
  "id": 328084,
  "title": "Advanced EDA - UMAP/Hdbscan",
  "url": "/competitions/amex-default-prediction/discussion/328084",
  "author_name": "Lucas Morin",
  "post_date": "2022-05-30T19:55:03.391000",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<p>From past competitions I've learned that some advanced techniques can produce beautiful results, sometimes even useful. </p>\n<p><strong>I've shared a simple UMAP/Hdbscan notebook for data exploration:</strong> <a href=\"https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration</a></p>\n<p>I think we can see different things:</p>\n<ul>\n<li>There are some distinct clusters. </li>\n<li>The distribution doesn't seems to change much between train and test. </li>\n<li>Unsupervised clusters show very different historical default rate:</li>\n</ul>\n<p><img src=\"https://i.imgur.com/RGa9Nvj.png\" alt=\"embedding\"></p>\n<ul>\n<li>Some features present expected correlation with target</li>\n<li>The Public / Private data (distinction based on date) seems relatively well distributed.</li>\n</ul>\n<p><strong>Let me know if you see something else and don't hesitate to upvote the notebook if you find this beautiful or interesting :)</strong></p>",
  "messages": [
    {
      "id": 1806118,
      "postDate": "2022-05-30T19:55:03.390Z",
      "content": "<p>From past competitions I've learned that some advanced techniques can produce beautiful results, sometimes even useful. </p>\n<p><strong>I've shared a simple UMAP/Hdbscan notebook for data exploration:</strong> <a href=\"https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration</a></p>\n<p>I think we can see different things:</p>\n<ul>\n<li>There are some distinct clusters. </li>\n<li>The distribution doesn't seems to change much between train and test. </li>\n<li>Unsupervised clusters show very different historical default rate:</li>\n</ul>\n<p><img src=\"https://i.imgur.com/RGa9Nvj.png\" alt=\"embedding\"></p>\n<ul>\n<li>Some features present expected correlation with target</li>\n<li>The Public / Private data (distinction based on date) seems relatively well distributed.</li>\n</ul>\n<p><strong>Let me know if you see something else and don't hesitate to upvote the notebook if you find this beautiful or interesting :)</strong></p>",
      "rawMarkdown": "From past competitions I've learned that some advanced techniques can produce beautiful results, sometimes even useful. \n\n**I've shared a simple UMAP/Hdbscan notebook for data exploration:** https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration\n\nI think we can see different things:\n- There are some distinct clusters. \n- The distribution doesn't seems to change much between train and test. \n- Unsupervised clusters show very different historical default rate:\n\n![embedding](https://i.imgur.com/RGa9Nvj.png)\n\n- Some features present expected correlation with target\n- The Public / Private data (distinction based on date) seems relatively well distributed.\n\n**Let me know if you see something else and don't hesitate to upvote the notebook if you find this beautiful or interesting :)**",
      "votes": 13
    },
    {
      "id": 1806141,
      "postDate": "2022-05-30T20:20:06.653Z",
      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a>; I was trying to learn more about the topics, and I came across this post:<br>\n<a href=\"https://pberba.github.io/stats/2020/01/17/hdbscan/\" target=\"_blank\">https://pberba.github.io/stats/2020/01/17/hdbscan/</a><br>\nprobably will be helpful for someone trying to understand Hdbscan in more detail</p>",
      "rawMarkdown": "Thanks for sharing, @lucasmorin; I was trying to learn more about the topics, and I came across this post:\nhttps://pberba.github.io/stats/2020/01/17/hdbscan/\nprobably will be helpful for someone trying to understand Hdbscan in more detail\n",
      "votes": 2,
      "replies": [
        {
          "id": 1806147,
          "postDate": "2022-05-30T20:32:26.890Z",
          "content": "<p>Mmmh I had abandonned the idea of understanding how TDA works after going trough the UMAP paper, and learned to just enjoy the nice pictures. It seems that your link has enough nice pictures for me to get back to it :-)</p>",
          "rawMarkdown": "Mmmh I had abandonned the idea of understanding how TDA works after going trough the UMAP paper, and learned to just enjoy the nice pictures. It seems that your link has enough nice pictures for me to get back to it :-)"
        },
        {
          "id": 1806200,
          "postDate": "2022-05-30T21:48:32.053Z",
          "content": "<p>😂 Once I had to apply DBScan, this seemed similar; I agree that the post has excellent pictures; those are the best for learning.</p>",
          "rawMarkdown": "😂 Once I had to apply DBScan, this seemed similar; I agree that the post has excellent pictures; those are the best for learning."
        }
      ]
    },
    {
      "id": 1814109,
      "postDate": "2022-06-07T14:08:41.320Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1814338,
          "postDate": "2022-06-07T18:47:39.007Z",
          "content": "<p>Yes the color represent average target, with red being higher. And I put everyone's last statement in the UMAP.</p>",
          "rawMarkdown": "Yes the color represent average target, with red being higher. And I put everyone's last statement in the UMAP."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1806141,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-05-30T20:20:06.653000",
      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a>; I was trying to learn more about the topics, and I came across this post:<br>\n<a href=\"https://pberba.github.io/stats/2020/01/17/hdbscan/\" target=\"_blank\">https://pberba.github.io/stats/2020/01/17/hdbscan/</a><br>\nprobably will be helpful for someone trying to understand Hdbscan in more detail</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1806147,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-05-30T20:32:26.890000",
          "content": "<p>Mmmh I had abandonned the idea of understanding how TDA works after going trough the UMAP paper, and learned to just enjoy the nice pictures. It seems that your link has enough nice pictures for me to get back to it :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1806200,
          "author_name": "C4rl05/V",
          "author_url": "",
          "post_date": "2022-05-30T21:48:32.053000",
          "content": "<p>😂 Once I had to apply DBScan, this seemed similar; I agree that the post has excellent pictures; those are the best for learning.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1814109,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-07T14:08:41.320000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1814338,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-06-07T18:47:39.007000",
          "content": "<p>Yes the color represent average target, with red being higher. And I put everyone's last statement in the UMAP.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1806118": "From past competitions I've learned that some advanced techniques can produce beautiful results, sometimes even useful. \n\n**I've shared a simple UMAP/Hdbscan notebook for data exploration:** https://www.kaggle.com/code/lucasmorin/amex-umap-hdbscan-data-exploration\n\nI think we can see different things:\n- There are some distinct clusters. \n- The distribution doesn't seems to change much between train and test. \n- Unsupervised clusters show very different historical default rate:\n\n![embedding](https://i.imgur.com/RGa9Nvj.png)\n\n- Some features present expected correlation with target\n- The Public / Private data (distinction based on date) seems relatively well distributed.\n\n**Let me know if you see something else and don't hesitate to upvote the notebook if you find this beautiful or interesting :)**",
    "1806141": "Thanks for sharing, @lucasmorin; I was trying to learn more about the topics, and I came across this post:\nhttps://pberba.github.io/stats/2020/01/17/hdbscan/\nprobably will be helpful for someone trying to understand Hdbscan in more detail\n",
    "1814109": ""
  }
}