{
  "id": 473954,
  "title": "what are some insights that can be acquired from viewing distribution plots?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/473954",
  "author_name": "",
  "post_date": "2024-02-06T15:49:19.727792800Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Most of the depth=0 numerical variable's distplot seems either left skewed or mostly single bar.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3855eeb1e01c79be6c98c0055fd9aaca%2Fd.PNG?generation=1707234454704674&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3665ca0f47da287b969cfc731cb6dbe5%2F111.PNG?generation=1707234477205115&amp;alt=media\"></p>\n<p>but last one tries to look like gaussian distribution.<br>\nAny reference or idea what these dist plots show to us?<br>\n-pics from my <a href=\"https://www.kaggle.com/code/beckpro/depth-0-complete-eda-with-nice-descriptions\" target=\"_blank\">notebook</a>-</p>",
  "messages": [
    {
      "id": "2638983",
      "postDate": "02/06/2024 15:49:19",
      "content": "<p>Most of the depth=0 numerical variable's distplot seems either left skewed or mostly single bar.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3855eeb1e01c79be6c98c0055fd9aaca%2Fd.PNG?generation=1707234454704674&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3665ca0f47da287b969cfc731cb6dbe5%2F111.PNG?generation=1707234477205115&amp;alt=media\"></p>\n<p>but last one tries to look like gaussian distribution.<br>\nAny reference or idea what these dist plots show to us?<br>\n-pics from my <a href=\"https://www.kaggle.com/code/beckpro/depth-0-complete-eda-with-nice-descriptions\" target=\"_blank\">notebook</a>-</p>",
      "rawMarkdown": "Most of the depth=0 numerical variable's distplot seems either left skewed or mostly single bar.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3855eeb1e01c79be6c98c0055fd9aaca%2Fd.PNG?generation=1707234454704674&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3665ca0f47da287b969cfc731cb6dbe5%2F111.PNG?generation=1707234477205115&alt=media)\n\nbut last one tries to look like gaussian distribution.\nAny reference or idea what these dist plots show to us?\n-pics from my [notebook](https://www.kaggle.com/code/beckpro/depth-0-complete-eda-with-nice-descriptions)-",
      "votes": null
    },
    {
      "id": "2639056",
      "postDate": "02/06/2024 16:52:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a> ! I think you have to deal with the outliers in this situation. First create boxplots and then use Isolation Forest. </p>",
      "rawMarkdown": "Hi @beckpro ! I think you have to deal with the outliers in this situation. First create boxplots and then use Isolation Forest.",
      "votes": null
    },
    {
      "id": "2639121",
      "postDate": "02/06/2024 17:15:25",
      "content": "<p><a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a>, I remembered 2 other options that might help with these distributions. This is the Box-Cox method <a href=\"https://www.statology.org/box-cox-transformation-python/\" target=\"_blank\">https://www.statology.org/box-cox-transformation-python/</a> and as an option you can also try binning, i.e. dividing the feature into several categories. For this you can use the optbinning library <a href=\"https://gnpalencia.org/optbinning/tutorials/tutorial_binary.html\" target=\"_blank\">https://gnpalencia.org/optbinning/tutorials/tutorial_binary.html</a> . I mainly use optbinning in my work when I build a model based on logistic regression.</p>",
      "rawMarkdown": "beckpro, I remembered 2 other options that might help with these distributions. This is the Box-Cox method https://www.statology.org/box-cox-transformation-python/ and as an option you can also try binning, i.e. dividing the feature into several categories. For this you can use the optbinning library https://gnpalencia.org/optbinning/tutorials/tutorial_binary.html . I mainly use optbinning in my work when I build a model based on logistic regression.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2639056,
      "author_name": "bratkovskyevgeny",
      "author_url": "",
      "post_date": "02/06/2024 16:52:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a> ! I think you have to deal with the outliers in this situation. First create boxplots and then use Isolation Forest. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2639121,
      "author_name": "bratkovskyevgeny",
      "author_url": "",
      "post_date": "02/06/2024 17:15:25",
      "content": "<p><a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a>, I remembered 2 other options that might help with these distributions. This is the Box-Cox method <a href=\"https://www.statology.org/box-cox-transformation-python/\" target=\"_blank\">https://www.statology.org/box-cox-transformation-python/</a> and as an option you can also try binning, i.e. dividing the feature into several categories. For this you can use the optbinning library <a href=\"https://gnpalencia.org/optbinning/tutorials/tutorial_binary.html\" target=\"_blank\">https://gnpalencia.org/optbinning/tutorials/tutorial_binary.html</a> . I mainly use optbinning in my work when I build a model based on logistic regression.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2638983": "Most of the depth=0 numerical variable's distplot seems either left skewed or mostly single bar.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3855eeb1e01c79be6c98c0055fd9aaca%2Fd.PNG?generation=1707234454704674&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4921379%2F3665ca0f47da287b969cfc731cb6dbe5%2F111.PNG?generation=1707234477205115&alt=media)\n\nbut last one tries to look like gaussian distribution.\nAny reference or idea what these dist plots show to us?\n-pics from my [notebook](https://www.kaggle.com/code/beckpro/depth-0-complete-eda-with-nice-descriptions)-",
    "2639056": "Hi @beckpro ! I think you have to deal with the outliers in this situation. First create boxplots and then use Isolation Forest.",
    "2639121": "beckpro, I remembered 2 other options that might help with these distributions. This is the Box-Cox method https://www.statology.org/box-cox-transformation-python/ and as an option you can also try binning, i.e. dividing the feature into several categories. For this you can use the optbinning library https://gnpalencia.org/optbinning/tutorials/tutorial_binary.html . I mainly use optbinning in my work when I build a model based on logistic regression."
  },
  "source": "meta"
}