{
  "id": 332446,
  "title": "Understanding Features with Quantiles",
  "url": "/competitions/amex-default-prediction/discussion/332446",
  "author_name": "",
  "post_date": "2022-06-21T17:21:33.560588100Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for valuble insights into the data patterns. <br>\nReference: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>\n<p>I have created a notebook to illustrate the following patterns from the data through quantile graph representation.</p>\n<ol>\n<li>Highly discreete values </li>\n<li>Outlier detection.</li>\n<li>follows gaussian distribution/not</li>\n<li>skewness and kurtosis </li>\n</ol>\n<h5>Quick Insights</h5>\n<p>The values are clipped as follows for the series .<br>\nq01 = np.quantile(s, 0.01)<br>\nq99 = np.quantile(s, 0.99)<br>\nd = q99 - q01<br>\ns = np.clip(s, q01-2<em>d, q99+2</em>d)</p>\n<p><strong>Discreete values</strong><br>\nWe are able to see few quantile graphs to have a sharp elbows in the quantile graph in case of outliers or having discreete values.<br>\nExample: S15 -&gt; looks from the quantile graph to have atleast 7 discreete values.</p>\n<p><strong>Gaussian Distribution</strong><br>\nP_2, D_47  from the Normal QQ plot looks to follow gaussian plot.<br>\nwhere as few other features R_13, R_17 will have high kurtosis</p>\n<p>B_31 have single value  across all the quantiles.</p>\n<p>with the reference from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> comment on R_26 has 6 discreete values when zooming-in. which is not showing clearly from my notebook. May need to remove the outliers and check, since the gap is so small.</p>\n<p>Like R_26, many other features have a big jump in the outliers, that are inhibiting to show the other values like D_94, S_26.</p>\n<p>please let me know if any furthur processing, insights and interpretation which might be interesting to look at.</p>\n<p>link to my notebook <br>\n<a href=\"https://www.kaggle.com/code/narendra/understanding-features-with-quantiles\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "1828438",
      "postDate": "06/21/2022 17:21:33",
      "content": "<p>thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for valuble insights into the data patterns. <br>\nReference: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>\n<p>I have created a notebook to illustrate the following patterns from the data through quantile graph representation.</p>\n<ol>\n<li>Highly discreete values </li>\n<li>Outlier detection.</li>\n<li>follows gaussian distribution/not</li>\n<li>skewness and kurtosis </li>\n</ol>\n<h5>Quick Insights</h5>\n<p>The values are clipped as follows for the series .<br>\nq01 = np.quantile(s, 0.01)<br>\nq99 = np.quantile(s, 0.99)<br>\nd = q99 - q01<br>\ns = np.clip(s, q01-2<em>d, q99+2</em>d)</p>\n<p><strong>Discreete values</strong><br>\nWe are able to see few quantile graphs to have a sharp elbows in the quantile graph in case of outliers or having discreete values.<br>\nExample: S15 -&gt; looks from the quantile graph to have atleast 7 discreete values.</p>\n<p><strong>Gaussian Distribution</strong><br>\nP_2, D_47  from the Normal QQ plot looks to follow gaussian plot.<br>\nwhere as few other features R_13, R_17 will have high kurtosis</p>\n<p>B_31 have single value  across all the quantiles.</p>\n<p>with the reference from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> comment on R_26 has 6 discreete values when zooming-in. which is not showing clearly from my notebook. May need to remove the outliers and check, since the gap is so small.</p>\n<p>Like R_26, many other features have a big jump in the outliers, that are inhibiting to show the other values like D_94, S_26.</p>\n<p>please let me know if any furthur processing, insights and interpretation which might be interesting to look at.</p>\n<p>link to my notebook <br>\n<a href=\"https://www.kaggle.com/code/narendra/understanding-features-with-quantiles\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "thanks to @raddar for valuble insights into the data patterns. \nReference: [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)\n\nI have created a notebook to illustrate the following patterns from the data through quantile graph representation.\n1. Highly discreete values \n2. Outlier detection.\n3. follows gaussian distribution/not\n4. skewness and kurtosis \n\n\n\n##### Quick Insights\n\nThe values are clipped as follows for the series .\nq01 = np.quantile(s, 0.01)\nq99 = np.quantile(s, 0.99)\nd = q99 - q01\ns = np.clip(s, q01-2*d, q99+2*d)\n\n\n**Discreete values**\nWe are able to see few quantile graphs to have a sharp elbows in the quantile graph in case of outliers or having discreete values.\nExample: S15 -> looks from the quantile graph to have atleast 7 discreete values.\n\n**Gaussian Distribution**\nP_2, D_47  from the Normal QQ plot looks to follow gaussian plot.\nwhere as few other features R_13, R_17 will have high kurtosis\n\nB_31 have single value  across all the quantiles.\n\n\n\nwith the reference from @cdeotte comment on R_26 has 6 discreete values when zooming-in. which is not showing clearly from my notebook. May need to remove the outliers and check, since the gap is so small.\n\nLike R_26, many other features have a big jump in the outliers, that are inhibiting to show the other values like D_94, S_26.\n\n\nplease let me know if any furthur processing, insights and interpretation which might be interesting to look at.\n\nlink to my notebook \n[here](https://www.kaggle.com/code/narendra/understanding-features-with-quantiles)",
      "votes": null
    },
    {
      "id": "1828450",
      "postDate": "06/21/2022 17:35:35",
      "content": "<p>Link to the notebook is broken :)</p>",
      "rawMarkdown": "Link to the notebook is broken :)",
      "votes": null
    },
    {
      "id": "1828458",
      "postDate": "06/21/2022 17:55:46",
      "content": "<p>fixed. thnks</p>",
      "rawMarkdown": "fixed. thnks",
      "votes": null
    },
    {
      "id": "1831133",
      "postDate": "06/24/2022 01:26:42",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">doteeee</a>, excellent work! As a statistician, I'm very happy when I see distributions, quantiles and so on.</p>\n<p>Some ideas popped to my mind while I was reading your notebook:</p>\n<ol>\n<li><p>Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x**2 or exp(x) for the same reasons.</p></li>\n<li><p>From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?</p></li>\n<li><p>One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?</p></li>\n</ol>\n<p>Again, excellent analysis. Upvoted!</p>",
      "rawMarkdown": "Hello [doteeee](https://www.kaggle.com/narendra), excellent work! As a statistician, I'm very happy when I see distributions, quantiles and so on.\n\nSome ideas popped to my mind while I was reading your notebook:\n\n1. Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x**2 or exp(x) for the same reasons.\n\n2. From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?\n\n3. One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?\n\nAgain, excellent analysis. Upvoted!",
      "votes": null
    },
    {
      "id": "1832471",
      "postDate": "06/25/2022 04:38:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlosasdesouza\" target=\"_blank\">@carlosasdesouza</a> ,  </p>\n<ol>\n<li><p><strong>Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x^^2 or exp(x) for the same reasons.</strong></p>\n<p>yes, i observe that few variables are deviating from the normal distribution due to the outliers. Along with the above said transaformations we can use quantile transformation and check. these yet to try.</p></li>\n\n\n<li><p><strong>From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?</strong><br>\nApart from S23, the rest looks to be counting variables. outliers effected the visual representation of the graph, you can look at the normal plot and say many of them are not zeros.<br>\nwhen we zoom-in [0.131, 0.14] --&gt; uniform distribution, &lt; 0 -&gt; Left-Skewed.<br>\nThe sharp bends from the graph also represents possible outliers.</p></li>\n</ol>\n<p>3.<strong>One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?</strong></p>\n<p>D_87 --&gt; I think we can remove this variable and check the performance of the model.<br>\n0.926% are missing values. <br>\nOf the missing values 0.751335 (target=0), 0.248665(target=1), which had lesser distingushing capacity, atleast from analysis. <a href=\"https://www.kaggle.com/code/narendra/amex-high-level-analysis\" target=\"_blank\">here</a>.</p>\n<p>B_31 --&gt; some mistake of me representing the values, it actually had 2 values(0,1). we can check from the z-score, the range of values are different.</p>",
      "rawMarkdown": "Hi @carlosasdesouza ,  \n\n1. **Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x^^2 or exp(x) for the same reasons.**\n\n       yes, i observe that few variables are deviating from the normal distribution due to the outliers. Along with the above said transaformations we can use quantile transformation and check. these yet to try.\n\n\n\n\n2. **From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?**\nApart from S23, the rest looks to be counting variables. outliers effected the visual representation of the graph, you can look at the normal plot and say many of them are not zeros.\nwhen we zoom-in [0.131, 0.14] --> uniform distribution, < 0 -> Left-Skewed.\nThe sharp bends from the graph also represents possible outliers.\n\n\n\n\n3.**One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?**\n\n\n D_87 --> I think we can remove this variable and check the performance of the model.\n0.926% are missing values. \nOf the missing values 0.751335 (target=0), 0.248665(target=1), which had lesser distingushing capacity, atleast from analysis. [here](https://www.kaggle.com/code/narendra/amex-high-level-analysis).\n\nB_31 --> some mistake of me representing the values, it actually had 2 values(0,1). we can check from the z-score, the range of values are different.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1828450,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/21/2022 17:35:35",
      "content": "<p>Link to the notebook is broken :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1828458,
          "author_name": "narendra",
          "author_url": "",
          "post_date": "06/21/2022 17:55:46",
          "content": "<p>fixed. thnks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1831133,
      "author_name": "carlosasdesouza",
      "author_url": "",
      "post_date": "06/24/2022 01:26:42",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">doteeee</a>, excellent work! As a statistician, I'm very happy when I see distributions, quantiles and so on.</p>\n<p>Some ideas popped to my mind while I was reading your notebook:</p>\n<ol>\n<li><p>Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x**2 or exp(x) for the same reasons.</p></li>\n<li><p>From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?</p></li>\n<li><p>One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?</p></li>\n</ol>\n<p>Again, excellent analysis. Upvoted!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1832471,
          "author_name": "narendra",
          "author_url": "",
          "post_date": "06/25/2022 04:38:32",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/carlosasdesouza\" target=\"_blank\">@carlosasdesouza</a> ,  </p>\n<ol>\n<li><p><strong>Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x^^2 or exp(x) for the same reasons.</strong></p>\n<p>yes, i observe that few variables are deviating from the normal distribution due to the outliers. Along with the above said transaformations we can use quantile transformation and check. these yet to try.</p></li>\n\n\n<li><p><strong>From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?</strong><br>\nApart from S23, the rest looks to be counting variables. outliers effected the visual representation of the graph, you can look at the normal plot and say many of them are not zeros.<br>\nwhen we zoom-in [0.131, 0.14] --&gt; uniform distribution, &lt; 0 -&gt; Left-Skewed.<br>\nThe sharp bends from the graph also represents possible outliers.</p></li>\n</ol>\n<p>3.<strong>One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?</strong></p>\n<p>D_87 --&gt; I think we can remove this variable and check the performance of the model.<br>\n0.926% are missing values. <br>\nOf the missing values 0.751335 (target=0), 0.248665(target=1), which had lesser distingushing capacity, atleast from analysis. <a href=\"https://www.kaggle.com/code/narendra/amex-high-level-analysis\" target=\"_blank\">here</a>.</p>\n<p>B_31 --&gt; some mistake of me representing the values, it actually had 2 values(0,1). we can check from the z-score, the range of values are different.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1828438": "thanks to @raddar for valuble insights into the data patterns. \nReference: [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)\n\nI have created a notebook to illustrate the following patterns from the data through quantile graph representation.\n1. Highly discreete values \n2. Outlier detection.\n3. follows gaussian distribution/not\n4. skewness and kurtosis \n\n\n\n##### Quick Insights\n\nThe values are clipped as follows for the series .\nq01 = np.quantile(s, 0.01)\nq99 = np.quantile(s, 0.99)\nd = q99 - q01\ns = np.clip(s, q01-2*d, q99+2*d)\n\n\n**Discreete values**\nWe are able to see few quantile graphs to have a sharp elbows in the quantile graph in case of outliers or having discreete values.\nExample: S15 -> looks from the quantile graph to have atleast 7 discreete values.\n\n**Gaussian Distribution**\nP_2, D_47  from the Normal QQ plot looks to follow gaussian plot.\nwhere as few other features R_13, R_17 will have high kurtosis\n\nB_31 have single value  across all the quantiles.\n\n\n\nwith the reference from @cdeotte comment on R_26 has 6 discreete values when zooming-in. which is not showing clearly from my notebook. May need to remove the outliers and check, since the gap is so small.\n\nLike R_26, many other features have a big jump in the outliers, that are inhibiting to show the other values like D_94, S_26.\n\n\nplease let me know if any furthur processing, insights and interpretation which might be interesting to look at.\n\nlink to my notebook \n[here](https://www.kaggle.com/code/narendra/understanding-features-with-quantiles)",
    "1828450": "Link to the notebook is broken :)",
    "1828458": "fixed. thnks",
    "1831133": "Hello [doteeee](https://www.kaggle.com/narendra), excellent work! As a statistician, I'm very happy when I see distributions, quantiles and so on.\n\nSome ideas popped to my mind while I was reading your notebook:\n\n1. Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x**2 or exp(x) for the same reasons.\n\n2. From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?\n\n3. One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?\n\nAgain, excellent analysis. Upvoted!",
    "1832471": "Hi @carlosasdesouza ,  \n\n1. **Should we apply concave transformations (log, sqrt) to several variables ? A lot of them seems to have a quadratic -like shaped, all of them with a very discrepant point (in most cases, this point is 1.0). Applying this transformation,we could , in a way, \"soft\" the distributions, vanquishing the effects of the outliers and skewness. For a very few variables we could make the opposite: apply some transformation like x^^2 or exp(x) for the same reasons.**\n\n       yes, i observe that few variables are deviating from the normal distribution due to the outliers. Along with the above said transaformations we can use quantile transformation and check. these yet to try.\n\n\n\n\n2. **From the plots, it seems that R9, S23, D51, D70, D82, D122, D136, D138 are counting variables, aren't they?**\nApart from S23, the rest looks to be counting variables. outliers effected the visual representation of the graph, you can look at the normal plot and say many of them are not zeros.\nwhen we zoom-in [0.131, 0.14] --> uniform distribution, < 0 -> Left-Skewed.\nThe sharp bends from the graph also represents possible outliers.\n\n\n\n\n3.**One could say we drop B_31, D_87 for having an unique value across all the quantiles but, if we fill the missing values of this variable, the distribution of the target change for missing/non-missing values ? Have you already done this analysis ?**\n\n\n D_87 --> I think we can remove this variable and check the performance of the model.\n0.926% are missing values. \nOf the missing values 0.751335 (target=0), 0.248665(target=1), which had lesser distingushing capacity, atleast from analysis. [here](https://www.kaggle.com/code/narendra/amex-high-level-analysis).\n\nB_31 --> some mistake of me representing the values, it actually had 2 values(0,1). we can check from the z-score, the range of values are different."
  },
  "source": "meta"
}