{
  "id": 546547,
  "title": "How to handle features like 'BMI' in linear regression models?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546547",
  "author_name": "Marius Heuser",
  "post_date": "2024-11-16T14:10:37.294000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nwhile working on this problem I stumbled upon a problem/question. How do you implement features like 'BMI', if you make the assumption that both extremes (very high/very low) BMI have a strong influence on the final classification?<br>\nThis is probably not relevant for the models used in this competition, since tree's can just make another split, for example:<br>\nBMI over 25 -&gt; no -&gt; BMI under 15 -&gt; …<br>\nBut for linear models, where the value incluences the model to a certain degree it would neglect one side of the extremes.<br>\nMy thoughts are: <br>\nEither normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)<br>\nOr split the feature into two. One where it goes from very low to normal and the other from normal to very high.<br>\nIs any of these methods also useful for tree models? Does it confuse or help the model? What's the standard approach if it is used for a linear model?<br>\nMaybe it's not a problem and I just made an incorrect assumption ;)<br>\nThanks!</p>",
  "messages": [
    {
      "id": 3047773,
      "postDate": "2024-11-17T06:52:20.267Z",
      "content": "<p>From my experience, I would just cut the distribution of BMI to some certain percentiles (for example 0.01 for the lower part - 0.99 for the upper). In a linear regression model the extremes (outliers) can influence too much the result, and that's not normally what you want in a non kaggle environment (in kaggle you can just try it and test if you get better results)</p>\n<blockquote>\n  <p>Either normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)</p>\n</blockquote>\n<p>This is an assumption you can make about the data, but think that you are changing the distribution (in this case the order the data points are positioned) so this would also affect algorithms based on trees. In a linear regression normally what you would do is apply transformations that look for linearity between the predictors and the target (as you are making this assumption at the moment of using a linear regressor -at least from the coefficients, as for the data you can preprocess it before with non linear transformations-)</p>",
      "rawMarkdown": "From my experience, I would just cut the distribution of BMI to some certain percentiles (for example 0.01 for the lower part - 0.99 for the upper). In a linear regression model the extremes (outliers) can influence too much the result, and that's not normally what you want in a non kaggle environment (in kaggle you can just try it and test if you get better results)\n\n>Either normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)\n\nThis is an assumption you can make about the data, but think that you are changing the distribution (in this case the order the data points are positioned) so this would also affect algorithms based on trees. In a linear regression normally what you would do is apply transformations that look for linearity between the predictors and the target (as you are making this assumption at the moment of using a linear regressor -at least from the coefficients, as for the data you can preprocess it before with non linear transformations-)",
      "votes": 1
    },
    {
      "id": 3062675,
      "postDate": "2024-12-03T20:00:50.597Z",
      "content": "<p>In multiple linear regression, a common way to include/allow a \"double sided\" effect of \\(x_{\\rm BMI}\\) on \\(y\\) is to add another, quadratic feature: \\(x_{BMI2} = (x_{\\rm BMI}-x_{\\rm const})^2\\); the subtracted constant is to help with the fitting and can be any middle-of-xs value. The model can adjust the linear coefficients of these two BMI features for best fit, i.e., the best quadratic relation for the contribution of the x value to y.</p>",
      "rawMarkdown": "In multiple linear regression, a common way to include/allow a \"double sided\" effect of \\\\(x_{\\rm BMI}\\\\) on \\\\(y\\\\) is to add another, quadratic feature: \\\\(x_{BMI2} = (x_{\\rm BMI}-x_{\\rm const})^2\\\\); the subtracted constant is to help with the fitting and can be any middle-of-xs value. The model can adjust the linear coefficients of these two BMI features for best fit, i.e., the best quadratic relation for the contribution of the x value to y."
    },
    {
      "id": 3047295,
      "postDate": "2024-11-16T14:10:37.293Z",
      "content": "<p>Hello everyone,<br>\nwhile working on this problem I stumbled upon a problem/question. How do you implement features like 'BMI', if you make the assumption that both extremes (very high/very low) BMI have a strong influence on the final classification?<br>\nThis is probably not relevant for the models used in this competition, since tree's can just make another split, for example:<br>\nBMI over 25 -&gt; no -&gt; BMI under 15 -&gt; …<br>\nBut for linear models, where the value incluences the model to a certain degree it would neglect one side of the extremes.<br>\nMy thoughts are: <br>\nEither normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)<br>\nOr split the feature into two. One where it goes from very low to normal and the other from normal to very high.<br>\nIs any of these methods also useful for tree models? Does it confuse or help the model? What's the standard approach if it is used for a linear model?<br>\nMaybe it's not a problem and I just made an incorrect assumption ;)<br>\nThanks!</p>",
      "rawMarkdown": "Hello everyone,\nwhile working on this problem I stumbled upon a problem/question. How do you implement features like 'BMI', if you make the assumption that both extremes (very high/very low) BMI have a strong influence on the final classification?\nThis is probably not relevant for the models used in this competition, since tree's can just make another split, for example:\nBMI over 25 -> no -> BMI under 15 -> ...\nBut for linear models, where the value incluences the model to a certain degree it would neglect one side of the extremes.\nMy thoughts are: \nEither normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)\nOr split the feature into two. One where it goes from very low to normal and the other from normal to very high.\nIs any of these methods also useful for tree models? Does it confuse or help the model? What's the standard approach if it is used for a linear model?\nMaybe it's not a problem and I just made an incorrect assumption ;)\nThanks!"
    },
    {
      "id": 3047322,
      "postDate": "2024-11-16T14:51:34.463Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3047773,
      "author_name": "dib",
      "author_url": "",
      "post_date": "2024-11-17T06:52:20.267000",
      "content": "<p>From my experience, I would just cut the distribution of BMI to some certain percentiles (for example 0.01 for the lower part - 0.99 for the upper). In a linear regression model the extremes (outliers) can influence too much the result, and that's not normally what you want in a non kaggle environment (in kaggle you can just try it and test if you get better results)</p>\n<blockquote>\n  <p>Either normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)</p>\n</blockquote>\n<p>This is an assumption you can make about the data, but think that you are changing the distribution (in this case the order the data points are positioned) so this would also affect algorithms based on trees. In a linear regression normally what you would do is apply transformations that look for linearity between the predictors and the target (as you are making this assumption at the moment of using a linear regressor -at least from the coefficients, as for the data you can preprocess it before with non linear transformations-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3062675,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-12-03T20:00:50.597000",
      "content": "<p>In multiple linear regression, a common way to include/allow a \"double sided\" effect of \\(x_{\\rm BMI}\\) on \\(y\\) is to add another, quadratic feature: \\(x_{BMI2} = (x_{\\rm BMI}-x_{\\rm const})^2\\); the subtracted constant is to help with the fitting and can be any middle-of-xs value. The model can adjust the linear coefficients of these two BMI features for best fit, i.e., the best quadratic relation for the contribution of the x value to y.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3047322,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-16T14:51:34.463000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3047773": "From my experience, I would just cut the distribution of BMI to some certain percentiles (for example 0.01 for the lower part - 0.99 for the upper). In a linear regression model the extremes (outliers) can influence too much the result, and that's not normally what you want in a non kaggle environment (in kaggle you can just try it and test if you get better results)\n\n>Either normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)\n\nThis is an assumption you can make about the data, but think that you are changing the distribution (in this case the order the data points are positioned) so this would also affect algorithms based on trees. In a linear regression normally what you would do is apply transformations that look for linearity between the predictors and the target (as you are making this assumption at the moment of using a linear regressor -at least from the coefficients, as for the data you can preprocess it before with non linear transformations-)",
    "3062675": "In multiple linear regression, a common way to include/allow a \"double sided\" effect of \\\\(x_{\\rm BMI}\\\\) on \\\\(y\\\\) is to add another, quadratic feature: \\\\(x_{BMI2} = (x_{\\rm BMI}-x_{\\rm const})^2\\\\); the subtracted constant is to help with the fitting and can be any middle-of-xs value. The model can adjust the linear coefficients of these two BMI features for best fit, i.e., the best quadratic relation for the contribution of the x value to y.",
    "3047295": "Hello everyone,\nwhile working on this problem I stumbled upon a problem/question. How do you implement features like 'BMI', if you make the assumption that both extremes (very high/very low) BMI have a strong influence on the final classification?\nThis is probably not relevant for the models used in this competition, since tree's can just make another split, for example:\nBMI over 25 -> no -> BMI under 15 -> ...\nBut for linear models, where the value incluences the model to a certain degree it would neglect one side of the extremes.\nMy thoughts are: \nEither normalize the values and then mirror the values below '1', so very low and very high values influence the model the same way. (for example 0.5 becomes 1.5)\nOr split the feature into two. One where it goes from very low to normal and the other from normal to very high.\nIs any of these methods also useful for tree models? Does it confuse or help the model? What's the standard approach if it is used for a linear model?\nMaybe it's not a problem and I just made an incorrect assumption ;)\nThanks!",
    "3047322": ""
  }
}