{
  "id": 497934,
  "title": "Help me! Not related to this competition question! An interesting question about feature engineering",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/497934",
  "author_name": "humbleyll",
  "post_date": "2024-04-26T09:28:21.897000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Help me! Not related to this competition question! An interesting question about feature engineering:</p>\n<ol>\n<li><p>How is it more appropriate to handle continuous data with a very large range of values for certain features (ranging from one to several hundred million in numerical terms)?</p></li>\n<li><p>Similarly, if the value range of a feature ranges from a few tenths to one billionth, what should be done?</p></li>\n</ol>\n<p>Is there any big shot who can help analyze this problem? appreciate😀</p>",
  "messages": [
    {
      "id": 2789236,
      "postDate": "2024-05-02T16:04:17.527Z",
      "content": "<p>As an option: sklearn.preprocessing.QuantileTransformer</p>",
      "rawMarkdown": "As an option: sklearn.preprocessing.QuantileTransformer",
      "votes": 1
    },
    {
      "id": 2777930,
      "postDate": "2024-04-26T21:31:54.237Z",
      "content": "<p>It is also important to consider in what context this data is being used, also the domain expertise, few of the options are, binning, logarithmic transformation or scaling techniques in sklearn. </p>",
      "rawMarkdown": "It is also important to consider in what context this data is being used, also the domain expertise, few of the options are, binning, logarithmic transformation or scaling techniques in sklearn. ",
      "votes": 1
    },
    {
      "id": 2776845,
      "postDate": "2024-04-26T11:25:59.593Z",
      "content": "<p>You can make classes based on percentiles, for example pandas qcut, or you can define your own intervals and use pandas cut. The choice of these ranges/intervals depend on the relation the variable holds with the response variable. e.g., for linear relation you can cut the range on equal intervals, for exponential relation the interval should also scale exponentially.</p>",
      "rawMarkdown": "You can make classes based on percentiles, for example pandas qcut, or you can define your own intervals and use pandas cut. The choice of these ranges/intervals depend on the relation the variable holds with the response variable. e.g., for linear relation you can cut the range on equal intervals, for exponential relation the interval should also scale exponentially.",
      "votes": 1,
      "replies": [
        {
          "id": 2776948,
          "postDate": "2024-04-26T12:22:44.077Z",
          "content": "<p>Thank you for your reply! Your viewpoint is good, considering the linear or logarithmic relationship of the original features, and then performing box splitting or logarithmic scaling.</p>\n<p>But the new problem is: as described in question 2, if the value of the feature is very small, such that after taking the logarithm, it exceeds the data types supported by memory, such as floating-point and double precision floating-point, what should be done? Can the processed data be directly input into the neural network for calculation? Or what are the pre operations?</p>",
          "rawMarkdown": "Thank you for your reply! Your viewpoint is good, considering the linear or logarithmic relationship of the original features, and then performing box splitting or logarithmic scaling.\n\nBut the new problem is: as described in question 2, if the value of the feature is very small, such that after taking the logarithm, it exceeds the data types supported by memory, such as floating-point and double precision floating-point, what should be done? Can the processed data be directly input into the neural network for calculation? Or what are the pre operations?",
          "replies": [
            {
              "id": 2777036,
              "postDate": "2024-04-26T13:26:29.623Z",
              "content": "<p>If I understand properly, the point 1 and 2 are quite similar as both suffer from very high standard deviation. Such data can not be effectively scaled or standardized without loosing information. If you try to place all of these points on a zero mean and unit variance, then loss of information is certain, as the relative difference between points will become so small.</p>",
              "rawMarkdown": "If I understand properly, the point 1 and 2 are quite similar as both suffer from very high standard deviation. Such data can not be effectively scaled or standardized without loosing information. If you try to place all of these points on a zero mean and unit variance, then loss of information is certain, as the relative difference between points will become so small.",
              "votes": 1
            },
            {
              "id": 2777099,
              "postDate": "2024-04-26T14:14:32.290Z",
              "content": "<p>Well, there is something wrong with my expression just now, and your understanding is correct. An intuitive idea is to enlarge the very, very small values in the feature by the same multiple. Without changing the trend of the feature data.</p>",
              "rawMarkdown": "Well, there is something wrong with my expression just now, and your understanding is correct. An intuitive idea is to enlarge the very, very small values in the feature by the same multiple. Without changing the trend of the feature data."
            }
          ]
        }
      ]
    },
    {
      "id": 2776762,
      "postDate": "2024-04-26T10:19:22.100Z",
      "content": "<p>How about performing logarithmic or exponential operations？np.log(),np.exp()</p>",
      "rawMarkdown": "How about performing logarithmic or exponential operations？np.log(),np.exp()",
      "votes": 1,
      "replies": [
        {
          "id": 2776944,
          "postDate": "2024-04-26T12:15:46.570Z",
          "content": "<p>Hi, Yes, I also thought of this point and used logarithms to reduce it. A number with a value of 100 million does indeed decrease after taking the logarithm, and that's not wrong.</p>\n<p>However, the new problem is that there are both numbers with a value of 100 million and numbers with a value of 1 in this interval. What I mean is that if all the numbers in the interval are operated uniformly according to the logarithm, the originally small numbers may become very, very small, so that the Float and Double data types cannot handle them</p>",
          "rawMarkdown": "Hi, Yes, I also thought of this point and used logarithms to reduce it. A number with a value of 100 million does indeed decrease after taking the logarithm, and that's not wrong.\n\nHowever, the new problem is that there are both numbers with a value of 100 million and numbers with a value of 1 in this interval. What I mean is that if all the numbers in the interval are operated uniformly according to the logarithm, the originally small numbers may become very, very small, so that the Float and Double data types cannot handle them"
        }
      ]
    },
    {
      "id": 2776680,
      "postDate": "2024-04-26T09:28:21.897Z",
      "content": "<p>Help me! Not related to this competition question! An interesting question about feature engineering:</p>\n<ol>\n<li><p>How is it more appropriate to handle continuous data with a very large range of values for certain features (ranging from one to several hundred million in numerical terms)?</p></li>\n<li><p>Similarly, if the value range of a feature ranges from a few tenths to one billionth, what should be done?</p></li>\n</ol>\n<p>Is there any big shot who can help analyze this problem? appreciate😀</p>",
      "rawMarkdown": "Help me! Not related to this competition question! An interesting question about feature engineering:\n1. How is it more appropriate to handle continuous data with a very large range of values for certain features (ranging from one to several hundred million in numerical terms)?\n\n2. Similarly, if the value range of a feature ranges from a few tenths to one billionth, what should be done?\n\nIs there any big shot who can help analyze this problem? appreciate😀",
      "votes": 2
    },
    {
      "id": 2776947,
      "postDate": "2024-04-26T12:22:27.423Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2789236,
      "author_name": "Andrey Nesterov",
      "author_url": "",
      "post_date": "2024-05-02T16:04:17.527000",
      "content": "<p>As an option: sklearn.preprocessing.QuantileTransformer</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2777930,
      "author_name": "Shreyas Bhatt",
      "author_url": "",
      "post_date": "2024-04-26T21:31:54.237000",
      "content": "<p>It is also important to consider in what context this data is being used, also the domain expertise, few of the options are, binning, logarithmic transformation or scaling techniques in sklearn. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2776845,
      "author_name": "Zam",
      "author_url": "",
      "post_date": "2024-04-26T11:25:59.593000",
      "content": "<p>You can make classes based on percentiles, for example pandas qcut, or you can define your own intervals and use pandas cut. The choice of these ranges/intervals depend on the relation the variable holds with the response variable. e.g., for linear relation you can cut the range on equal intervals, for exponential relation the interval should also scale exponentially.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2776948,
          "author_name": "humbleyll",
          "author_url": "",
          "post_date": "2024-04-26T12:22:44.077000",
          "content": "<p>Thank you for your reply! Your viewpoint is good, considering the linear or logarithmic relationship of the original features, and then performing box splitting or logarithmic scaling.</p>\n<p>But the new problem is: as described in question 2, if the value of the feature is very small, such that after taking the logarithm, it exceeds the data types supported by memory, such as floating-point and double precision floating-point, what should be done? Can the processed data be directly input into the neural network for calculation? Or what are the pre operations?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2777036,
              "author_name": "Zam",
              "author_url": "",
              "post_date": "2024-04-26T13:26:29.623000",
              "content": "<p>If I understand properly, the point 1 and 2 are quite similar as both suffer from very high standard deviation. Such data can not be effectively scaled or standardized without loosing information. If you try to place all of these points on a zero mean and unit variance, then loss of information is certain, as the relative difference between points will become so small.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2777099,
              "author_name": "humbleyll",
              "author_url": "",
              "post_date": "2024-04-26T14:14:32.290000",
              "content": "<p>Well, there is something wrong with my expression just now, and your understanding is correct. An intuitive idea is to enlarge the very, very small values in the feature by the same multiple. Without changing the trend of the feature data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2776762,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-04-26T10:19:22.100000",
      "content": "<p>How about performing logarithmic or exponential operations？np.log(),np.exp()</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2776944,
          "author_name": "humbleyll",
          "author_url": "",
          "post_date": "2024-04-26T12:15:46.570000",
          "content": "<p>Hi, Yes, I also thought of this point and used logarithms to reduce it. A number with a value of 100 million does indeed decrease after taking the logarithm, and that's not wrong.</p>\n<p>However, the new problem is that there are both numbers with a value of 100 million and numbers with a value of 1 in this interval. What I mean is that if all the numbers in the interval are operated uniformly according to the logarithm, the originally small numbers may become very, very small, so that the Float and Double data types cannot handle them</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2776947,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-26T12:22:27.423000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2789236": "As an option: sklearn.preprocessing.QuantileTransformer",
    "2777930": "It is also important to consider in what context this data is being used, also the domain expertise, few of the options are, binning, logarithmic transformation or scaling techniques in sklearn. ",
    "2776845": "You can make classes based on percentiles, for example pandas qcut, or you can define your own intervals and use pandas cut. The choice of these ranges/intervals depend on the relation the variable holds with the response variable. e.g., for linear relation you can cut the range on equal intervals, for exponential relation the interval should also scale exponentially.",
    "2776762": "How about performing logarithmic or exponential operations？np.log(),np.exp()",
    "2776680": "Help me! Not related to this competition question! An interesting question about feature engineering:\n1. How is it more appropriate to handle continuous data with a very large range of values for certain features (ranging from one to several hundred million in numerical terms)?\n\n2. Similarly, if the value range of a feature ranges from a few tenths to one billionth, what should be done?\n\nIs there any big shot who can help analyze this problem? appreciate😀",
    "2776947": ""
  }
}