{
  "id": 337920,
  "title": "Is the dataset normalized?",
  "url": "/competitions/amex-default-prediction/discussion/337920",
  "author_name": "",
  "post_date": "2022-07-18T09:38:23.674920400Z",
  "votes": 1,
  "comment_count": 14,
  "views": 0,
  "content": "<p>In data description, the dataset's features are already normalized.<br>\nBut when I print them, there values are not between 0~1 or -1~1, and their variances are not around 1. Some feature's variance is too small (0.00036..) or too big (65.3…). <br>\nI am a student who has not yet received a master's degree, so I wonder that I misunderstand it. <br>\nThank You.</p>",
  "messages": [
    {
      "id": "1860381",
      "postDate": "07/18/2022 09:38:23",
      "content": "<p>In data description, the dataset's features are already normalized.<br>\nBut when I print them, there values are not between 0~1 or -1~1, and their variances are not around 1. Some feature's variance is too small (0.00036..) or too big (65.3…). <br>\nI am a student who has not yet received a master's degree, so I wonder that I misunderstand it. <br>\nThank You.</p>",
      "rawMarkdown": "In data description, the dataset's features are already normalized.\nBut when I print them, there values are not between 0~1 or -1~1, and their variances are not around 1. Some feature's variance is too small (0.00036..) or too big (65.3...). \nI am a student who has not yet received a master's degree, so I wonder that I misunderstand it. \nThank You.",
      "votes": null
    },
    {
      "id": "1860463",
      "postDate": "07/18/2022 10:40:39",
      "content": "<p>From the data decription page: <code>Features are anonymized and normalized ...</code></p>",
      "rawMarkdown": "From the data decription page: `Features are anonymized and normalized ...`",
      "votes": null
    },
    {
      "id": "1860489",
      "postDate": "07/18/2022 11:03:40",
      "content": "<p>Yes, I know that.<br>\nBut, I learned that normalization means adjusting to a value of 0 mean and 1 variance, but the provided dataset doesn't fit.</p>",
      "rawMarkdown": "Yes, I know that.\nBut, I learned that normalization means adjusting to a value of 0 mean and 1 variance, but the provided dataset doesn't fit.",
      "votes": null
    },
    {
      "id": "1860559",
      "postDate": "07/18/2022 11:39:08",
      "content": "<p>maybe it does not mean all features are normalized. Some features need to denormalize first before feature engineering. One example is \"D_141\".  df['D_141'] = df['D_141'] - mean_norm(df['D_141']) where mean_norm = (col - col.mean()) / (col.max() - col.min())</p>",
      "rawMarkdown": "maybe it does not mean all features are normalized. Some features need to denormalize first before feature engineering. One example is \"D_141\".  df['D_141'] = df['D_141'] - mean_norm(df['D_141']) where mean_norm = (col - col.mean()) / (col.max() - col.min())",
      "votes": null
    },
    {
      "id": "1860566",
      "postDate": "07/18/2022 11:43:21",
      "content": "<p>Why do we need denormalization? <br>\nShouldn't we normalize the features which are not normalized?</p>",
      "rawMarkdown": "Why do we need denormalization? \nShouldn't we normalize the features which are not normalized?",
      "votes": null
    },
    {
      "id": "1860578",
      "postDate": "07/18/2022 11:56:20",
      "content": "<p>Example of Anonymization Algorithm:<br>\nStage 1: Data Normalization<br>\nStage 2: Hiding (\"masking\") the class labels<br>\nStage 3: Hiding the order of data samples<br>\nStage 4: Hiding the order of dimensions (data attributes)<br>\nStage 5: Homeomorphic data space transformation<br>\nStage 6: Adding secret bias. <br>\nStage 7: Applying activation function</p>",
      "rawMarkdown": "Example of Anonymization Algorithm:\nStage 1: Data Normalization\nStage 2: Hiding (\"masking\") the class labels\nStage 3: Hiding the order of data samples\nStage 4: Hiding the order of dimensions (data attributes)\nStage 5: Homeomorphic data space transformation\nStage 6: Adding secret bias. \nStage 7: Applying activation function",
      "votes": null
    },
    {
      "id": "1860601",
      "postDate": "07/18/2022 12:14:42",
      "content": "<p>There are many many ways to normalize the data, i.e.:</p>\n<ul>\n<li>standardization (mean=0,std=1)</li>\n<li>min-max normalization</li>\n<li>box-cox transformation</li>\n<li>quantile normalization</li>\n<li>rank normalization</li>\n<li>…</li>\n</ul>\n<p>you are probably only familiar with the first one :)</p>\n<p>The organizers may have applied different normalization techniques to different columns.</p>",
      "rawMarkdown": "There are many many ways to normalize the data, i.e.:\n\n- standardization (mean=0,std=1)\n- min-max normalization\n- box-cox transformation\n- quantile normalization\n- rank normalization\n- ...\n\nyou are probably only familiar with the first one :)\n\nThe organizers may have applied different normalization techniques to different columns.",
      "votes": null
    },
    {
      "id": "1860623",
      "postDate": "07/18/2022 12:32:08",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> you are correct. most observable normalizations are standardization, mean-norm, and max abs:</p>\n<p>delinquency<br>\n    28 max abs<br>\n    35 mean norm<br>\n    24 standardization<br>\nspend<br>\n    1 max abs<br>\n    9 mean norm<br>\n    11 standardization<br>\npayment<br>\n    1 mean norm<br>\n    2 standardization<br>\nbalance<br>\n    4 max abs<br>\n    12 mean norm<br>\n    22 standardization<br>\nrisk<br>\n    3 max abs<br>\n    18 mean norm<br>\n    6 standardization</p>",
      "rawMarkdown": "raddar you are correct. most observable normalizations are standardization, mean-norm, and max abs:\n\ndelinquency\n\t28 max abs\n\t35 mean norm\n\t24 standardization\nspend\n\t1 max abs\n\t9 mean norm\n\t11 standardization\npayment\n\t1 mean norm\n\t2 standardization\nbalance\n\t4 max abs\n\t12 mean norm\n\t22 standardization\nrisk\n\t3 max abs\n\t18 mean norm\n\t6 standardization",
      "votes": null
    },
    {
      "id": "1860690",
      "postDate": "07/18/2022 13:13:47",
      "content": "<p><a href=\"https://www.kaggle.com/FGPC\" target=\"_blank\">@FGPC</a>, <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a><br>\nThank you for your answers! But I have a few more questions.<br>\nTo sum it up, there are many kinds of normalization, and anonymization includes normalization, and different normalization techniques have applied to different columns except categorical data.<br>\nBut I'm still confused about denormalization…..<br>\nAre there any negative effects when using anonymized and normalized features to learn our models? If not, I think there is no need to denormalize features.<br>\nAnd how can I specify which normalization was used for which column? Should I use statistics like mean, std? </p>\n<p>Since this is my first time participating in a kaggle competition, your answers are very helpful. Sorry for my poor English.</p>",
      "rawMarkdown": "FGPC, @raddar\nThank you for your answers! But I have a few more questions.\nTo sum it up, there are many kinds of normalization, and anonymization includes normalization, and different normalization techniques have applied to different columns except categorical data.\nBut I'm still confused about denormalization.....\nAre there any negative effects when using anonymized and normalized features to learn our models? If not, I think there is no need to denormalize features.\nAnd how can I specify which normalization was used for which column? Should I use statistics like mean, std? \n\nSince this is my first time participating in a kaggle competition, your answers are very helpful. Sorry for my poor English.",
      "votes": null
    },
    {
      "id": "1860707",
      "postDate": "07/18/2022 13:25:16",
      "content": "<p>You are completely right - denormalization is not necessary. However for this competition, denormalization allowed us to remove artificial noise added to integer features.</p>",
      "rawMarkdown": "You are completely right - denormalization is not necessary. However for this competition, denormalization allowed us to remove artificial noise added to integer features.",
      "votes": null
    },
    {
      "id": "1861316",
      "postDate": "07/19/2022 00:59:10",
      "content": "<p>本次比赛主要使用树模型，暂时不需要对数据进行统一的标准化或规范化等处理，正如@raddar 说的一样，有些特征是整数列，进行相应的处理，可以获得一些先验性知识，让模型能够学到更多的差异化内容。但是，如果你使用NN模型，需要考虑进行相应的处理。</p>",
      "rawMarkdown": "本次比赛主要使用树模型，暂时不需要对数据进行统一的标准化或规范化等处理，正如@raddar 说的一样，有些特征是整数列，进行相应的处理，可以获得一些先验性知识，让模型能够学到更多的差异化内容。但是，如果你使用NN模型，需要考虑进行相应的处理。",
      "votes": null
    },
    {
      "id": "1861646",
      "postDate": "07/19/2022 06:28:32",
      "content": "<p>Does that mean tree models don't need to standardize or normalize uniformly? I've never used tree models for machine learning, I just know what they are. (for example, Light GBM, random forest, etc….)</p>",
      "rawMarkdown": "Does that mean tree models don't need to standardize or normalize uniformly? I've never used tree models for machine learning, I just know what they are. (for example, Light GBM, random forest, etc....)",
      "votes": null
    },
    {
      "id": "1861652",
      "postDate": "07/19/2022 06:32:41",
      "content": "<p>tree methods are robust to any input. that's why they rock :)</p>",
      "rawMarkdown": "tree methods are robust to any input. that's why they rock :)",
      "votes": null
    },
    {
      "id": "1861655",
      "postDate": "07/19/2022 06:40:00",
      "content": "<p>Thank you for your answers! I understand clearly.</p>",
      "rawMarkdown": "Thank you for your answers! I understand clearly.",
      "votes": null
    },
    {
      "id": "1871991",
      "postDate": "07/26/2022 15:33:45",
      "content": "<p>i think there is probably no need to normalize the data for neural networks (i tried it and the diff is negligible) and there definitely is no need for tree based models. but if you are planning to use distance based algorithms (kmeans / pca / svm), then i think standardization will help. </p>",
      "rawMarkdown": "i think there is probably no need to normalize the data for neural networks (i tried it and the diff is negligible) and there definitely is no need for tree based models. but if you are planning to use distance based algorithms (kmeans / pca / svm), then i think standardization will help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1860463,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "07/18/2022 10:40:39",
      "content": "<p>From the data decription page: <code>Features are anonymized and normalized ...</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1860489,
          "author_name": "jangyoungchan",
          "author_url": "",
          "post_date": "07/18/2022 11:03:40",
          "content": "<p>Yes, I know that.<br>\nBut, I learned that normalization means adjusting to a value of 0 mean and 1 variance, but the provided dataset doesn't fit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860559,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "07/18/2022 11:39:08",
          "content": "<p>maybe it does not mean all features are normalized. Some features need to denormalize first before feature engineering. One example is \"D_141\".  df['D_141'] = df['D_141'] - mean_norm(df['D_141']) where mean_norm = (col - col.mean()) / (col.max() - col.min())</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860566,
          "author_name": "jangyoungchan",
          "author_url": "",
          "post_date": "07/18/2022 11:43:21",
          "content": "<p>Why do we need denormalization? <br>\nShouldn't we normalize the features which are not normalized?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860578,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "07/18/2022 11:56:20",
          "content": "<p>Example of Anonymization Algorithm:<br>\nStage 1: Data Normalization<br>\nStage 2: Hiding (\"masking\") the class labels<br>\nStage 3: Hiding the order of data samples<br>\nStage 4: Hiding the order of dimensions (data attributes)<br>\nStage 5: Homeomorphic data space transformation<br>\nStage 6: Adding secret bias. <br>\nStage 7: Applying activation function</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860601,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "07/18/2022 12:14:42",
          "content": "<p>There are many many ways to normalize the data, i.e.:</p>\n<ul>\n<li>standardization (mean=0,std=1)</li>\n<li>min-max normalization</li>\n<li>box-cox transformation</li>\n<li>quantile normalization</li>\n<li>rank normalization</li>\n<li>…</li>\n</ul>\n<p>you are probably only familiar with the first one :)</p>\n<p>The organizers may have applied different normalization techniques to different columns.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860623,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "07/18/2022 12:32:08",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> you are correct. most observable normalizations are standardization, mean-norm, and max abs:</p>\n<p>delinquency<br>\n    28 max abs<br>\n    35 mean norm<br>\n    24 standardization<br>\nspend<br>\n    1 max abs<br>\n    9 mean norm<br>\n    11 standardization<br>\npayment<br>\n    1 mean norm<br>\n    2 standardization<br>\nbalance<br>\n    4 max abs<br>\n    12 mean norm<br>\n    22 standardization<br>\nrisk<br>\n    3 max abs<br>\n    18 mean norm<br>\n    6 standardization</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860690,
          "author_name": "jangyoungchan",
          "author_url": "",
          "post_date": "07/18/2022 13:13:47",
          "content": "<p><a href=\"https://www.kaggle.com/FGPC\" target=\"_blank\">@FGPC</a>, <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a><br>\nThank you for your answers! But I have a few more questions.<br>\nTo sum it up, there are many kinds of normalization, and anonymization includes normalization, and different normalization techniques have applied to different columns except categorical data.<br>\nBut I'm still confused about denormalization…..<br>\nAre there any negative effects when using anonymized and normalized features to learn our models? If not, I think there is no need to denormalize features.<br>\nAnd how can I specify which normalization was used for which column? Should I use statistics like mean, std? </p>\n<p>Since this is my first time participating in a kaggle competition, your answers are very helpful. Sorry for my poor English.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860707,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "07/18/2022 13:25:16",
          "content": "<p>You are completely right - denormalization is not necessary. However for this competition, denormalization allowed us to remove artificial noise added to integer features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1861316,
          "author_name": "kgxiao",
          "author_url": "",
          "post_date": "07/19/2022 00:59:10",
          "content": "<p>本次比赛主要使用树模型，暂时不需要对数据进行统一的标准化或规范化等处理，正如@raddar 说的一样，有些特征是整数列，进行相应的处理，可以获得一些先验性知识，让模型能够学到更多的差异化内容。但是，如果你使用NN模型，需要考虑进行相应的处理。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1861646,
          "author_name": "jangyoungchan",
          "author_url": "",
          "post_date": "07/19/2022 06:28:32",
          "content": "<p>Does that mean tree models don't need to standardize or normalize uniformly? I've never used tree models for machine learning, I just know what they are. (for example, Light GBM, random forest, etc….)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1861652,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "07/19/2022 06:32:41",
          "content": "<p>tree methods are robust to any input. that's why they rock :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1861655,
          "author_name": "jangyoungchan",
          "author_url": "",
          "post_date": "07/19/2022 06:40:00",
          "content": "<p>Thank you for your answers! I understand clearly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1871991,
      "author_name": "evanning",
      "author_url": "",
      "post_date": "07/26/2022 15:33:45",
      "content": "<p>i think there is probably no need to normalize the data for neural networks (i tried it and the diff is negligible) and there definitely is no need for tree based models. but if you are planning to use distance based algorithms (kmeans / pca / svm), then i think standardization will help. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1860381": "In data description, the dataset's features are already normalized.\nBut when I print them, there values are not between 0~1 or -1~1, and their variances are not around 1. Some feature's variance is too small (0.00036..) or too big (65.3...). \nI am a student who has not yet received a master's degree, so I wonder that I misunderstand it. \nThank You.",
    "1860463": "From the data decription page: `Features are anonymized and normalized ...`",
    "1860489": "Yes, I know that.\nBut, I learned that normalization means adjusting to a value of 0 mean and 1 variance, but the provided dataset doesn't fit.",
    "1860559": "maybe it does not mean all features are normalized. Some features need to denormalize first before feature engineering. One example is \"D_141\".  df['D_141'] = df['D_141'] - mean_norm(df['D_141']) where mean_norm = (col - col.mean()) / (col.max() - col.min())",
    "1860566": "Why do we need denormalization? \nShouldn't we normalize the features which are not normalized?",
    "1860578": "Example of Anonymization Algorithm:\nStage 1: Data Normalization\nStage 2: Hiding (\"masking\") the class labels\nStage 3: Hiding the order of data samples\nStage 4: Hiding the order of dimensions (data attributes)\nStage 5: Homeomorphic data space transformation\nStage 6: Adding secret bias. \nStage 7: Applying activation function",
    "1860601": "There are many many ways to normalize the data, i.e.:\n\n- standardization (mean=0,std=1)\n- min-max normalization\n- box-cox transformation\n- quantile normalization\n- rank normalization\n- ...\n\nyou are probably only familiar with the first one :)\n\nThe organizers may have applied different normalization techniques to different columns.",
    "1860623": "raddar you are correct. most observable normalizations are standardization, mean-norm, and max abs:\n\ndelinquency\n\t28 max abs\n\t35 mean norm\n\t24 standardization\nspend\n\t1 max abs\n\t9 mean norm\n\t11 standardization\npayment\n\t1 mean norm\n\t2 standardization\nbalance\n\t4 max abs\n\t12 mean norm\n\t22 standardization\nrisk\n\t3 max abs\n\t18 mean norm\n\t6 standardization",
    "1860690": "FGPC, @raddar\nThank you for your answers! But I have a few more questions.\nTo sum it up, there are many kinds of normalization, and anonymization includes normalization, and different normalization techniques have applied to different columns except categorical data.\nBut I'm still confused about denormalization.....\nAre there any negative effects when using anonymized and normalized features to learn our models? If not, I think there is no need to denormalize features.\nAnd how can I specify which normalization was used for which column? Should I use statistics like mean, std? \n\nSince this is my first time participating in a kaggle competition, your answers are very helpful. Sorry for my poor English.",
    "1860707": "You are completely right - denormalization is not necessary. However for this competition, denormalization allowed us to remove artificial noise added to integer features.",
    "1861316": "本次比赛主要使用树模型，暂时不需要对数据进行统一的标准化或规范化等处理，正如@raddar 说的一样，有些特征是整数列，进行相应的处理，可以获得一些先验性知识，让模型能够学到更多的差异化内容。但是，如果你使用NN模型，需要考虑进行相应的处理。",
    "1861646": "Does that mean tree models don't need to standardize or normalize uniformly? I've never used tree models for machine learning, I just know what they are. (for example, Light GBM, random forest, etc....)",
    "1861652": "tree methods are robust to any input. that's why they rock :)",
    "1861655": "Thank you for your answers! I understand clearly.",
    "1871991": "i think there is probably no need to normalize the data for neural networks (i tried it and the diff is negligible) and there definitely is no need for tree based models. but if you are planning to use distance based algorithms (kmeans / pca / svm), then i think standardization will help."
  },
  "source": "meta"
}