{
  "id": 193141,
  "title": "INGV - Working with tabular data - normalization.",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/193141",
  "author_name": "",
  "post_date": "2020-10-25T13:57:26.491107900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi I just started to getting along with INGV - Volcanic Eruption Prediction challenge, and I've got quick question. </p>\n<p>Is it possible to normalize all inputs in train data (sensor_1,sensor_2 etc.) to be distributed from [-1, to 1]?  and If so It will have any positive or negative effect on build models and predictions?</p>",
  "messages": [
    {
      "id": "1059823",
      "postDate": "10/25/2020 13:57:26",
      "content": "<p>Hi I just started to getting along with INGV - Volcanic Eruption Prediction challenge, and I've got quick question. </p>\n<p>Is it possible to normalize all inputs in train data (sensor_1,sensor_2 etc.) to be distributed from [-1, to 1]?  and If so It will have any positive or negative effect on build models and predictions?</p>",
      "rawMarkdown": "Hi I just started to getting along with INGV - Volcanic Eruption Prediction challenge, and I've got quick question. \n\nIs it possible to normalize all inputs in train data (sensor_1,sensor_2 etc.) to be distributed from [-1, to 1]?  and If so It will have any positive or negative effect on build models and predictions?",
      "votes": null
    },
    {
      "id": "1128263",
      "postDate": "12/27/2020 09:31:46",
      "content": "<p>There are several ways you can do such normalization : </p>\n<ul>\n<li>Normalize the data of each sensor for each volcano \"one by one\", i.e. make each sensor data belong to the [-1,1] interval exactly. In my opinion (but I might be wrong) this would not be a good option because the original range might be important, and the range of each sensor relatively to other sensors might be too. This is only an assumption, I haven't checked it.</li>\n<li>Normalize the data of each volcano, i.e. make all the values of each volcano belong to the [-1,1] interval. This way, you keep the range of each sensor relatively to others unchanged (up to the scale of course). But the information of the range of each volcano relatively to other volcanoes is lost, since all volcanoes now belong to the [-1,1] interval.</li>\n<li>Normalize the entire dataset so that it belongs to the [-1,1] interval. You only change the global range of your data, and keep the information of each sensor and each volcano relatively to each other. This is (in my opinion) the best option. This might be beneficial for some algorithms (e.g. SVM and neural networks which work better with normalized inputs) or might have no effect at all (for tree based algorithms for example). But working with data belonging to a given interval might be easier than dealing with the original range.</li>\n</ul>",
      "rawMarkdown": "There are several ways you can do such normalization : \n- Normalize the data of each sensor for each volcano \"one by one\", i.e. make each sensor data belong to the [-1,1] interval exactly. In my opinion (but I might be wrong) this would not be a good option because the original range might be important, and the range of each sensor relatively to other sensors might be too. This is only an assumption, I haven't checked it.\n- Normalize the data of each volcano, i.e. make all the values of each volcano belong to the [-1,1] interval. This way, you keep the range of each sensor relatively to others unchanged (up to the scale of course). But the information of the range of each volcano relatively to other volcanoes is lost, since all volcanoes now belong to the [-1,1] interval.\n- Normalize the entire dataset so that it belongs to the [-1,1] interval. You only change the global range of your data, and keep the information of each sensor and each volcano relatively to each other. This is (in my opinion) the best option. This might be beneficial for some algorithms (e.g. SVM and neural networks which work better with normalized inputs) or might have no effect at all (for tree based algorithms for example). But working with data belonging to a given interval might be easier than dealing with the original range.",
      "votes": null
    },
    {
      "id": "1128933",
      "postDate": "12/27/2020 22:32:43",
      "content": "<p>thank you so much for reply, could you maybe provide me with some code how to exactly execute that ? </p>",
      "rawMarkdown": "thank you so much for reply, could you maybe provide me with some code how to exactly execute that ?",
      "votes": null
    },
    {
      "id": "1129249",
      "postDate": "12/28/2020 07:17:18",
      "content": "<p>You can rescale data to a given range [a,b] using the following formula : <br>\nx_rescaled = a+(x-min(x))*(b-a)/(max(x)-min(x))<br>\nIf you want to use the last option, assuming X is a t * 10 * n array (where t is the number of measures realized by each sensor and n is the number of volcanoes), </p>\n<ul>\n<li>if you use Python, you can compute the following : <br>\nimport numpy as np<br>\na = -1<br>\nb = 1<br>\nX_rescaled = a+(X-np.min(X))*(b-a)/(np.max(X)-np.min(X))</li>\n<li>if you use R :<br>\na = -1<br>\nb = 1<br>\nX_rescaled = a+(X-min(X))*(b-a)/(max(X)-min(X))<br>\nIf you want to use one of the two first options, it is a bit less straightforward but you can do it using the same formula, and using the apply_along_axis / apply_over_axes functions (in Python) or the apply function (in R).</li>\n</ul>",
      "rawMarkdown": "You can rescale data to a given range [a,b] using the following formula : \nx_rescaled = a+(x-min(x))*(b-a)/(max(x)-min(x))\nIf you want to use the last option, assuming X is a t * 10 * n array (where t is the number of measures realized by each sensor and n is the number of volcanoes), \n- if you use Python, you can compute the following : \nimport numpy as np\na = -1\nb = 1\nX_rescaled = a+(X-np.min(X))*(b-a)/(np.max(X)-np.min(X))\n- if you use R :\na = -1\nb = 1\nX_rescaled = a+(X-min(X))*(b-a)/(max(X)-min(X))\nIf you want to use one of the two first options, it is a bit less straightforward but you can do it using the same formula, and using the apply_along_axis / apply_over_axes functions (in Python) or the apply function (in R).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1128263,
      "author_name": "florian12",
      "author_url": "",
      "post_date": "12/27/2020 09:31:46",
      "content": "<p>There are several ways you can do such normalization : </p>\n<ul>\n<li>Normalize the data of each sensor for each volcano \"one by one\", i.e. make each sensor data belong to the [-1,1] interval exactly. In my opinion (but I might be wrong) this would not be a good option because the original range might be important, and the range of each sensor relatively to other sensors might be too. This is only an assumption, I haven't checked it.</li>\n<li>Normalize the data of each volcano, i.e. make all the values of each volcano belong to the [-1,1] interval. This way, you keep the range of each sensor relatively to others unchanged (up to the scale of course). But the information of the range of each volcano relatively to other volcanoes is lost, since all volcanoes now belong to the [-1,1] interval.</li>\n<li>Normalize the entire dataset so that it belongs to the [-1,1] interval. You only change the global range of your data, and keep the information of each sensor and each volcano relatively to each other. This is (in my opinion) the best option. This might be beneficial for some algorithms (e.g. SVM and neural networks which work better with normalized inputs) or might have no effect at all (for tree based algorithms for example). But working with data belonging to a given interval might be easier than dealing with the original range.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1128933,
          "author_name": "maciejgronczynski",
          "author_url": "",
          "post_date": "12/27/2020 22:32:43",
          "content": "<p>thank you so much for reply, could you maybe provide me with some code how to exactly execute that ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129249,
          "author_name": "florian12",
          "author_url": "",
          "post_date": "12/28/2020 07:17:18",
          "content": "<p>You can rescale data to a given range [a,b] using the following formula : <br>\nx_rescaled = a+(x-min(x))*(b-a)/(max(x)-min(x))<br>\nIf you want to use the last option, assuming X is a t * 10 * n array (where t is the number of measures realized by each sensor and n is the number of volcanoes), </p>\n<ul>\n<li>if you use Python, you can compute the following : <br>\nimport numpy as np<br>\na = -1<br>\nb = 1<br>\nX_rescaled = a+(X-np.min(X))*(b-a)/(np.max(X)-np.min(X))</li>\n<li>if you use R :<br>\na = -1<br>\nb = 1<br>\nX_rescaled = a+(X-min(X))*(b-a)/(max(X)-min(X))<br>\nIf you want to use one of the two first options, it is a bit less straightforward but you can do it using the same formula, and using the apply_along_axis / apply_over_axes functions (in Python) or the apply function (in R).</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1059823": "Hi I just started to getting along with INGV - Volcanic Eruption Prediction challenge, and I've got quick question. \n\nIs it possible to normalize all inputs in train data (sensor_1,sensor_2 etc.) to be distributed from [-1, to 1]?  and If so It will have any positive or negative effect on build models and predictions?",
    "1128263": "There are several ways you can do such normalization : \n- Normalize the data of each sensor for each volcano \"one by one\", i.e. make each sensor data belong to the [-1,1] interval exactly. In my opinion (but I might be wrong) this would not be a good option because the original range might be important, and the range of each sensor relatively to other sensors might be too. This is only an assumption, I haven't checked it.\n- Normalize the data of each volcano, i.e. make all the values of each volcano belong to the [-1,1] interval. This way, you keep the range of each sensor relatively to others unchanged (up to the scale of course). But the information of the range of each volcano relatively to other volcanoes is lost, since all volcanoes now belong to the [-1,1] interval.\n- Normalize the entire dataset so that it belongs to the [-1,1] interval. You only change the global range of your data, and keep the information of each sensor and each volcano relatively to each other. This is (in my opinion) the best option. This might be beneficial for some algorithms (e.g. SVM and neural networks which work better with normalized inputs) or might have no effect at all (for tree based algorithms for example). But working with data belonging to a given interval might be easier than dealing with the original range.",
    "1128933": "thank you so much for reply, could you maybe provide me with some code how to exactly execute that ?",
    "1129249": "You can rescale data to a given range [a,b] using the following formula : \nx_rescaled = a+(x-min(x))*(b-a)/(max(x)-min(x))\nIf you want to use the last option, assuming X is a t * 10 * n array (where t is the number of measures realized by each sensor and n is the number of volcanoes), \n- if you use Python, you can compute the following : \nimport numpy as np\na = -1\nb = 1\nX_rescaled = a+(X-np.min(X))*(b-a)/(np.max(X)-np.min(X))\n- if you use R :\na = -1\nb = 1\nX_rescaled = a+(X-min(X))*(b-a)/(max(X)-min(X))\nIf you want to use one of the two first options, it is a bit less straightforward but you can do it using the same formula, and using the apply_along_axis / apply_over_axes functions (in Python) or the apply function (in R)."
  },
  "source": "meta"
}