{
  "id": 250624,
  "title": "Data normalization without data distortion",
  "url": "/competitions/seti-breakthrough-listen/discussion/250624",
  "author_name": "Vadim",
  "post_date": "2021-07-03T15:01:20.150000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>All methods about data normalization that I've used either had data distortion or didn't make sense to improve the model. I found the best normalization algorithm (for me) with distribution [0-1] and improved model convergence. I guess it also removed some noise from the data, but I'm not sure. Please upvote if it helped you.</p>\n<pre><code>img = np.vstack(np.load(file)).transpose((1, 0))\nnormalized = (img - np.min(img)) / np.ptp(img)\n</code></pre>",
  "messages": [
    {
      "id": 1374744,
      "postDate": "2021-07-03T15:01:20.150Z",
      "content": "<p>All methods about data normalization that I've used either had data distortion or didn't make sense to improve the model. I found the best normalization algorithm (for me) with distribution [0-1] and improved model convergence. I guess it also removed some noise from the data, but I'm not sure. Please upvote if it helped you.</p>\n<pre><code>img = np.vstack(np.load(file)).transpose((1, 0))\nnormalized = (img - np.min(img)) / np.ptp(img)\n</code></pre>",
      "rawMarkdown": "All methods about data normalization that I've used either had data distortion or didn't make sense to improve the model. I found the best normalization algorithm (for me) with distribution [0-1] and improved model convergence. I guess it also removed some noise from the data, but I'm not sure. Please upvote if it helped you.\n```\nimg = np.vstack(np.load(file)).transpose((1, 0))\nnormalized = (img - np.min(img)) / np.ptp(img)\n```",
      "votes": 5
    },
    {
      "id": 1669482,
      "postDate": "2022-01-30T18:04:42.927Z",
      "content": "<p>Thanks for sharing! </p>\n<p>On data leak: A common error is to apply it to the entire data before splitting it into training and test sets.</p>\n<p>To check index and leakage and identifier leakage, you can use <a href=\"https://docs.deepchecks.com/en/stable/examples/guides/quickstart_in_5_minutes.html``\" target=\"_blank\">deepchecks</a></p>\n<pre><code>from deepchecks.base import Dataset\nfrom deepchecks.checks import IndexTrainTestLeakage\nimport pandas as pd\n\ndef dataset_from_dict(d: dict, index_name: str = None) -&gt; Dataset:\n    dataframe = pd.DataFrame(data=d)\n    return Dataset(dataframe, index_name=index_name)\n</code></pre>\n<p>Synthetic leakage </p>\n<pre><code>train_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')\ncheck_obj = IndexTrainTestLeakage()\ncheck_obj.run(train_ds, test_ds)\n</code></pre>\n<p>Index train test leakage </p>\n<p>train_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')<br>\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')<br>\ncheck_obj = IndexTrainTestLeakage(n_index_to_show=1)<br>\ncheck_obj.run(train_ds, test_ds)</p>\n<p>[<a href=\"https://docs.deepchecks.com/en/stable/examples/checks/integrity/data_duplicates.html[](url)\" target=\"_blank\">See detailed codes</a>] </p>",
      "rawMarkdown": "Thanks for sharing! \n\nOn data leak: A common error is to apply it to the entire data before splitting it into training and test sets.\n\nTo check index and leakage and identifier leakage, you can use [deepchecks](https://docs.deepchecks.com/en/stable/examples/guides/quickstart_in_5_minutes.html``)\n\n```\nfrom deepchecks.base import Dataset\nfrom deepchecks.checks import IndexTrainTestLeakage\nimport pandas as pd\n\ndef dataset_from_dict(d: dict, index_name: str = None) -> Dataset:\n    dataframe = pd.DataFrame(data=d)\n    return Dataset(dataframe, index_name=index_name)\n```\n\n\nSynthetic leakage \n\n```\ntrain_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')\ncheck_obj = IndexTrainTestLeakage()\ncheck_obj.run(train_ds, test_ds)\n```\n\nIndex train test leakage \n\ntrain_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')\ncheck_obj = IndexTrainTestLeakage(n_index_to_show=1)\ncheck_obj.run(train_ds, test_ds)\n\n[[See detailed codes](https://docs.deepchecks.com/en/stable/examples/checks/integrity/data_duplicates.html[](url))] \n",
      "votes": 2
    },
    {
      "id": 1376873,
      "postDate": "2021-07-05T12:01:43.287Z",
      "content": "<p>your normalization does not remove some leak, and it helps the models compared to the normalization I have shared.  You need to do it for each of the 6 channels separately to remove the leak.</p>\n<p>Maybe your way will be better on the new data, though.</p>",
      "rawMarkdown": "your normalization does not remove some leak, and it helps the models compared to the normalization I have shared.  You need to do it for each of the 6 channels separately to remove the leak.\n\nMaybe your way will be better on the new data, though.",
      "votes": 1
    },
    {
      "id": 1731640,
      "postDate": "2022-03-22T15:04:06.783Z",
      "content": "<p><a href=\"https://www.kaggle.com/richardkleins\" target=\"_blank\">@richardkleins</a> thank you for sharing</p>",
      "rawMarkdown": "@richardkleins thank you for sharing"
    },
    {
      "id": 1374753,
      "postDate": "2021-07-03T15:10:25.543Z",
      "content": "<p>Helpful, thanks for sharing</p>",
      "rawMarkdown": "Helpful, thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 1669482,
      "author_name": "Richard Kleins",
      "author_url": "",
      "post_date": "2022-01-30T18:04:42.927000",
      "content": "<p>Thanks for sharing! </p>\n<p>On data leak: A common error is to apply it to the entire data before splitting it into training and test sets.</p>\n<p>To check index and leakage and identifier leakage, you can use <a href=\"https://docs.deepchecks.com/en/stable/examples/guides/quickstart_in_5_minutes.html``\" target=\"_blank\">deepchecks</a></p>\n<pre><code>from deepchecks.base import Dataset\nfrom deepchecks.checks import IndexTrainTestLeakage\nimport pandas as pd\n\ndef dataset_from_dict(d: dict, index_name: str = None) -&gt; Dataset:\n    dataframe = pd.DataFrame(data=d)\n    return Dataset(dataframe, index_name=index_name)\n</code></pre>\n<p>Synthetic leakage </p>\n<pre><code>train_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')\ncheck_obj = IndexTrainTestLeakage()\ncheck_obj.run(train_ds, test_ds)\n</code></pre>\n<p>Index train test leakage </p>\n<p>train_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')<br>\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')<br>\ncheck_obj = IndexTrainTestLeakage(n_index_to_show=1)<br>\ncheck_obj.run(train_ds, test_ds)</p>\n<p>[<a href=\"https://docs.deepchecks.com/en/stable/examples/checks/integrity/data_duplicates.html[](url)\" target=\"_blank\">See detailed codes</a>] </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1376873,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-07-05T12:01:43.287000",
      "content": "<p>your normalization does not remove some leak, and it helps the models compared to the normalization I have shared.  You need to do it for each of the 6 channels separately to remove the leak.</p>\n<p>Maybe your way will be better on the new data, though.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1731640,
      "author_name": "Greg Andersons",
      "author_url": "",
      "post_date": "2022-03-22T15:04:06.783000",
      "content": "<p><a href=\"https://www.kaggle.com/richardkleins\" target=\"_blank\">@richardkleins</a> thank you for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1374753,
      "author_name": "Samarth Gupta",
      "author_url": "",
      "post_date": "2021-07-03T15:10:25.543000",
      "content": "<p>Helpful, thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1374744": "All methods about data normalization that I've used either had data distortion or didn't make sense to improve the model. I found the best normalization algorithm (for me) with distribution [0-1] and improved model convergence. I guess it also removed some noise from the data, but I'm not sure. Please upvote if it helped you.\n```\nimg = np.vstack(np.load(file)).transpose((1, 0))\nnormalized = (img - np.min(img)) / np.ptp(img)\n```",
    "1669482": "Thanks for sharing! \n\nOn data leak: A common error is to apply it to the entire data before splitting it into training and test sets.\n\nTo check index and leakage and identifier leakage, you can use [deepchecks](https://docs.deepchecks.com/en/stable/examples/guides/quickstart_in_5_minutes.html``)\n\n```\nfrom deepchecks.base import Dataset\nfrom deepchecks.checks import IndexTrainTestLeakage\nimport pandas as pd\n\ndef dataset_from_dict(d: dict, index_name: str = None) -> Dataset:\n    dataframe = pd.DataFrame(data=d)\n    return Dataset(dataframe, index_name=index_name)\n```\n\n\nSynthetic leakage \n\n```\ntrain_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')\ncheck_obj = IndexTrainTestLeakage()\ncheck_obj.run(train_ds, test_ds)\n```\n\nIndex train test leakage \n\ntrain_ds = dataset_from_dict({'col1': [1, 2, 3, 4, 10, 11]}, 'col1')\ntest_ds = dataset_from_dict({'col1': [4, 3, 5, 6, 7]}, 'col1')\ncheck_obj = IndexTrainTestLeakage(n_index_to_show=1)\ncheck_obj.run(train_ds, test_ds)\n\n[[See detailed codes](https://docs.deepchecks.com/en/stable/examples/checks/integrity/data_duplicates.html[](url))] \n",
    "1376873": "your normalization does not remove some leak, and it helps the models compared to the normalization I have shared.  You need to do it for each of the 6 channels separately to remove the leak.\n\nMaybe your way will be better on the new data, though.",
    "1731640": "@richardkleins thank you for sharing",
    "1374753": "Helpful, thanks for sharing"
  }
}