{
  "id": 338313,
  "title": "Data Preprocessing for Neural Networks",
  "url": "/competitions/amex-default-prediction/discussion/338313",
  "author_name": "",
  "post_date": "2022-07-20T04:17:16.261163500Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello Kagglers,</p>\n<p>I learned that it might be a good idea to make your data to be \"gaussian-like\" for it to work well with NN. This is stated separately from normalizing the input and I understand this should be done prior to feature engineering/extraction. In the solutions of other competitions, I haven't seen this practice consistently carried out. I was wondering from your experience if transforming the raw data to gaussian-like distribution is usually helpful?</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "1862869",
      "postDate": "07/20/2022 04:17:16",
      "content": "<p>Hello Kagglers,</p>\n<p>I learned that it might be a good idea to make your data to be \"gaussian-like\" for it to work well with NN. This is stated separately from normalizing the input and I understand this should be done prior to feature engineering/extraction. In the solutions of other competitions, I haven't seen this practice consistently carried out. I was wondering from your experience if transforming the raw data to gaussian-like distribution is usually helpful?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hello Kagglers,\n\nI learned that it might be a good idea to make your data to be \"gaussian-like\" for it to work well with NN. This is stated separately from normalizing the input and I understand this should be done prior to feature engineering/extraction. In the solutions of other competitions, I haven't seen this practice consistently carried out. I was wondering from your experience if transforming the raw data to gaussian-like distribution is usually helpful?\n\nThank you!",
      "votes": null
    },
    {
      "id": "1862898",
      "postDate": "07/20/2022 04:40:00",
      "content": "<p>经验告诉我，在超大的数据集上它的影响很小，但在小数据集是必须要处理的。</p>",
      "rawMarkdown": "经验告诉我，在超大的数据集上它的影响很小，但在小数据集是必须要处理的。",
      "votes": null
    },
    {
      "id": "1865403",
      "postDate": "07/21/2022 19:33:39",
      "content": "<p>I pasted the above comment into Google translate, and this is what it said:</p>\n<blockquote>\n  <p>Experience tells me that it has little impact on very large datasets, but it must be dealt with on small datasets.</p>\n</blockquote>\n<p>I think the downvotes are a little harsh although I guess people want to be able to read in English. Maybe Kaggle could add a translate feature of their own in the future?</p>",
      "rawMarkdown": "I pasted the above comment into Google translate, and this is what it said:\n> Experience tells me that it has little impact on very large datasets, but it must be dealt with on small datasets.\n\nI think the downvotes are a little harsh although I guess people want to be able to read in English. Maybe Kaggle could add a translate feature of their own in the future?",
      "votes": null
    },
    {
      "id": "1867604",
      "postDate": "07/23/2022 11:19:57",
      "content": "<p>I would recommend you to read ‘efficient backdrop’ from lecun to understand what is needed for NNs and why. Main problem here might be asymmetric features.</p>",
      "rawMarkdown": "I would recommend you to read ‘efficient backdrop’ from lecun to understand what is needed for NNs and why. Main problem here might be asymmetric features.",
      "votes": null
    },
    {
      "id": "1868336",
      "postDate": "07/23/2022 23:14:47",
      "content": "<p>I think it depends on the data and the neural network architecture. Some data sets and architectures benefit from having gaussian-like data, while others don't. If you're not sure, it might be worth experimenting with both to see what works best on this data set.</p>",
      "rawMarkdown": "I think it depends on the data and the neural network architecture. Some data sets and architectures benefit from having gaussian-like data, while others don't. If you're not sure, it might be worth experimenting with both to see what works best on this data set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1862898,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "07/20/2022 04:40:00",
      "content": "<p>经验告诉我，在超大的数据集上它的影响很小，但在小数据集是必须要处理的。</p>",
      "votes": null,
      "replies": [
        {
          "id": 1865403,
          "author_name": "datahobbit",
          "author_url": "",
          "post_date": "07/21/2022 19:33:39",
          "content": "<p>I pasted the above comment into Google translate, and this is what it said:</p>\n<blockquote>\n  <p>Experience tells me that it has little impact on very large datasets, but it must be dealt with on small datasets.</p>\n</blockquote>\n<p>I think the downvotes are a little harsh although I guess people want to be able to read in English. Maybe Kaggle could add a translate feature of their own in the future?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1867604,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "07/23/2022 11:19:57",
      "content": "<p>I would recommend you to read ‘efficient backdrop’ from lecun to understand what is needed for NNs and why. Main problem here might be asymmetric features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1868336,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/23/2022 23:14:47",
      "content": "<p>I think it depends on the data and the neural network architecture. Some data sets and architectures benefit from having gaussian-like data, while others don't. If you're not sure, it might be worth experimenting with both to see what works best on this data set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1862869": "Hello Kagglers,\n\nI learned that it might be a good idea to make your data to be \"gaussian-like\" for it to work well with NN. This is stated separately from normalizing the input and I understand this should be done prior to feature engineering/extraction. In the solutions of other competitions, I haven't seen this practice consistently carried out. I was wondering from your experience if transforming the raw data to gaussian-like distribution is usually helpful?\n\nThank you!",
    "1862898": "经验告诉我，在超大的数据集上它的影响很小，但在小数据集是必须要处理的。",
    "1865403": "I pasted the above comment into Google translate, and this is what it said:\n> Experience tells me that it has little impact on very large datasets, but it must be dealt with on small datasets.\n\nI think the downvotes are a little harsh although I guess people want to be able to read in English. Maybe Kaggle could add a translate feature of their own in the future?",
    "1867604": "I would recommend you to read ‘efficient backdrop’ from lecun to understand what is needed for NNs and why. Main problem here might be asymmetric features.",
    "1868336": "I think it depends on the data and the neural network architecture. Some data sets and architectures benefit from having gaussian-like data, while others don't. If you're not sure, it might be worth experimenting with both to see what works best on this data set."
  },
  "source": "meta"
}