{
  "id": 494117,
  "title": "Anyone successful with NN? ",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/494117",
  "author_name": "",
  "post_date": "2024-04-16T03:09:00.850126Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I tried using Neural Networks and found that NN training is quite slow, and the performance is still somewhat inferior to LGB. Has anyone achieved good results with NN?</p>",
  "messages": [
    {
      "id": "2754409",
      "postDate": "04/16/2024 03:09:00",
      "content": "<p>I tried using Neural Networks and found that NN training is quite slow, and the performance is still somewhat inferior to LGB. Has anyone achieved good results with NN?</p>",
      "rawMarkdown": "I tried using Neural Networks and found that NN training is quite slow, and the performance is still somewhat inferior to LGB. Has anyone achieved good results with NN?",
      "votes": null
    },
    {
      "id": "2756165",
      "postDate": "04/16/2024 22:37:24",
      "content": "<p>I've gotten close in terms of AUC. I haven't checked the NN's LB score yet.</p>\n<p>One trick to improve training speed that worked really well for me is to load the training data into the gpu all at once (as opposed to putting it on the gpu in batches). Training 1 epoch went from taking around 2 hours to between 20 and 30 minutes.</p>",
      "rawMarkdown": "I've gotten close in terms of AUC. I haven't checked the NN's LB score yet.\n\nOne trick to improve training speed that worked really well for me is to load the training data into the gpu all at once (as opposed to putting it on the gpu in batches). Training 1 epoch went from taking around 2 hours to between 20 and 30 minutes.",
      "votes": null
    },
    {
      "id": "2756390",
      "postDate": "04/17/2024 02:36:20",
      "content": "<p>Amazing results！What method did you use to fill in the NaN values？</p>",
      "rawMarkdown": "Amazing results！What method did you use to fill in the NaN values？",
      "votes": null
    },
    {
      "id": "2757560",
      "postDate": "04/17/2024 15:24:50",
      "content": "<p>For this kind of table data, and the amount of data is millions of cases, I do not think that the neural network can be better than the tree model, just a simple attempt after I gave up the idea of using neural networks</p>",
      "rawMarkdown": "For this kind of table data, and the amount of data is millions of cases, I do not think that the neural network can be better than the tree model, just a simple attempt after I gave up the idea of using neural networks",
      "votes": null
    },
    {
      "id": "2758052",
      "postDate": "04/17/2024 22:43:42",
      "content": "<p>I first standardised the data and then just filled all NaNs/Nulls with -20. The idea being that the right architecture and use of activations should be able to discriminate Nans/Nulls. There might be better approaches (:</p>",
      "rawMarkdown": "I first standardised the data and then just filled all NaNs/Nulls with -20. The idea being that the right architecture and use of activations should be able to discriminate Nans/Nulls. There might be better approaches (:",
      "votes": null
    },
    {
      "id": "2759932",
      "postDate": "04/19/2024 02:59:12",
      "content": "<p>Can I ask why the -20 fill and not something else?</p>",
      "rawMarkdown": "Can I ask why the -20 fill and not something else?",
      "votes": null
    },
    {
      "id": "2767036",
      "postDate": "04/22/2024 04:48:05",
      "content": "<p>Well after standardisation 1sd is just 1 so I just picked a number that should hopefully be far enough away from the distribution of the data, i.e., easily distinguishable. And I chose a negative number as perhaps most features before removing the mean would be positive numbers though it may not matter.</p>\n<p>I originally chose -1000 but that increased the ram used a bunch so I made it smaller.</p>",
      "rawMarkdown": "Well after standardisation 1sd is just 1 so I just picked a number that should hopefully be far enough away from the distribution of the data, i.e., easily distinguishable. And I chose a negative number as perhaps most features before removing the mean would be positive numbers though it may not matter.\n\nI originally chose -1000 but that increased the ram used a bunch so I made it smaller.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2756165,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "04/16/2024 22:37:24",
      "content": "<p>I've gotten close in terms of AUC. I haven't checked the NN's LB score yet.</p>\n<p>One trick to improve training speed that worked really well for me is to load the training data into the gpu all at once (as opposed to putting it on the gpu in batches). Training 1 epoch went from taking around 2 hours to between 20 and 30 minutes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2756390,
          "author_name": "act18l",
          "author_url": "",
          "post_date": "04/17/2024 02:36:20",
          "content": "<p>Amazing results！What method did you use to fill in the NaN values？</p>",
          "votes": null,
          "replies": [
            {
              "id": 2758052,
              "author_name": "caelhasse",
              "author_url": "",
              "post_date": "04/17/2024 22:43:42",
              "content": "<p>I first standardised the data and then just filled all NaNs/Nulls with -20. The idea being that the right architecture and use of activations should be able to discriminate Nans/Nulls. There might be better approaches (:</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2759932,
                  "author_name": "majiaqi111",
                  "author_url": "",
                  "post_date": "04/19/2024 02:59:12",
                  "content": "<p>Can I ask why the -20 fill and not something else?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2767036,
                      "author_name": "caelhasse",
                      "author_url": "",
                      "post_date": "04/22/2024 04:48:05",
                      "content": "<p>Well after standardisation 1sd is just 1 so I just picked a number that should hopefully be far enough away from the distribution of the data, i.e., easily distinguishable. And I chose a negative number as perhaps most features before removing the mean would be positive numbers though it may not matter.</p>\n<p>I originally chose -1000 but that increased the ram used a bunch so I made it smaller.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2757560,
      "author_name": "wenjhuang",
      "author_url": "",
      "post_date": "04/17/2024 15:24:50",
      "content": "<p>For this kind of table data, and the amount of data is millions of cases, I do not think that the neural network can be better than the tree model, just a simple attempt after I gave up the idea of using neural networks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2754409": "I tried using Neural Networks and found that NN training is quite slow, and the performance is still somewhat inferior to LGB. Has anyone achieved good results with NN?",
    "2756165": "I've gotten close in terms of AUC. I haven't checked the NN's LB score yet.\n\nOne trick to improve training speed that worked really well for me is to load the training data into the gpu all at once (as opposed to putting it on the gpu in batches). Training 1 epoch went from taking around 2 hours to between 20 and 30 minutes.",
    "2756390": "Amazing results！What method did you use to fill in the NaN values？",
    "2757560": "For this kind of table data, and the amount of data is millions of cases, I do not think that the neural network can be better than the tree model, just a simple attempt after I gave up the idea of using neural networks",
    "2758052": "I first standardised the data and then just filled all NaNs/Nulls with -20. The idea being that the right architecture and use of activations should be able to discriminate Nans/Nulls. There might be better approaches (:",
    "2759932": "Can I ask why the -20 fill and not something else?",
    "2767036": "Well after standardisation 1sd is just 1 so I just picked a number that should hopefully be far enough away from the distribution of the data, i.e., easily distinguishable. And I chose a negative number as perhaps most features before removing the mean would be positive numbers though it may not matter.\n\nI originally chose -1000 but that increased the ram used a bunch so I made it smaller."
  },
  "source": "meta"
}