{
  "id": 359489,
  "title": "Using a simple NN for this competition",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/359489",
  "author_name": "",
  "post_date": "2022-10-12T09:08:17.911000200Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi there !<br>\nI wanted to practice building a simple NN, since I'm quite new. I didn't want to use online learning, neither special techniques to handle the amount of data (so I worked only with train0 file).</p>\n<p>Here I share you my notebook: <a href=\"https://www.kaggle.com/code/sgduran/rocket-league/notebook?scriptVersionId=107766340\" target=\"_blank\">https://www.kaggle.com/code/sgduran/rocket-league/notebook?scriptVersionId=107766340</a>. I built a multiclass classification problem, where for any entry one of the following three results should happen: either Team A scores within the next 10 secs, either Team B does, either none of them does (so I built an extra column for the target set).</p>\n<p>Although I got to submit a solution which performed quite well, there are many things to improve. To begin with, the loss function and accuracy metrics stopped improving after epoch 10, when they started growing consistently until epoch 50, which was the last one. Should I have stopped at epoch 10?</p>\n<p>I have some other questions: how can I choose an appropriate network architecture? Here I used 5 layers with 64 units and ReLu activations. Should I add batch normalization? And regularization?</p>\n<p>If you have articles to read, or other competitions that would work better for practice, I'd be glad to have them.<br>\nThanks for your help!</p>",
  "messages": [
    {
      "id": "1983827",
      "postDate": "10/12/2022 09:08:17",
      "content": "<p>Hi there !<br>\nI wanted to practice building a simple NN, since I'm quite new. I didn't want to use online learning, neither special techniques to handle the amount of data (so I worked only with train0 file).</p>\n<p>Here I share you my notebook: <a href=\"https://www.kaggle.com/code/sgduran/rocket-league/notebook?scriptVersionId=107766340\" target=\"_blank\">https://www.kaggle.com/code/sgduran/rocket-league/notebook?scriptVersionId=107766340</a>. I built a multiclass classification problem, where for any entry one of the following three results should happen: either Team A scores within the next 10 secs, either Team B does, either none of them does (so I built an extra column for the target set).</p>\n<p>Although I got to submit a solution which performed quite well, there are many things to improve. To begin with, the loss function and accuracy metrics stopped improving after epoch 10, when they started growing consistently until epoch 50, which was the last one. Should I have stopped at epoch 10?</p>\n<p>I have some other questions: how can I choose an appropriate network architecture? Here I used 5 layers with 64 units and ReLu activations. Should I add batch normalization? And regularization?</p>\n<p>If you have articles to read, or other competitions that would work better for practice, I'd be glad to have them.<br>\nThanks for your help!</p>",
      "rawMarkdown": "Hi there !\nI wanted to practice building a simple NN, since I'm quite new. I didn't want to use online learning, neither special techniques to handle the amount of data (so I worked only with train0 file).\n\nHere I share you my notebook: https://www.kaggle.com/code/sgduran/rocket-league/notebook?scriptVersionId=107766340. I built a multiclass classification problem, where for any entry one of the following three results should happen: either Team A scores within the next 10 secs, either Team B does, either none of them does (so I built an extra column for the target set).\n\nAlthough I got to submit a solution which performed quite well, there are many things to improve. To begin with, the loss function and accuracy metrics stopped improving after epoch 10, when they started growing consistently until epoch 50, which was the last one. Should I have stopped at epoch 10?\n\nI have some other questions: how can I choose an appropriate network architecture? Here I used 5 layers with 64 units and ReLu activations. Should I add batch normalization? And regularization?\n\nIf you have articles to read, or other competitions that would work better for practice, I'd be glad to have them.\nThanks for your help!",
      "votes": null
    },
    {
      "id": "1986215",
      "postDate": "10/13/2022 22:20:32",
      "content": "<p>I havent made a model for this one yet, but 5 layers of only 64 nodes each seems quite small. especially when you have such a large dataset to work with. \"You gotta pump those numbers up, those are rookie numbers\"</p>",
      "rawMarkdown": "I havent made a model for this one yet, but 5 layers of only 64 nodes each seems quite small. especially when you have such a large dataset to work with. \"You gotta pump those numbers up, those are rookie numbers\"",
      "votes": null
    },
    {
      "id": "2005027",
      "postDate": "10/26/2022 17:13:03",
      "content": "<p>Hi, thanks for your advice <a href=\"https://www.kaggle.com/grahambroughton\" target=\"_blank\">@grahambroughton</a>!<br>\nI did try to pump those numbers up, look: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/sgduran/rocket-league?scriptVersionId=108070430</a>,<br>\nbut it didn't work. The score was even worse.<br>\nFurthermore, which is even weirder, the predictions for the test set are almost constant! (at least, that's what I saw in the submission.csv). Why would it be?</p>",
      "rawMarkdown": "Hi, thanks for your advice @grahambroughton!\nI did try to pump those numbers up, look: [https://www.kaggle.com/code/sgduran/rocket-league?scriptVersionId=108070430](url),\nbut it didn't work. The score was even worse.\nFurthermore, which is even weirder, the predictions for the test set are almost constant! (at least, that's what I saw in the submission.csv). Why would it be?",
      "votes": null
    },
    {
      "id": "2005028",
      "postDate": "10/26/2022 17:13:25",
      "content": "<p>by the way, did you also tried a NN in your model?</p>",
      "rawMarkdown": "by the way, did you also tried a NN in your model?",
      "votes": null
    },
    {
      "id": "2005289",
      "postDate": "10/26/2022 20:32:34",
      "content": "<p>Hi Santi,</p>\n<p>For information, with an NN (based on Tensorflow) I get a mediocre board score of 0.20045. But I admit that on this competition I mainly tried to work on how to manage large volumes of data ( TensorFlow Dataset with csv file or TFRecord). No model optimization, no use of oof, …</p>\n<pre><code>model = keras.models.Sequential([\n    keras.layers.InputLayer(input_shape=[ X.shape[1]]),\n    keras.layers.BatchNormalization(),\n    keras.layers.Dense(neurons, activation=\"swish\"), #, input_shape=[n_inputs-2]),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dropout(0.25),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dropout(0.25),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dropout(0.25),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dense(y.shape[1],activation=\"sigmoid\"),\n])\n</code></pre>\n<p>with neuron from 100 to 200.</p>\n<p>Only a BatchNormalization at the beginning. No others because the NN is not deep. </p>\n<p>Regards<br>\nThierry </p>",
      "rawMarkdown": "Hi Santi,\n\nFor information, with an NN (based on Tensorflow) I get a mediocre board score of 0.20045. But I admit that on this competition I mainly tried to work on how to manage large volumes of data ( TensorFlow Dataset with csv file or TFRecord). No model optimization, no use of oof, ...\n\n    model = keras.models.Sequential([\n        keras.layers.InputLayer(input_shape=[ X.shape[1]]),\n        keras.layers.BatchNormalization(),\n        keras.layers.Dense(neurons, activation=\"swish\"), #, input_shape=[n_inputs-2]),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dropout(0.25),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dropout(0.25),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dropout(0.25),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dense(y.shape[1],activation=\"sigmoid\"),\n    ])\nwith neuron from 100 to 200.\n\nOnly a BatchNormalization at the beginning. No others because the NN is not deep. \n\nRegards\nThierry",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1986215,
      "author_name": "grahambroughton",
      "author_url": "",
      "post_date": "10/13/2022 22:20:32",
      "content": "<p>I havent made a model for this one yet, but 5 layers of only 64 nodes each seems quite small. especially when you have such a large dataset to work with. \"You gotta pump those numbers up, those are rookie numbers\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2005027,
      "author_name": "sgduran",
      "author_url": "",
      "post_date": "10/26/2022 17:13:03",
      "content": "<p>Hi, thanks for your advice <a href=\"https://www.kaggle.com/grahambroughton\" target=\"_blank\">@grahambroughton</a>!<br>\nI did try to pump those numbers up, look: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/sgduran/rocket-league?scriptVersionId=108070430</a>,<br>\nbut it didn't work. The score was even worse.<br>\nFurthermore, which is even weirder, the predictions for the test set are almost constant! (at least, that's what I saw in the submission.csv). Why would it be?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2005028,
      "author_name": "sgduran",
      "author_url": "",
      "post_date": "10/26/2022 17:13:25",
      "content": "<p>by the way, did you also tried a NN in your model?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2005289,
      "author_name": "thierryneusius",
      "author_url": "",
      "post_date": "10/26/2022 20:32:34",
      "content": "<p>Hi Santi,</p>\n<p>For information, with an NN (based on Tensorflow) I get a mediocre board score of 0.20045. But I admit that on this competition I mainly tried to work on how to manage large volumes of data ( TensorFlow Dataset with csv file or TFRecord). No model optimization, no use of oof, …</p>\n<pre><code>model = keras.models.Sequential([\n    keras.layers.InputLayer(input_shape=[ X.shape[1]]),\n    keras.layers.BatchNormalization(),\n    keras.layers.Dense(neurons, activation=\"swish\"), #, input_shape=[n_inputs-2]),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dropout(0.25),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dropout(0.25),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dropout(0.25),\n    keras.layers.Dense(neurons, activation=\"swish\"),\n    keras.layers.Dense(y.shape[1],activation=\"sigmoid\"),\n])\n</code></pre>\n<p>with neuron from 100 to 200.</p>\n<p>Only a BatchNormalization at the beginning. No others because the NN is not deep. </p>\n<p>Regards<br>\nThierry </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1983827": "Hi there !\nI wanted to practice building a simple NN, since I'm quite new. I didn't want to use online learning, neither special techniques to handle the amount of data (so I worked only with train0 file).\n\nHere I share you my notebook: https://www.kaggle.com/code/sgduran/rocket-league/notebook?scriptVersionId=107766340. I built a multiclass classification problem, where for any entry one of the following three results should happen: either Team A scores within the next 10 secs, either Team B does, either none of them does (so I built an extra column for the target set).\n\nAlthough I got to submit a solution which performed quite well, there are many things to improve. To begin with, the loss function and accuracy metrics stopped improving after epoch 10, when they started growing consistently until epoch 50, which was the last one. Should I have stopped at epoch 10?\n\nI have some other questions: how can I choose an appropriate network architecture? Here I used 5 layers with 64 units and ReLu activations. Should I add batch normalization? And regularization?\n\nIf you have articles to read, or other competitions that would work better for practice, I'd be glad to have them.\nThanks for your help!",
    "1986215": "I havent made a model for this one yet, but 5 layers of only 64 nodes each seems quite small. especially when you have such a large dataset to work with. \"You gotta pump those numbers up, those are rookie numbers\"",
    "2005027": "Hi, thanks for your advice @grahambroughton!\nI did try to pump those numbers up, look: [https://www.kaggle.com/code/sgduran/rocket-league?scriptVersionId=108070430](url),\nbut it didn't work. The score was even worse.\nFurthermore, which is even weirder, the predictions for the test set are almost constant! (at least, that's what I saw in the submission.csv). Why would it be?",
    "2005028": "by the way, did you also tried a NN in your model?",
    "2005289": "Hi Santi,\n\nFor information, with an NN (based on Tensorflow) I get a mediocre board score of 0.20045. But I admit that on this competition I mainly tried to work on how to manage large volumes of data ( TensorFlow Dataset with csv file or TFRecord). No model optimization, no use of oof, ...\n\n    model = keras.models.Sequential([\n        keras.layers.InputLayer(input_shape=[ X.shape[1]]),\n        keras.layers.BatchNormalization(),\n        keras.layers.Dense(neurons, activation=\"swish\"), #, input_shape=[n_inputs-2]),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dropout(0.25),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dropout(0.25),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dropout(0.25),\n        keras.layers.Dense(neurons, activation=\"swish\"),\n        keras.layers.Dense(y.shape[1],activation=\"sigmoid\"),\n    ])\nwith neuron from 100 to 200.\n\nOnly a BatchNormalization at the beginning. No others because the NN is not deep. \n\nRegards\nThierry"
  },
  "source": "meta"
}