{
  "id": 201488,
  "title": "What is the problem with my approach",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/201488",
  "author_name": "",
  "post_date": "2020-12-05T08:54:11.800582600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My approach was : </p>\n<p>Feature Extraction Process : For each sensor of each segment file I calculated, Max, Min, Mean ,Standard deviation , Kurtosis, Skewness, Signal to Noise Ratio. Therefore I had a tabular data with 4400 rows and 72 colums.</p>\n<p>Classfication : Dense Neural network for regression, 1000 epochs</p>\n<p>It is showing a 11 digit result and in leaderboard, I am standing in the last place. Please help.</p>",
  "messages": [
    {
      "id": "1102744",
      "postDate": "12/05/2020 08:54:11",
      "content": "<p>My approach was : </p>\n<p>Feature Extraction Process : For each sensor of each segment file I calculated, Max, Min, Mean ,Standard deviation , Kurtosis, Skewness, Signal to Noise Ratio. Therefore I had a tabular data with 4400 rows and 72 colums.</p>\n<p>Classfication : Dense Neural network for regression, 1000 epochs</p>\n<p>It is showing a 11 digit result and in leaderboard, I am standing in the last place. Please help.</p>",
      "rawMarkdown": "My approach was : \n\nFeature Extraction Process : For each sensor of each segment file I calculated, Max, Min, Mean ,Standard deviation , Kurtosis, Skewness, Signal to Noise Ratio. Therefore I had a tabular data with 4400 rows and 72 colums.\n\nClassfication : Dense Neural network for regression, 1000 epochs\n\nIt is showing a 11 digit result and in leaderboard, I am standing in the last place. Please help.",
      "votes": null
    },
    {
      "id": "1103258",
      "postDate": "12/05/2020 19:07:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dynamite2055\" target=\"_blank\">@dynamite2055</a>,</p>\n<p>The result you obtained depends on several factors, the first thing that comes to mind is a problem with the normalization of the data and/or the outliers.</p>\n<p>Another cause could be that with the calculated features you are not extracting the necessary information for the prediction, you should try to add more. As far as I've worked with the data, this is mostly a feature engineering problem.</p>\n<p>In general it is a good idea to use Short Time Fourier Transform or Continuous Wavelet Transform, with this you extract information from the spectrum of the signal. From these spectrograms you can calculate some statistics.</p>\n<p>Then you should check the NAs and outliers to see how they affect the data, I'm still having quite a bit of trouble with this.</p>\n<p>Also, as you are possibly going to have a lot of variables, it is necessary to select the ones that serve you the most (this is another big problem).</p>\n<p>ps: sorry for my English, I only speak Spanish (thanks google translate 😁)</p>",
      "rawMarkdown": "Hi @dynamite2055,\n\nThe result you obtained depends on several factors, the first thing that comes to mind is a problem with the normalization of the data and/or the outliers.\n\nAnother cause could be that with the calculated features you are not extracting the necessary information for the prediction, you should try to add more. As far as I've worked with the data, this is mostly a feature engineering problem.\n\nIn general it is a good idea to use Short Time Fourier Transform or Continuous Wavelet Transform, with this you extract information from the spectrum of the signal. From these spectrograms you can calculate some statistics.\n\nThen you should check the NAs and outliers to see how they affect the data, I'm still having quite a bit of trouble with this.\n\nAlso, as you are possibly going to have a lot of variables, it is necessary to select the ones that serve you the most (this is another big problem).\n\nps: sorry for my English, I only speak Spanish (thanks google translate 😁)",
      "votes": null
    },
    {
      "id": "1103792",
      "postDate": "12/06/2020 09:42:32",
      "content": "<p>Well … It is hard to say without looking at the notebook but I believe you might have a bug in the generation of your submission file, since predicting only zeros yields a better score than the one you reported.</p>\n<p>Within your dataset, what is your calculated MAE ?? Have you diagnosed your model using cross validation?<br>\nLet me know if you want to make your notebook available so I can help you diagnose the issue better.</p>\n<p>Best! </p>",
      "rawMarkdown": "Well ... It is hard to say without looking at the notebook but I believe you might have a bug in the generation of your submission file, since predicting only zeros yields a better score than the one you reported.\n\nWithin your dataset, what is your calculated MAE ?? Have you diagnosed your model using cross validation?\nLet me know if you want to make your notebook available so I can help you diagnose the issue better.\n\nBest!",
      "votes": null
    },
    {
      "id": "1108619",
      "postDate": "12/10/2020 20:43:39",
      "content": "<p>Well I can say from a personal experience that Neural Network regressors are not suitable for this problem. I tried NN for different structures/number of neurones/optimizers etc I tried it with different validation schemes, and different datasets and the best I could have is 9.800.000 something. I compared it to XGBoost on the same dataset XGboost gives 5.000.000 while the neural network as I said near the 9.800.000. </p>\n<p>However, make sure you check the following points: </p>\n<ul>\n<li>stop training on the validation minimum loss value to avoid overfitting</li>\n<li>make sure you are using normalization and standardization of your data before feeding it to the NN</li>\n<li>start with a small network and then extend do not start big </li>\n<li>use dropout layers to reduce the overfitting chances. </li>\n</ul>\n<p>Best of luck</p>",
      "rawMarkdown": "Well I can say from a personal experience that Neural Network regressors are not suitable for this problem. I tried NN for different structures/number of neurones/optimizers etc I tried it with different validation schemes, and different datasets and the best I could have is 9.800.000 something. I compared it to XGBoost on the same dataset XGboost gives 5.000.000 while the neural network as I said near the 9.800.000. \n\nHowever, make sure you check the following points: \n- stop training on the validation minimum loss value to avoid overfitting\n- make sure you are using normalization and standardization of your data before feeding it to the NN\n- start with a small network and then extend do not start big \n- use dropout layers to reduce the overfitting chances. \n\nBest of luck",
      "votes": null
    },
    {
      "id": "1109159",
      "postDate": "12/11/2020 11:25:11",
      "content": "<p><a href=\"https://www.kaggle.com/desareca\" target=\"_blank\">@desareca</a> <a href=\"https://www.kaggle.com/sebastianpaez\" target=\"_blank\">@sebastianpaez</a> <a href=\"https://www.kaggle.com/obougacha\" target=\"_blank\">@obougacha</a>  Thank you all for your help. I found the mistake. The bad error value was due to normalizing the train data but not test data. It is now giving a decent 8 digit value.</p>",
      "rawMarkdown": "desareca @sebastianpaez @obougacha  Thank you all for your help. I found the mistake. The bad error value was due to normalizing the train data but not test data. It is now giving a decent 8 digit value.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103258,
      "author_name": "desareca",
      "author_url": "",
      "post_date": "12/05/2020 19:07:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dynamite2055\" target=\"_blank\">@dynamite2055</a>,</p>\n<p>The result you obtained depends on several factors, the first thing that comes to mind is a problem with the normalization of the data and/or the outliers.</p>\n<p>Another cause could be that with the calculated features you are not extracting the necessary information for the prediction, you should try to add more. As far as I've worked with the data, this is mostly a feature engineering problem.</p>\n<p>In general it is a good idea to use Short Time Fourier Transform or Continuous Wavelet Transform, with this you extract information from the spectrum of the signal. From these spectrograms you can calculate some statistics.</p>\n<p>Then you should check the NAs and outliers to see how they affect the data, I'm still having quite a bit of trouble with this.</p>\n<p>Also, as you are possibly going to have a lot of variables, it is necessary to select the ones that serve you the most (this is another big problem).</p>\n<p>ps: sorry for my English, I only speak Spanish (thanks google translate 😁)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103792,
      "author_name": "sebastianpaez",
      "author_url": "",
      "post_date": "12/06/2020 09:42:32",
      "content": "<p>Well … It is hard to say without looking at the notebook but I believe you might have a bug in the generation of your submission file, since predicting only zeros yields a better score than the one you reported.</p>\n<p>Within your dataset, what is your calculated MAE ?? Have you diagnosed your model using cross validation?<br>\nLet me know if you want to make your notebook available so I can help you diagnose the issue better.</p>\n<p>Best! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1108619,
      "author_name": "obougacha",
      "author_url": "",
      "post_date": "12/10/2020 20:43:39",
      "content": "<p>Well I can say from a personal experience that Neural Network regressors are not suitable for this problem. I tried NN for different structures/number of neurones/optimizers etc I tried it with different validation schemes, and different datasets and the best I could have is 9.800.000 something. I compared it to XGBoost on the same dataset XGboost gives 5.000.000 while the neural network as I said near the 9.800.000. </p>\n<p>However, make sure you check the following points: </p>\n<ul>\n<li>stop training on the validation minimum loss value to avoid overfitting</li>\n<li>make sure you are using normalization and standardization of your data before feeding it to the NN</li>\n<li>start with a small network and then extend do not start big </li>\n<li>use dropout layers to reduce the overfitting chances. </li>\n</ul>\n<p>Best of luck</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1109159,
      "author_name": "dynamite2055",
      "author_url": "",
      "post_date": "12/11/2020 11:25:11",
      "content": "<p><a href=\"https://www.kaggle.com/desareca\" target=\"_blank\">@desareca</a> <a href=\"https://www.kaggle.com/sebastianpaez\" target=\"_blank\">@sebastianpaez</a> <a href=\"https://www.kaggle.com/obougacha\" target=\"_blank\">@obougacha</a>  Thank you all for your help. I found the mistake. The bad error value was due to normalizing the train data but not test data. It is now giving a decent 8 digit value.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1102744": "My approach was : \n\nFeature Extraction Process : For each sensor of each segment file I calculated, Max, Min, Mean ,Standard deviation , Kurtosis, Skewness, Signal to Noise Ratio. Therefore I had a tabular data with 4400 rows and 72 colums.\n\nClassfication : Dense Neural network for regression, 1000 epochs\n\nIt is showing a 11 digit result and in leaderboard, I am standing in the last place. Please help.",
    "1103258": "Hi @dynamite2055,\n\nThe result you obtained depends on several factors, the first thing that comes to mind is a problem with the normalization of the data and/or the outliers.\n\nAnother cause could be that with the calculated features you are not extracting the necessary information for the prediction, you should try to add more. As far as I've worked with the data, this is mostly a feature engineering problem.\n\nIn general it is a good idea to use Short Time Fourier Transform or Continuous Wavelet Transform, with this you extract information from the spectrum of the signal. From these spectrograms you can calculate some statistics.\n\nThen you should check the NAs and outliers to see how they affect the data, I'm still having quite a bit of trouble with this.\n\nAlso, as you are possibly going to have a lot of variables, it is necessary to select the ones that serve you the most (this is another big problem).\n\nps: sorry for my English, I only speak Spanish (thanks google translate 😁)",
    "1103792": "Well ... It is hard to say without looking at the notebook but I believe you might have a bug in the generation of your submission file, since predicting only zeros yields a better score than the one you reported.\n\nWithin your dataset, what is your calculated MAE ?? Have you diagnosed your model using cross validation?\nLet me know if you want to make your notebook available so I can help you diagnose the issue better.\n\nBest!",
    "1108619": "Well I can say from a personal experience that Neural Network regressors are not suitable for this problem. I tried NN for different structures/number of neurones/optimizers etc I tried it with different validation schemes, and different datasets and the best I could have is 9.800.000 something. I compared it to XGBoost on the same dataset XGboost gives 5.000.000 while the neural network as I said near the 9.800.000. \n\nHowever, make sure you check the following points: \n- stop training on the validation minimum loss value to avoid overfitting\n- make sure you are using normalization and standardization of your data before feeding it to the NN\n- start with a small network and then extend do not start big \n- use dropout layers to reduce the overfitting chances. \n\nBest of luck",
    "1109159": "desareca @sebastianpaez @obougacha  Thank you all for your help. I found the mistake. The bad error value was due to normalizing the train data but not test data. It is now giving a decent 8 digit value."
  },
  "source": "meta"
}