{
  "id": 194300,
  "title": "Adversarial Validation - There is something fishy about the test data ",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/194300",
  "author_name": "",
  "post_date": "2020-11-01T02:22:56.625069300Z",
  "votes": 11,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi all,<br>\nI performed adversarial validation notebook here -&gt; <a href=\"https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a> . I identify a set of 1557 segments that do not have good proxies in train and I just do some views of the distributions of <code>time_to_eruptions</code> to see when the sensor failures happen. I then assume my predictions for the 1557 segments are directional correct, but remap the distribution to the distribution where sensor 10 is down in train.</p>\n<p>Using this simple PDF remap I am able to take the predictions from my original notebook: <a href=\"https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline</a> with a score in the <code>6.6e6</code> range and improve it to <code>5.8e6</code>.</p>\n<p>The notebook only takes half an hour to run, the modeling section with mini NNs takes 5 minutes.</p>",
  "messages": [
    {
      "id": "1065900",
      "postDate": "11/01/2020 02:22:56",
      "content": "<p>Hi all,<br>\nI performed adversarial validation notebook here -&gt; <a href=\"https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a> . I identify a set of 1557 segments that do not have good proxies in train and I just do some views of the distributions of <code>time_to_eruptions</code> to see when the sensor failures happen. I then assume my predictions for the 1557 segments are directional correct, but remap the distribution to the distribution where sensor 10 is down in train.</p>\n<p>Using this simple PDF remap I am able to take the predictions from my original notebook: <a href=\"https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline</a> with a score in the <code>6.6e6</code> range and improve it to <code>5.8e6</code>.</p>\n<p>The notebook only takes half an hour to run, the modeling section with mini NNs takes 5 minutes.</p>",
      "rawMarkdown": "Hi all,\nI performed adversarial validation notebook here -> https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences . I identify a set of 1557 segments that do not have good proxies in train and I just do some views of the distributions of `time_to_eruptions` to see when the sensor failures happen. I then assume my predictions for the 1557 segments are directional correct, but remap the distribution to the distribution where sensor 10 is down in train.\n\nUsing this simple PDF remap I am able to take the predictions from my original notebook: https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline with a score in the `6.6e6` range and improve it to `5.8e6`.\n\nThe notebook only takes half an hour to run, the modeling section with mini NNs takes 5 minutes.",
      "votes": null
    },
    {
      "id": "1078497",
      "postDate": "11/14/2020 21:52:15",
      "content": "<p>Thank you for your post. I generated plenty of features for each sensor. I have computed KS statistics between Train and Test sets for each feature. And I discovered this :<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F275525%2Fabb57aca6d06681297d213eda1ea7b9c%2FScreenshot_20201114_224332.png?generation=1605390475333869&amp;alt=media\" alt=\"\"><br>\nEach point is the KS statistics for a feature. It confirms that sensor 10 is dangerous (because KS statistics are high for features from this sensor). But it is not the only one…</p>",
      "rawMarkdown": "Thank you for your post. I generated plenty of features for each sensor. I have computed KS statistics between Train and Test sets for each feature. And I discovered this :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F275525%2Fabb57aca6d06681297d213eda1ea7b9c%2FScreenshot_20201114_224332.png?generation=1605390475333869&alt=media)\nEach point is the KS statistics for a feature. It confirms that sensor 10 is dangerous (because KS statistics are high for features from this sensor). But it is not the only one...",
      "votes": null
    },
    {
      "id": "1085633",
      "postDate": "11/21/2020 03:27:07",
      "content": "<p>Has anyone made any progress on this? I've tried a number of methods to identify and remove 'bad actors' but have not had much luck on converging between the MAE on the CV and the test set. When I've reduced the features with poor distribution, it has continued to improve the MAE within cross validation but still no progress on test set.</p>",
      "rawMarkdown": "Has anyone made any progress on this? I've tried a number of methods to identify and remove 'bad actors' but have not had much luck on converging between the MAE on the CV and the test set. When I've reduced the features with poor distribution, it has continued to improve the MAE within cross validation but still no progress on test set.",
      "votes": null
    },
    {
      "id": "1085640",
      "postDate": "11/21/2020 03:46:23",
      "content": "<p>hmm the KS stats for sensors 2, 5, 9 are not fantastic. We would hope to be &lt; 0.05 so we can reject the null hypothesis</p>",
      "rawMarkdown": "hmm the KS stats for sensors 2, 5, 9 are not fantastic. We would hope to be < 0.05 so we can reject the null hypothesis",
      "votes": null
    },
    {
      "id": "1088579",
      "postDate": "11/23/2020 18:57:09",
      "content": "<p>Hi Kyle, </p>\n<p>I tried the same at first, but no improvments were obtained by filtering some sensors. <br>\nI tried also a reduction of the number of variables in the model. The results were fine on train set, but at the end on the LB the results were quite disappointing.</p>\n<p>Currently i'm working on a simple data set, based on the \"INGV Volcanic : Basic solution (STFT)\" notebook and got better results on LB by exploring cross validation methods rather than trying to remove features. By the way this notebook is very interesting because it is really reducing the eruption aftershocks noise saw by using the standard indicators (sum, std, mean,….).</p>\n<p>In my case, i was blocked around 5.5m on LB but trying to reproduce the candy box experiment ( show a transparent candy box full of candies and ask to 1 000 persons to give you an estimation of the number of candies inside, and at the end, you will notice that the mean of their predictions will converge to the real number of candies inside the box - \"wisdom of the crowd\" if think-), I've got 4.7m results now. <br>\nThe idea, because the dataset is small when using aggregate datas on each segment, is to multiple the evaluation made by the model. So i gave up the Kfold/RepeatedKfold/…. and using now only the shuffleSplit ( splitting randomly the train set 2000 times with a repartition of 70-30 between the train and validation dataset) with aggressive values for the LGBM model ( for instance n_estimators set to 800).  It gave me a sensible improvment of the LB, because this method seems to make the outliers disapearing and helps with the overfitting issue by having a more suitable way of generalize the features estimators.</p>\n<p>The huge default is the fact that this method is requiring lot of computation. The second default, it will not help to improve the lack of of consistency between the train and test datasets as shown by the notebook created by Adam James. SInce the beginning of the competition, i'm surprized by the test dataset size, and by the fact that the train an test dataset are not homogenous. But this is i think the key point of this competition.</p>\n<p>May be a stack of model can help to improve also the LB ( I only used it for classification problem, and i don't know if it's working on regression problem)</p>\n<p>Hope sharing my vision will help and give you ideas.</p>",
      "rawMarkdown": "Hi Kyle, \n\nI tried the same at first, but no improvments were obtained by filtering some sensors. \nI tried also a reduction of the number of variables in the model. The results were fine on train set, but at the end on the LB the results were quite disappointing.\n\nCurrently i'm working on a simple data set, based on the \"INGV Volcanic : Basic solution (STFT)\" notebook and got better results on LB by exploring cross validation methods rather than trying to remove features. By the way this notebook is very interesting because it is really reducing the eruption aftershocks noise saw by using the standard indicators (sum, std, mean,....).\n\nIn my case, i was blocked around 5.5m on LB but trying to reproduce the candy box experiment ( show a transparent candy box full of candies and ask to 1 000 persons to give you an estimation of the number of candies inside, and at the end, you will notice that the mean of their predictions will converge to the real number of candies inside the box - \"wisdom of the crowd\" if think-), I've got 4.7m results now. \nThe idea, because the dataset is small when using aggregate datas on each segment, is to multiple the evaluation made by the model. So i gave up the Kfold/RepeatedKfold/.... and using now only the shuffleSplit ( splitting randomly the train set 2000 times with a repartition of 70-30 between the train and validation dataset) with aggressive values for the LGBM model ( for instance n_estimators set to 800).  It gave me a sensible improvment of the LB, because this method seems to make the outliers disapearing and helps with the overfitting issue by having a more suitable way of generalize the features estimators.\n\nThe huge default is the fact that this method is requiring lot of computation. The second default, it will not help to improve the lack of of consistency between the train and test datasets as shown by the notebook created by Adam James. SInce the beginning of the competition, i'm surprized by the test dataset size, and by the fact that the train an test dataset are not homogenous. But this is i think the key point of this competition.\n\nMay be a stack of model can help to improve also the LB ( I only used it for classification problem, and i don't know if it's working on regression problem)\n\nHope sharing my vision will help and give you ideas.",
      "votes": null
    },
    {
      "id": "1088796",
      "postDate": "11/24/2020 00:23:45",
      "content": "<p>Thank you for sharing. I have not explored the training method as you suggested which I'll play around with. I've been stuck at the 5.5m on LB for awhile. </p>\n<p>I had incremental success trying to expand the training set by creating a multi-output regression model. So looking at the time_to_eruption distribution, I generated more training data  to expand the training set. It had not resulted in the improvement you had suggested training the model multiple times but for sake of computational time, I hadn't produced a significant amount of additional training data. Another option to explore.</p>",
      "rawMarkdown": "Thank you for sharing. I have not explored the training method as you suggested which I'll play around with. I've been stuck at the 5.5m on LB for awhile. \n\nI had incremental success trying to expand the training set by creating a multi-output regression model. So looking at the time_to_eruption distribution, I generated more training data  to expand the training set. It had not resulted in the improvement you had suggested training the model multiple times but for sake of computational time, I hadn't produced a significant amount of additional training data. Another option to explore.",
      "votes": null
    },
    {
      "id": "1088808",
      "postDate": "11/24/2020 00:40:58",
      "content": "<p>Regarding the use of shuffleSplit, just be aware that it looks very sensitive to the features you are taking. For instance i just finished a 20hours calculations with a larger dataset than the previous one   but it gave me poor LB results at the end. I have also to learn and explore this method 😔. Just sharing the trick that helped me.</p>",
      "rawMarkdown": "Regarding the use of shuffleSplit, just be aware that it looks very sensitive to the features you are taking. For instance i just finished a 20hours calculations with a larger dataset than the previous one   but it gave me poor LB results at the end. I have also to learn and explore this method 😔. Just sharing the trick that helped me.",
      "votes": null
    },
    {
      "id": "1091108",
      "postDate": "11/25/2020 19:13:23",
      "content": "<p>Hello, Thank you both for sharing your ideas. I thought I was going crazy and that I missed something when creating my submission file but it seems that there is quite a difference between the results of the CV and those on the LB. </p>\n<p>To be honest, I spent quite a while trying different methods for instance I used neural networks with Tensorflow (That was not a good idea) I'm getting fairly low MAE on the CV around 4.9m but disappointingly insane results on the LB around 18.9m and 22m. I just wanted to explore the NN regressions. </p>\n<p>Also tried the stacking part, where I created several base model from RandomForest, XGBoost, LGBM, etc.. then tried several stacking levels. I was quite excited about this method when I saw the CV 2.5m with 0.89 r2 (I was quite careful not to fall into the overfitting trap). However, I was shocked with 22m. <br>\nSo I guess even stacking is not a great idea. </p>\n<p>Finally, I tried something quite crazy computationally. I created a bootstrapped NNs. So I made quite a 30 or 40 NN each trained with a random part of the Training part of the CV and also with a random selection of the features. (I thought to myself since we have random sensors failures at least that my hypothesis, so the model input should be robust to this). Well, this made the NN score a bit better 10.9m. Then I tried something even crazier. I used a 200% sample with a replacement method when I create the part of training for each NN sub-model. Well, It took forever to train only 40 dummy Neural Network but it managed to drop the LB to 8.7m. This got me thinking maybe a smaller part of this training set represents the test distribution and by sampling with replacement, I accidentally reinforced this part. </p>",
      "rawMarkdown": "Hello, Thank you both for sharing your ideas. I thought I was going crazy and that I missed something when creating my submission file but it seems that there is quite a difference between the results of the CV and those on the LB. \n\nTo be honest, I spent quite a while trying different methods for instance I used neural networks with Tensorflow (That was not a good idea) I'm getting fairly low MAE on the CV around 4.9m but disappointingly insane results on the LB around 18.9m and 22m. I just wanted to explore the NN regressions. \n\nAlso tried the stacking part, where I created several base model from RandomForest, XGBoost, LGBM, etc.. then tried several stacking levels. I was quite excited about this method when I saw the CV 2.5m with 0.89 r2 (I was quite careful not to fall into the overfitting trap). However, I was shocked with 22m. \nSo I guess even stacking is not a great idea. \n\nFinally, I tried something quite crazy computationally. I created a bootstrapped NNs. So I made quite a 30 or 40 NN each trained with a random part of the Training part of the CV and also with a random selection of the features. (I thought to myself since we have random sensors failures at least that my hypothesis, so the model input should be robust to this). Well, this made the NN score a bit better 10.9m. Then I tried something even crazier. I used a 200% sample with a replacement method when I create the part of training for each NN sub-model. Well, It took forever to train only 40 dummy Neural Network but it managed to drop the LB to 8.7m. This got me thinking maybe a smaller part of this training set represents the test distribution and by sampling with replacement, I accidentally reinforced this part.",
      "votes": null
    },
    {
      "id": "1091230",
      "postDate": "11/25/2020 21:59:23",
      "content": "<p>Hi, to be honest i didn't undertstood all the scenarios you had tried, especially for the NN part. But the values that you had encountered look like you're doing overfitting. Also, if your models (XGB, LGBM,…) doing overfitting, the stack model will do the same. May be restarting from a basic model ( for instance my preference is for this great notebook<a href=\"url\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a> will help you to restart from a good point. I'm not sure that a NN model will help regarding the datas in this competition, except if you are able to detect the eruption sequences as explained int the notebook of this post.</p>\n<p>BR</p>",
      "rawMarkdown": "Hi, to be honest i didn't undertstood all the scenarios you had tried, especially for the NN part. But the values that you had encountered look like you're doing overfitting. Also, if your models (XGB, LGBM,...) doing overfitting, the stack model will do the same. May be restarting from a basic model ( for instance my preference is for this great notebook[https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft](url) will help you to restart from a good point. I'm not sure that a NN model will help regarding the datas in this competition, except if you are able to detect the eruption sequences as explained int the notebook of this post.\n\nBR",
      "votes": null
    },
    {
      "id": "1091614",
      "postDate": "11/26/2020 06:21:29",
      "content": "<p>For the NN I created several dummy models with only two layers with very few neurons. Then, each model is trained on a subset of the features and a subset of the training set (random sampling with a replacement for the training set). It's like doing bootstrapping for random forests. I hope it's more clear.</p>\n<p>I'm quite sure I'm not overfitting. My CV training errors are around 4.5 - 3.2 m and the validation errors are around 5.8 - 4.5m depending on the model. I think the values are quite acceptable and do not alarm for overfitting.</p>\n<p>My guess is that those 1500 and so entries in the test set that were detected by Adam James are the ones that causing me problems. I think their real values of time to eruption is quite different from the one that we predict with our models.  Since these points are different from any other points the model has a hard time predicting their values. And the more sophisticated and well-fitted model to the training set the worst it is doing in predicting those outliers. This is quite the only explanation I can think of. </p>",
      "rawMarkdown": "For the NN I created several dummy models with only two layers with very few neurons. Then, each model is trained on a subset of the features and a subset of the training set (random sampling with a replacement for the training set). It's like doing bootstrapping for random forests. I hope it's more clear.\n\nI'm quite sure I'm not overfitting. My CV training errors are around 4.5 - 3.2 m and the validation errors are around 5.8 - 4.5m depending on the model. I think the values are quite acceptable and do not alarm for overfitting.\n\nMy guess is that those 1500 and so entries in the test set that were detected by Adam James are the ones that causing me problems. I think their real values of time to eruption is quite different from the one that we predict with our models.  Since these points are different from any other points the model has a hard time predicting their values. And the more sophisticated and well-fitted model to the training set the worst it is doing in predicting those outliers. This is quite the only explanation I can think of.",
      "votes": null
    },
    {
      "id": "1093130",
      "postDate": "11/27/2020 13:18:06",
      "content": "<p>Hi,<br>\nNow it's more clear in my mind regarding your method. But your problem is very weird, because even if the test dataset has a bias, it cannot explain such a difference in the LB results, especially if you taking care about the validation datatest results (  I'm doing the same, with a  validation dataset size of 25%-30% of the global number of segments) . May be your model relies too much on the sensors which have important numbers of NaN values in the test dataset. <br>\nMay be in your case it can be interesting to check with your method the results on only sensors 3,4,6 and 7, because in the train/test datasets they are the sensors with quite the same number of NaN values and for these sensors the NaN is less than 10% of the total number of segments for each of them. <br>\nOn my side, still to find solutions to this overfitting problem, i have in my todo list, to check if by taking only segments which don't have NaN sensors or less than 3, the results are better. The annoying thing is that we are using the LB score to help to improbe the model. Not a very nice statistics approach, but i think we don't have the choice in this competition.</p>\n<p>Also tried the NN on my side and having bad results too. I think that the NN are too sensitive to a change of values for a sensor on the global stats to generalize correctly the model. </p>\n<p>Br</p>",
      "rawMarkdown": "Hi,\nNow it's more clear in my mind regarding your method. But your problem is very weird, because even if the test dataset has a bias, it cannot explain such a difference in the LB results, especially if you taking care about the validation datatest results (  I'm doing the same, with a  validation dataset size of 25%-30% of the global number of segments) . May be your model relies too much on the sensors which have important numbers of NaN values in the test dataset. \nMay be in your case it can be interesting to check with your method the results on only sensors 3,4,6 and 7, because in the train/test datasets they are the sensors with quite the same number of NaN values and for these sensors the NaN is less than 10% of the total number of segments for each of them. \nOn my side, still to find solutions to this overfitting problem, i have in my todo list, to check if by taking only segments which don't have NaN sensors or less than 3, the results are better. The annoying thing is that we are using the LB score to help to improbe the model. Not a very nice statistics approach, but i think we don't have the choice in this competition.\n\nAlso tried the NN on my side and having bad results too. I think that the NN are too sensitive to a change of values for a sensor on the global stats to generalize correctly the model. \n\nBr",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1078497,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "11/14/2020 21:52:15",
      "content": "<p>Thank you for your post. I generated plenty of features for each sensor. I have computed KS statistics between Train and Test sets for each feature. And I discovered this :<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F275525%2Fabb57aca6d06681297d213eda1ea7b9c%2FScreenshot_20201114_224332.png?generation=1605390475333869&amp;alt=media\" alt=\"\"><br>\nEach point is the KS statistics for a feature. It confirms that sensor 10 is dangerous (because KS statistics are high for features from this sensor). But it is not the only one…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1085640,
          "author_name": "ajcostarino",
          "author_url": "",
          "post_date": "11/21/2020 03:46:23",
          "content": "<p>hmm the KS stats for sensors 2, 5, 9 are not fantastic. We would hope to be &lt; 0.05 so we can reject the null hypothesis</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1085633,
      "author_name": "kylesnyder",
      "author_url": "",
      "post_date": "11/21/2020 03:27:07",
      "content": "<p>Has anyone made any progress on this? I've tried a number of methods to identify and remove 'bad actors' but have not had much luck on converging between the MAE on the CV and the test set. When I've reduced the features with poor distribution, it has continued to improve the MAE within cross validation but still no progress on test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1088579,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/23/2020 18:57:09",
          "content": "<p>Hi Kyle, </p>\n<p>I tried the same at first, but no improvments were obtained by filtering some sensors. <br>\nI tried also a reduction of the number of variables in the model. The results were fine on train set, but at the end on the LB the results were quite disappointing.</p>\n<p>Currently i'm working on a simple data set, based on the \"INGV Volcanic : Basic solution (STFT)\" notebook and got better results on LB by exploring cross validation methods rather than trying to remove features. By the way this notebook is very interesting because it is really reducing the eruption aftershocks noise saw by using the standard indicators (sum, std, mean,….).</p>\n<p>In my case, i was blocked around 5.5m on LB but trying to reproduce the candy box experiment ( show a transparent candy box full of candies and ask to 1 000 persons to give you an estimation of the number of candies inside, and at the end, you will notice that the mean of their predictions will converge to the real number of candies inside the box - \"wisdom of the crowd\" if think-), I've got 4.7m results now. <br>\nThe idea, because the dataset is small when using aggregate datas on each segment, is to multiple the evaluation made by the model. So i gave up the Kfold/RepeatedKfold/…. and using now only the shuffleSplit ( splitting randomly the train set 2000 times with a repartition of 70-30 between the train and validation dataset) with aggressive values for the LGBM model ( for instance n_estimators set to 800).  It gave me a sensible improvment of the LB, because this method seems to make the outliers disapearing and helps with the overfitting issue by having a more suitable way of generalize the features estimators.</p>\n<p>The huge default is the fact that this method is requiring lot of computation. The second default, it will not help to improve the lack of of consistency between the train and test datasets as shown by the notebook created by Adam James. SInce the beginning of the competition, i'm surprized by the test dataset size, and by the fact that the train an test dataset are not homogenous. But this is i think the key point of this competition.</p>\n<p>May be a stack of model can help to improve also the LB ( I only used it for classification problem, and i don't know if it's working on regression problem)</p>\n<p>Hope sharing my vision will help and give you ideas.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088796,
          "author_name": "kylesnyder",
          "author_url": "",
          "post_date": "11/24/2020 00:23:45",
          "content": "<p>Thank you for sharing. I have not explored the training method as you suggested which I'll play around with. I've been stuck at the 5.5m on LB for awhile. </p>\n<p>I had incremental success trying to expand the training set by creating a multi-output regression model. So looking at the time_to_eruption distribution, I generated more training data  to expand the training set. It had not resulted in the improvement you had suggested training the model multiple times but for sake of computational time, I hadn't produced a significant amount of additional training data. Another option to explore.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088808,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/24/2020 00:40:58",
          "content": "<p>Regarding the use of shuffleSplit, just be aware that it looks very sensitive to the features you are taking. For instance i just finished a 20hours calculations with a larger dataset than the previous one   but it gave me poor LB results at the end. I have also to learn and explore this method 😔. Just sharing the trick that helped me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091108,
          "author_name": "obougacha",
          "author_url": "",
          "post_date": "11/25/2020 19:13:23",
          "content": "<p>Hello, Thank you both for sharing your ideas. I thought I was going crazy and that I missed something when creating my submission file but it seems that there is quite a difference between the results of the CV and those on the LB. </p>\n<p>To be honest, I spent quite a while trying different methods for instance I used neural networks with Tensorflow (That was not a good idea) I'm getting fairly low MAE on the CV around 4.9m but disappointingly insane results on the LB around 18.9m and 22m. I just wanted to explore the NN regressions. </p>\n<p>Also tried the stacking part, where I created several base model from RandomForest, XGBoost, LGBM, etc.. then tried several stacking levels. I was quite excited about this method when I saw the CV 2.5m with 0.89 r2 (I was quite careful not to fall into the overfitting trap). However, I was shocked with 22m. <br>\nSo I guess even stacking is not a great idea. </p>\n<p>Finally, I tried something quite crazy computationally. I created a bootstrapped NNs. So I made quite a 30 or 40 NN each trained with a random part of the Training part of the CV and also with a random selection of the features. (I thought to myself since we have random sensors failures at least that my hypothesis, so the model input should be robust to this). Well, this made the NN score a bit better 10.9m. Then I tried something even crazier. I used a 200% sample with a replacement method when I create the part of training for each NN sub-model. Well, It took forever to train only 40 dummy Neural Network but it managed to drop the LB to 8.7m. This got me thinking maybe a smaller part of this training set represents the test distribution and by sampling with replacement, I accidentally reinforced this part. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091230,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/25/2020 21:59:23",
          "content": "<p>Hi, to be honest i didn't undertstood all the scenarios you had tried, especially for the NN part. But the values that you had encountered look like you're doing overfitting. Also, if your models (XGB, LGBM,…) doing overfitting, the stack model will do the same. May be restarting from a basic model ( for instance my preference is for this great notebook<a href=\"url\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a> will help you to restart from a good point. I'm not sure that a NN model will help regarding the datas in this competition, except if you are able to detect the eruption sequences as explained int the notebook of this post.</p>\n<p>BR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091614,
          "author_name": "obougacha",
          "author_url": "",
          "post_date": "11/26/2020 06:21:29",
          "content": "<p>For the NN I created several dummy models with only two layers with very few neurons. Then, each model is trained on a subset of the features and a subset of the training set (random sampling with a replacement for the training set). It's like doing bootstrapping for random forests. I hope it's more clear.</p>\n<p>I'm quite sure I'm not overfitting. My CV training errors are around 4.5 - 3.2 m and the validation errors are around 5.8 - 4.5m depending on the model. I think the values are quite acceptable and do not alarm for overfitting.</p>\n<p>My guess is that those 1500 and so entries in the test set that were detected by Adam James are the ones that causing me problems. I think their real values of time to eruption is quite different from the one that we predict with our models.  Since these points are different from any other points the model has a hard time predicting their values. And the more sophisticated and well-fitted model to the training set the worst it is doing in predicting those outliers. This is quite the only explanation I can think of. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093130,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/27/2020 13:18:06",
          "content": "<p>Hi,<br>\nNow it's more clear in my mind regarding your method. But your problem is very weird, because even if the test dataset has a bias, it cannot explain such a difference in the LB results, especially if you taking care about the validation datatest results (  I'm doing the same, with a  validation dataset size of 25%-30% of the global number of segments) . May be your model relies too much on the sensors which have important numbers of NaN values in the test dataset. <br>\nMay be in your case it can be interesting to check with your method the results on only sensors 3,4,6 and 7, because in the train/test datasets they are the sensors with quite the same number of NaN values and for these sensors the NaN is less than 10% of the total number of segments for each of them. <br>\nOn my side, still to find solutions to this overfitting problem, i have in my todo list, to check if by taking only segments which don't have NaN sensors or less than 3, the results are better. The annoying thing is that we are using the LB score to help to improbe the model. Not a very nice statistics approach, but i think we don't have the choice in this competition.</p>\n<p>Also tried the NN on my side and having bad results too. I think that the NN are too sensitive to a change of values for a sensor on the global stats to generalize correctly the model. </p>\n<p>Br</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1065900": "Hi all,\nI performed adversarial validation notebook here -> https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences . I identify a set of 1557 segments that do not have good proxies in train and I just do some views of the distributions of `time_to_eruptions` to see when the sensor failures happen. I then assume my predictions for the 1557 segments are directional correct, but remap the distribution to the distribution where sensor 10 is down in train.\n\nUsing this simple PDF remap I am able to take the predictions from my original notebook: https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline with a score in the `6.6e6` range and improve it to `5.8e6`.\n\nThe notebook only takes half an hour to run, the modeling section with mini NNs takes 5 minutes.",
    "1078497": "Thank you for your post. I generated plenty of features for each sensor. I have computed KS statistics between Train and Test sets for each feature. And I discovered this :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F275525%2Fabb57aca6d06681297d213eda1ea7b9c%2FScreenshot_20201114_224332.png?generation=1605390475333869&alt=media)\nEach point is the KS statistics for a feature. It confirms that sensor 10 is dangerous (because KS statistics are high for features from this sensor). But it is not the only one...",
    "1085633": "Has anyone made any progress on this? I've tried a number of methods to identify and remove 'bad actors' but have not had much luck on converging between the MAE on the CV and the test set. When I've reduced the features with poor distribution, it has continued to improve the MAE within cross validation but still no progress on test set.",
    "1085640": "hmm the KS stats for sensors 2, 5, 9 are not fantastic. We would hope to be < 0.05 so we can reject the null hypothesis",
    "1088579": "Hi Kyle, \n\nI tried the same at first, but no improvments were obtained by filtering some sensors. \nI tried also a reduction of the number of variables in the model. The results were fine on train set, but at the end on the LB the results were quite disappointing.\n\nCurrently i'm working on a simple data set, based on the \"INGV Volcanic : Basic solution (STFT)\" notebook and got better results on LB by exploring cross validation methods rather than trying to remove features. By the way this notebook is very interesting because it is really reducing the eruption aftershocks noise saw by using the standard indicators (sum, std, mean,....).\n\nIn my case, i was blocked around 5.5m on LB but trying to reproduce the candy box experiment ( show a transparent candy box full of candies and ask to 1 000 persons to give you an estimation of the number of candies inside, and at the end, you will notice that the mean of their predictions will converge to the real number of candies inside the box - \"wisdom of the crowd\" if think-), I've got 4.7m results now. \nThe idea, because the dataset is small when using aggregate datas on each segment, is to multiple the evaluation made by the model. So i gave up the Kfold/RepeatedKfold/.... and using now only the shuffleSplit ( splitting randomly the train set 2000 times with a repartition of 70-30 between the train and validation dataset) with aggressive values for the LGBM model ( for instance n_estimators set to 800).  It gave me a sensible improvment of the LB, because this method seems to make the outliers disapearing and helps with the overfitting issue by having a more suitable way of generalize the features estimators.\n\nThe huge default is the fact that this method is requiring lot of computation. The second default, it will not help to improve the lack of of consistency between the train and test datasets as shown by the notebook created by Adam James. SInce the beginning of the competition, i'm surprized by the test dataset size, and by the fact that the train an test dataset are not homogenous. But this is i think the key point of this competition.\n\nMay be a stack of model can help to improve also the LB ( I only used it for classification problem, and i don't know if it's working on regression problem)\n\nHope sharing my vision will help and give you ideas.",
    "1088796": "Thank you for sharing. I have not explored the training method as you suggested which I'll play around with. I've been stuck at the 5.5m on LB for awhile. \n\nI had incremental success trying to expand the training set by creating a multi-output regression model. So looking at the time_to_eruption distribution, I generated more training data  to expand the training set. It had not resulted in the improvement you had suggested training the model multiple times but for sake of computational time, I hadn't produced a significant amount of additional training data. Another option to explore.",
    "1088808": "Regarding the use of shuffleSplit, just be aware that it looks very sensitive to the features you are taking. For instance i just finished a 20hours calculations with a larger dataset than the previous one   but it gave me poor LB results at the end. I have also to learn and explore this method 😔. Just sharing the trick that helped me.",
    "1091108": "Hello, Thank you both for sharing your ideas. I thought I was going crazy and that I missed something when creating my submission file but it seems that there is quite a difference between the results of the CV and those on the LB. \n\nTo be honest, I spent quite a while trying different methods for instance I used neural networks with Tensorflow (That was not a good idea) I'm getting fairly low MAE on the CV around 4.9m but disappointingly insane results on the LB around 18.9m and 22m. I just wanted to explore the NN regressions. \n\nAlso tried the stacking part, where I created several base model from RandomForest, XGBoost, LGBM, etc.. then tried several stacking levels. I was quite excited about this method when I saw the CV 2.5m with 0.89 r2 (I was quite careful not to fall into the overfitting trap). However, I was shocked with 22m. \nSo I guess even stacking is not a great idea. \n\nFinally, I tried something quite crazy computationally. I created a bootstrapped NNs. So I made quite a 30 or 40 NN each trained with a random part of the Training part of the CV and also with a random selection of the features. (I thought to myself since we have random sensors failures at least that my hypothesis, so the model input should be robust to this). Well, this made the NN score a bit better 10.9m. Then I tried something even crazier. I used a 200% sample with a replacement method when I create the part of training for each NN sub-model. Well, It took forever to train only 40 dummy Neural Network but it managed to drop the LB to 8.7m. This got me thinking maybe a smaller part of this training set represents the test distribution and by sampling with replacement, I accidentally reinforced this part.",
    "1091230": "Hi, to be honest i didn't undertstood all the scenarios you had tried, especially for the NN part. But the values that you had encountered look like you're doing overfitting. Also, if your models (XGB, LGBM,...) doing overfitting, the stack model will do the same. May be restarting from a basic model ( for instance my preference is for this great notebook[https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft](url) will help you to restart from a good point. I'm not sure that a NN model will help regarding the datas in this competition, except if you are able to detect the eruption sequences as explained int the notebook of this post.\n\nBR",
    "1091614": "For the NN I created several dummy models with only two layers with very few neurons. Then, each model is trained on a subset of the features and a subset of the training set (random sampling with a replacement for the training set). It's like doing bootstrapping for random forests. I hope it's more clear.\n\nI'm quite sure I'm not overfitting. My CV training errors are around 4.5 - 3.2 m and the validation errors are around 5.8 - 4.5m depending on the model. I think the values are quite acceptable and do not alarm for overfitting.\n\nMy guess is that those 1500 and so entries in the test set that were detected by Adam James are the ones that causing me problems. I think their real values of time to eruption is quite different from the one that we predict with our models.  Since these points are different from any other points the model has a hard time predicting their values. And the more sophisticated and well-fitted model to the training set the worst it is doing in predicting those outliers. This is quite the only explanation I can think of.",
    "1093130": "Hi,\nNow it's more clear in my mind regarding your method. But your problem is very weird, because even if the test dataset has a bias, it cannot explain such a difference in the LB results, especially if you taking care about the validation datatest results (  I'm doing the same, with a  validation dataset size of 25%-30% of the global number of segments) . May be your model relies too much on the sensors which have important numbers of NaN values in the test dataset. \nMay be in your case it can be interesting to check with your method the results on only sensors 3,4,6 and 7, because in the train/test datasets they are the sensors with quite the same number of NaN values and for these sensors the NaN is less than 10% of the total number of segments for each of them. \nOn my side, still to find solutions to this overfitting problem, i have in my todo list, to check if by taking only segments which don't have NaN sensors or less than 3, the results are better. The annoying thing is that we are using the LB score to help to improbe the model. Not a very nice statistics approach, but i think we don't have the choice in this competition.\n\nAlso tried the NN on my side and having bad results too. I think that the NN are too sensitive to a change of values for a sensor on the global stats to generalize correctly the model. \n\nBr"
  },
  "source": "meta"
}