{
  "id": 192766,
  "title": "CV problems",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/192766",
  "author_name": "",
  "post_date": "2020-10-23T06:40:09.193162500Z",
  "votes": 10,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I noticed that for many CV does not coincide with the LB (at least in public kernels).</p>\n<p>Besides, when tuning the model, as it turned out, the score in LB did not correlate at all(</p>\n<p>Who solved this problem how, if not secret?</p>",
  "messages": [
    {
      "id": "1057918",
      "postDate": "10/23/2020 06:40:09",
      "content": "<p>I noticed that for many CV does not coincide with the LB (at least in public kernels).</p>\n<p>Besides, when tuning the model, as it turned out, the score in LB did not correlate at all(</p>\n<p>Who solved this problem how, if not secret?</p>",
      "rawMarkdown": "I noticed that for many CV does not coincide with the LB (at least in public kernels).\n\nBesides, when tuning the model, as it turned out, the score in LB did not correlate at all(\n\nWho solved this problem how, if not secret?",
      "votes": null
    },
    {
      "id": "1057937",
      "postDate": "10/23/2020 07:00:21",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/mrmorj\" target=\"_blank\">@mrmorj</a>,</p>\n<p>As you well know, the key to avoiding the shakeup in a competition is a good CV setup such that ones personal CV \\( \\approx \\) LB (if ones CV does not correspond to the LB one is usually <a href=\"https://www.kaggle.com/carlmcbrideellis/overfitting-and-underfitting-the-titanic\" target=\"_blank\">overfitting</a>), and a good part of the problem is how one prepares the data to model, and what features are selected for modelling.<br>\nHowever, in this case I must agree that this dataset is quite tricky to pin down. Indeed perhaps the famously difficult nature of this particular problem is  maybe the reason the scientists at the INGV have been offered the data to kaggle in the first place; to see what novel approaches people come up with.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @mrmorj,\n\nAs you well know, the key to avoiding the shakeup in a competition is a good CV setup such that ones personal CV \\\\( \\approx \\\\) LB (if ones CV does not correspond to the LB one is usually [overfitting](https://www.kaggle.com/carlmcbrideellis/overfitting-and-underfitting-the-titanic)), and a good part of the problem is how one prepares the data to model, and what features are selected for modelling.\nHowever, in this case I must agree that this dataset is quite tricky to pin down. Indeed perhaps the famously difficult nature of this particular problem is  maybe the reason the scientists at the INGV have been offered the data to kaggle in the first place; to see what novel approaches people come up with.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1057972",
      "postDate": "10/23/2020 07:56:46",
      "content": "<p>The same problem<br>\nMy increase/decrease a result on the leaderboard is pure random, independent of the results of cross validation</p>",
      "rawMarkdown": "The same problem\nMy increase/decrease a result on the leaderboard is pure random, independent of the results of cross validation",
      "votes": null
    },
    {
      "id": "1058468",
      "postDate": "10/23/2020 18:42:04",
      "content": "<p>What are other peoples CV -&gt; LB scores. I'm probably overfit though. Early stopping conditions on LGBM probably not working as all the data comes from one volcano<br>\nMine is 2.5e6 -&gt; 6e6</p>",
      "rawMarkdown": "What are other peoples CV -> LB scores. I'm probably overfit though. Early stopping conditions on LGBM probably not working as all the data comes from one volcano\nMine is 2.5e6 -> 6e6",
      "votes": null
    },
    {
      "id": "1058496",
      "postDate": "10/23/2020 19:23:49",
      "content": "<p>xgb 4.0e6 -&gt; 6.8e6<br>\nrf 3.8e6 -&gt; 6.5e6<br>\nkNN 2.6e6 -&gt; 6e6</p>\n<p>That would be ok, but my CV score changes are not linearly related to the LB</p>",
      "rawMarkdown": "xgb 4.0e6 -> 6.8e6\nrf 3.8e6 -> 6.5e6\nkNN 2.6e6 -> 6e6\n\n\nThat would be ok, but my CV score changes are not linearly related to the LB",
      "votes": null
    },
    {
      "id": "1061893",
      "postDate": "10/27/2020 12:12:28",
      "content": "<p>Hi,<br>\nI think that this problem of overfit has two reasons :</p>\n<ul>\n<li>The size of training set vs the size of the test set.</li>\n<li>but most important, the fact that the sumarized statistics on each segment are showing that most of 50 %  of these segments have no usefull information. <br>\nFor instance if you look the std() of each segment for each sensor ordered by time-to-eruption, you will notice that some segment far from the eruption has the same 'activity' than a segment close to the eruption. You can do the exercice on others variables, same results. For instance on std() of sensor data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ff07140b67f1afb66b24afa306fa90a06%2Fstd.png?generation=1603804594848525&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>In the actual state of art for this kind of analysis, the researchers are using time series approach, with arma/sarima,… calculations in order to take into account the underlying physical phenomena and accumulation of usefull events.</p>\n<p>Because of this behavior regarding the global statistics of segments, you will always have overfitting problem(i'm guessing that you have used only global stats on each segments). Looking only to the global segment stats will lead to a false path. The analysis of inside data of each segment is only hope to have interesting results, but due the inactivity of most of them, there is a good chance to finish also in a dead end. My vision is that the seismic activity regarding the time to eruption is a non linear problem, and cannot give good results without a time series analysis and at least a AR/MA implementation.</p>\n<p>Best regards</p>",
      "rawMarkdown": "Hi,\nI think that this problem of overfit has two reasons :\n- The size of training set vs the size of the test set.\n- but most important, the fact that the sumarized statistics on each segment are showing that most of 50 %  of these segments have no usefull information. \nFor instance if you look the std() of each segment for each sensor ordered by time-to-eruption, you will notice that some segment far from the eruption has the same 'activity' than a segment close to the eruption. You can do the exercice on others variables, same results. For instance on std() of sensor data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ff07140b67f1afb66b24afa306fa90a06%2Fstd.png?generation=1603804594848525&alt=media)\n\nIn the actual state of art for this kind of analysis, the researchers are using time series approach, with arma/sarima,... calculations in order to take into account the underlying physical phenomena and accumulation of usefull events.\n\nBecause of this behavior regarding the global statistics of segments, you will always have overfitting problem(i'm guessing that you have used only global stats on each segments). Looking only to the global segment stats will lead to a false path. The analysis of inside data of each segment is only hope to have interesting results, but due the inactivity of most of them, there is a good chance to finish also in a dead end. My vision is that the seismic activity regarding the time to eruption is a non linear problem, and cannot give good results without a time series analysis and at least a AR/MA implementation.\n\nBest regards",
      "votes": null
    },
    {
      "id": "1062536",
      "postDate": "10/28/2020 00:21:51",
      "content": "<p>My CV was headed the same way as my LB most of the time (large changes to CV).</p>\n<p>Only change I made so far which made CV go up but LB go down (so reduced gap) was dropping features where there are a lot more missing values in Test vs Train. I think Sensor 10 for example has a lot of missing values in Test. But I've not done a lot of depth so this is just an idea…</p>",
      "rawMarkdown": "My CV was headed the same way as my LB most of the time (large changes to CV).\n\nOnly change I made so far which made CV go up but LB go down (so reduced gap) was dropping features where there are a lot more missing values in Test vs Train. I think Sensor 10 for example has a lot of missing values in Test. But I've not done a lot of depth so this is just an idea...",
      "votes": null
    },
    {
      "id": "1065118",
      "postDate": "10/30/2020 21:39:10",
      "content": "<p>Trying to add to this discussion with some Adversarial Validation -&gt; <a href=\"https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a></p>\n<p>Let me know what you all think</p>",
      "rawMarkdown": "Trying to add to this discussion with some Adversarial Validation -> https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\n\nLet me know what you all think",
      "votes": null
    },
    {
      "id": "1068700",
      "postDate": "11/03/2020 16:49:30",
      "content": "<p>So big difference between CV and LB also could be explained by data splitting method to train and test datasets. If data were splitted randomly, then we should have much smaller difference between CV and LB. Probably data were splitted by eruptions between train and test dataset. Moreover test dataset also could be splitted by eruption for public and private leaderboard  and we don't know which part of them contains most broken sensors lol</p>",
      "rawMarkdown": "So big difference between CV and LB also could be explained by data splitting method to train and test datasets. If data were splitted randomly, then we should have much smaller difference between CV and LB. Probably data were splitted by eruptions between train and test dataset. Moreover test dataset also could be splitted by eruption for public and private leaderboard  and we don't know which part of them contains most broken sensors lol",
      "votes": null
    },
    {
      "id": "1068771",
      "postDate": "11/03/2020 18:16:10",
      "content": "<p>Hi,<br>\nI'm currently trying to understand these differences, but currently no suitable solutions found. For instance, I splitted the train set in two subsets ( 80%/20%) in order to simulate the problem with only the train data ( i tried to respect the repartition of segments regarding the time_to_eruption target). <br>\nThe results are quite good in predictions ( around 1.5 M ) for this \"Homemade test set\", without using hyperparameters procedure, but on the LB it's around 6M with the same model.</p>\n<p>If i'm using hyperparameters procedure recommanded parameters, i have quite stupid results:<br>\n [3545 rows x 750 columns]</p>\n<p>rmse:  2089.373865027154<br>\nrmse:  4580.725797320655<br>\nrmse:  7067.023060134235<br>\nrmse:  6092.631826368005<br>\nrmse:  2124.6503030565063<br>\nSimple LGB model rmse:  2 124<br>\nSimple test LGB model rmse:  8 969 509</p>\n<p>My feeling is that the test set is containing some eruption sequence unknown in the train set ( as demonstrated in brilliant notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a>), and that is the core of this competition: Find a common \"indicator\" to different **profiles **of eruption sequence whatever the circumstances. As mentionned by the organizator, the broken sensors problem is part of real conditions for collecting the data, and we have to deal with it. I didn't find for the moment also a suitable way for handling the problem. </p>\n<p>And as i wrote before, with this competition dataset, there is also periods of inactivity in the underlying phenomena showed by sensors, which will make overfitting inevitable with segment global statisitics approach. I will be very surprized and quite admirative with people getting a score &lt;2M, or even 3M, as the problem seems particular complicated to manage with global stats on segment.</p>\n<p>Update:<br>\nwith the following results:</p>\n<p>rmse:  1293109.1611915014<br>\nrmse:  1305693.3730671261<br>\nrmse:  1338845.0306131486<br>\nrmse:  1324890.647749942<br>\nrmse:  1283450.642537269<br>\nrmse:  1340519.5381940873<br>\nrmse:  1301327.6153771526<br>\nrmse:  1331438.680263962<br>\nrmse:  1248229.1251629512<br>\nrmse:  1174334.53324965<br>\nSimple LGB model rmse:  1174334.53324965<br>\nSimple test LGB model rmse:  6654154.783135601</p>\n<p>it gives the following results with homemade test set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ffa7d74cb6cd91694e262a34dc9df4563%2Fres.png?generation=1604432179916381&amp;alt=media\" alt=\"\"></p>\n<p>I have to test with different samples of the home made test set, but it looks like that some events after the eruption ( hypothesis that the segments in the dataset are continuous between eruption sequence) are troubleshooting the model. May be my hypothesis that there is not enough informations in some segments is false, and that the important point if the seismic event after eruption.</p>\n<p>BR</p>",
      "rawMarkdown": "Hi,\nI'm currently trying to understand these differences, but currently no suitable solutions found. For instance, I splitted the train set in two subsets ( 80%/20%) in order to simulate the problem with only the train data ( i tried to respect the repartition of segments regarding the time_to_eruption target). \nThe results are quite good in predictions ( around 1.5 M ) for this \"Homemade test set\", without using hyperparameters procedure, but on the LB it's around 6M with the same model.\n\nIf i'm using hyperparameters procedure recommanded parameters, i have quite stupid results:\n [3545 rows x 750 columns]\n\nrmse:  2089.373865027154\nrmse:  4580.725797320655\nrmse:  7067.023060134235\nrmse:  6092.631826368005\nrmse:  2124.6503030565063\nSimple LGB model rmse:  2 124\nSimple test LGB model rmse:  8 969 509\n\nMy feeling is that the test set is containing some eruption sequence unknown in the train set ( as demonstrated in brilliant notebook [https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences](url)), and that is the core of this competition: Find a common \"indicator\" to different **profiles **of eruption sequence whatever the circumstances. As mentionned by the organizator, the broken sensors problem is part of real conditions for collecting the data, and we have to deal with it. I didn't find for the moment also a suitable way for handling the problem. \n\nAnd as i wrote before, with this competition dataset, there is also periods of inactivity in the underlying phenomena showed by sensors, which will make overfitting inevitable with segment global statisitics approach. I will be very surprized and quite admirative with people getting a score <2M, or even 3M, as the problem seems particular complicated to manage with global stats on segment.\n\nUpdate:\nwith the following results:\n\nrmse:  1293109.1611915014\nrmse:  1305693.3730671261\nrmse:  1338845.0306131486\nrmse:  1324890.647749942\nrmse:  1283450.642537269\nrmse:  1340519.5381940873\nrmse:  1301327.6153771526\nrmse:  1331438.680263962\nrmse:  1248229.1251629512\nrmse:  1174334.53324965\nSimple LGB model rmse:  1174334.53324965\nSimple test LGB model rmse:  6654154.783135601\n\nit gives the following results with homemade test set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ffa7d74cb6cd91694e262a34dc9df4563%2Fres.png?generation=1604432179916381&alt=media)\n\nI have to test with different samples of the home made test set, but it looks like that some events after the eruption ( hypothesis that the segments in the dataset are continuous between eruption sequence) are troubleshooting the model. May be my hypothesis that there is not enough informations in some segments is false, and that the important point if the seismic event after eruption.\n\n\nBR",
      "votes": null
    },
    {
      "id": "1068851",
      "postDate": "11/03/2020 20:03:19",
      "content": "<p>Great chart! I'm not sure <a href=\"https://www.kaggle.com/jace2005\" target=\"_blank\">@jace2005</a> but I believe you posted about aftershocks, which could be related to the dip here?</p>\n<p>I believe that there is a sequence of multiple eruptions, at least that's my hypothesis, and therefore some measurements must be right after an eruption.</p>",
      "rawMarkdown": "Great chart! I'm not sure @jace2005 but I believe you posted about aftershocks, which could be related to the dip here?\n\nI believe that there is a sequence of multiple eruptions, at least that's my hypothesis, and therefore some measurements must be right after an eruption.",
      "votes": null
    },
    {
      "id": "1068959",
      "postDate": "11/03/2020 22:54:14",
      "content": "<p>Hi,<br>\nfor the moment my idea is to check the hypothesis that there is a confusion for the model between the sensor's datas just before the eruption and just after, and because it was just a single test, i will generate others subsets from the train set, randomly, but keeping in mind to respect the repartition of segment on the time_to_eruption value. <br>\nif the hypothesis that there are some aftershocks after the eruption is validated, i think i will try to isolate the just after/before population ( by using only the segments in a specific window) to check if there a way to distinguish them. <br>\nI think now it's the main issue, if the hypothesis is confirmed, that can explain the difference between training results and tests results. For this task, I will probably use TSFRESH, because it has a lot of advanced function that can help to do so.   <br>\nCurrently, for performance issues on my laptop, my model is quite basic: i'm using the [sensor value], [sensor value - sensor value t-1], abs[sensor value - sensor value t-1], rolling[sensor value], rolling[abs[sensor value - sensor value t-1]], declined by sum, mean, std,….</p>\n<p>There are also some peaks in the middle of the chart but my intuition is that they are related to the previous problem and missing data on sensors and for the moment they are less important than the first issue.</p>\n<p>BR</p>",
      "rawMarkdown": "Hi,\nfor the moment my idea is to check the hypothesis that there is a confusion for the model between the sensor's datas just before the eruption and just after, and because it was just a single test, i will generate others subsets from the train set, randomly, but keeping in mind to respect the repartition of segment on the time_to_eruption value. \nif the hypothesis that there are some aftershocks after the eruption is validated, i think i will try to isolate the just after/before population ( by using only the segments in a specific window) to check if there a way to distinguish them. \nI think now it's the main issue, if the hypothesis is confirmed, that can explain the difference between training results and tests results. For this task, I will probably use TSFRESH, because it has a lot of advanced function that can help to do so.   \nCurrently, for performance issues on my laptop, my model is quite basic: i'm using the [sensor value], [sensor value - sensor value t-1], abs[sensor value - sensor value t-1], rolling[sensor value], rolling[abs[sensor value - sensor value t-1]], declined by sum, mean, std,....\n\nThere are also some peaks in the middle of the chart but my intuition is that they are related to the previous problem and missing data on sensors and for the moment they are less important than the first issue.\n\nBR",
      "votes": null
    },
    {
      "id": "1069749",
      "postDate": "11/04/2020 21:33:12",
      "content": "<p>After generating 60 randomly splitted train/test set from the original train set, and aggregating by the mean of prediction, i had the following results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F5a42e0c7a4076b5598bc91921c9a5e3b%2Ffinal.png?generation=1604523888385144&amp;alt=media\" alt=\"\"></p>\n<p>The aftershocks after eruption are clearly interpreted by my small model as events before eruption. It is also affecting the prediction segments closed to the eruption. In order to identify the kind of population each segment belongs, it worth trying to create indicators for instance to compare \"the activity\" ( to be defined) between the begin and the end of a segment &amp; removing features creating this gap if they exist.</p>\n<p>Any suggestion?</p>\n<p>BR</p>",
      "rawMarkdown": "After generating 60 randomly splitted train/test set from the original train set, and aggregating by the mean of prediction, i had the following results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F5a42e0c7a4076b5598bc91921c9a5e3b%2Ffinal.png?generation=1604523888385144&alt=media)\n\nThe aftershocks after eruption are clearly interpreted by my small model as events before eruption. It is also affecting the prediction segments closed to the eruption. In order to identify the kind of population each segment belongs, it worth trying to create indicators for instance to compare \"the activity\" ( to be defined) between the begin and the end of a segment & removing features creating this gap if they exist.\n\nAny suggestion?\n\nBR",
      "votes": null
    },
    {
      "id": "1070902",
      "postDate": "11/06/2020 09:47:06",
      "content": "<p><a href=\"https://www.kaggle.com/jace2005\" target=\"_blank\">@jace2005</a> nice catch. For now I see only one way - to find good features to distinguish between normal data and data with aftershocks. I think it is good way to apply to it the same trick like here for finding difference between train and test dataset <a href=\"https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a></p>",
      "rawMarkdown": "jace2005 nice catch. For now I see only one way - to find good features to distinguish between normal data and data with aftershocks. I think it is good way to apply to it the same trick like here for finding difference between train and test dataset https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences",
      "votes": null
    },
    {
      "id": "1071764",
      "postDate": "11/07/2020 11:40:50",
      "content": "<p>Hi Sergey,<br>\nI did the a basic binary classification and the results were quite good:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F43fde33451b8ef6a740116f86b9df911%2Fmm.png?generation=1604748712994835&amp;alt=media\" alt=\"\"></p>\n<p>'1' means segments closed to the eruption ( I created the variable for segments having a time_to eruption less than 8688340). I applied the classification model with a new feature on the final Train/Test sets of the competition, but the results were worst than before. </p>\n<p>Two options at this point:</p>\n<ul>\n<li>The way i generated the classification is overfitting (again ;) ). May be i can check for a multiclass classification, balanced class, and shufftling the data for generating the classification .</li>\n<li>The way i use the new binary classification result in a feature is not good regarding the LGBM regression i'm doing.</li>\n</ul>\n<p>Keep searching….</p>",
      "rawMarkdown": "Hi Sergey,\nI did the a basic binary classification and the results were quite good:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F43fde33451b8ef6a740116f86b9df911%2Fmm.png?generation=1604748712994835&alt=media)\n\n'1' means segments closed to the eruption ( I created the variable for segments having a time_to eruption less than 8688340). I applied the classification model with a new feature on the final Train/Test sets of the competition, but the results were worst than before. \n\nTwo options at this point:\n- The way i generated the classification is overfitting (again ;) ). May be i can check for a multiclass classification, balanced class, and shufftling the data for generating the classification .\n- The way i use the new binary classification result in a feature is not good regarding the LGBM regression i'm doing.\n\nKeep searching....",
      "votes": null
    },
    {
      "id": "1078007",
      "postDate": "11/14/2020 08:01:20",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/ajcostarino\" target=\"_blank\">@ajcostarino</a>,</p>\n<p>4.13 M -&gt; 6.64 M using the new experimental scikit-learn <a href=\"https://www.kaggle.com/carlmcbrideellis/histogram-gradient-boosting-regression-example\" target=\"_blank\">histogram gradient boosting regressor</a> with 10-fold cross-validation (run time: 70 seconds) applied to your <a href=\"https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline\" target=\"_blank\"><em>\"LGBM Baseline\"</em></a> data.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @ajcostarino,\n\n4.13 M -> 6.64 M using the new experimental scikit-learn [histogram gradient boosting regressor](https://www.kaggle.com/carlmcbrideellis/histogram-gradient-boosting-regression-example) with 10-fold cross-validation (run time: 70 seconds) applied to your [*\"LGBM Baseline\"*](https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline) data.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1079119",
      "postDate": "11/15/2020 16:29:18",
      "content": "<p>Just an update.</p>\n<p>After reading the <a href=\"https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a> notebook, i decided to use the results datasets to check the behavior.</p>\n<p>Using the train/test sets created from the original train set, i had the following results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd404e8d0a808a6ac5c2aa8430abb97ec%2Ftft.png?generation=1605456510445095&amp;alt=media\" alt=\"\"></p>\n<p>Simple LGB model rmse:  1 698 735<br>\nSimple test LGB model rmse:  3 599 313</p>\n<p>The issue with the aftershocks disappeared which is a rather good news, but the predictions trend is more horizontal than the real eruption time. </p>\n<p>So i tried to add a binary classification ( segment time eruption&lt; 24 000 000 and &gt;24 000 000) and i add the following results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F56b3c37478f27907f15390e1a2420158%2Fbinary.png?generation=1605456764742946&amp;alt=media\" alt=\"\"><br>\nSimple LGB model rmse:  1 102 708<br>\nSimple test LGB model rmse:  2 787 922</p>\n<p>then i tried a multiclass classification ( creating 3 classes : 0-16M,16M-32M,32M-48M) and i obtained this result:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd1c8caff6d818ed507216f40201ca89a%2Fmuti.png?generation=1605456951654452&amp;alt=media\" alt=\"\"><br>\nSimple LGB model rmse:  954 478<br>\nSimple test LGB model rmse:  3 009 237</p>\n<p>That's annoying because the multiclass new variable allows in the model to generate better predictions for most of segments, but at the same time is creating unexpected peaks.</p>\n<p>The peaks are frequently close to the limit of each class. Does someone have any idea to fix the problem of peak with the introduction of a classification variable in the dataset ?</p>\n<p>Fyi the multi class results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fa7cfee3b9ceae3635ee743e25a3e68af%2Fclass.png?generation=1605457673012051&amp;alt=media\" alt=\"\"></p>\n<p>BR</p>",
      "rawMarkdown": "Just an update.\n\nAfter reading the https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft notebook, i decided to use the results datasets to check the behavior.\n\nUsing the train/test sets created from the original train set, i had the following results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd404e8d0a808a6ac5c2aa8430abb97ec%2Ftft.png?generation=1605456510445095&alt=media)\n\nSimple LGB model rmse:  1 698 735\nSimple test LGB model rmse:  3 599 313\n\nThe issue with the aftershocks disappeared which is a rather good news, but the predictions trend is more horizontal than the real eruption time. \n\nSo i tried to add a binary classification ( segment time eruption< 24 000 000 and >24 000 000) and i add the following results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F56b3c37478f27907f15390e1a2420158%2Fbinary.png?generation=1605456764742946&alt=media)\nSimple LGB model rmse:  1 102 708\nSimple test LGB model rmse:  2 787 922\n\nthen i tried a multiclass classification ( creating 3 classes : 0-16M,16M-32M,32M-48M) and i obtained this result:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd1c8caff6d818ed507216f40201ca89a%2Fmuti.png?generation=1605456951654452&alt=media)\nSimple LGB model rmse:  954 478\nSimple test LGB model rmse:  3 009 237\n\nThat's annoying because the multiclass new variable allows in the model to generate better predictions for most of segments, but at the same time is creating unexpected peaks.\n\nThe peaks are frequently close to the limit of each class. Does someone have any idea to fix the problem of peak with the introduction of a classification variable in the dataset ?\n\nFyi the multi class results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fa7cfee3b9ceae3635ee743e25a3e68af%2Fclass.png?generation=1605457673012051&alt=media)\n\n\nBR",
      "votes": null
    },
    {
      "id": "1079124",
      "postDate": "11/15/2020 16:37:21",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/jace2005\" target=\"_blank\">@jace2005</a>,</p>\n<p>I am sure you know this but just in case; the submissions in this particular competition are evaluated using the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.mean_absolute_error.html\" target=\"_blank\"><em>mean absolute error</em> (MAE)</a> between the predicted loss and the actual loss, not the root mean squared error (RMSE).</p>\n<p>All the best, <br>\ncarl</p>",
      "rawMarkdown": "Dear @jace2005,\n\nI am sure you know this but just in case; the submissions in this particular competition are evaluated using the [*mean absolute error* (MAE)](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.mean_absolute_error.html) between the predicted loss and the actual loss, not the root mean squared error (RMSE).\n\nAll the best, \ncarl",
      "votes": null
    },
    {
      "id": "1079168",
      "postDate": "11/15/2020 17:50:56",
      "content": "<p>Thanks for the remark.<br>\nIt's an error in my script…😔</p>",
      "rawMarkdown": "Thanks for the remark.\nIt's an error in my script...😔",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1057937,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "10/23/2020 07:00:21",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/mrmorj\" target=\"_blank\">@mrmorj</a>,</p>\n<p>As you well know, the key to avoiding the shakeup in a competition is a good CV setup such that ones personal CV \\( \\approx \\) LB (if ones CV does not correspond to the LB one is usually <a href=\"https://www.kaggle.com/carlmcbrideellis/overfitting-and-underfitting-the-titanic\" target=\"_blank\">overfitting</a>), and a good part of the problem is how one prepares the data to model, and what features are selected for modelling.<br>\nHowever, in this case I must agree that this dataset is quite tricky to pin down. Indeed perhaps the famously difficult nature of this particular problem is  maybe the reason the scientists at the INGV have been offered the data to kaggle in the first place; to see what novel approaches people come up with.</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1057972,
      "author_name": "batyazhizni",
      "author_url": "",
      "post_date": "10/23/2020 07:56:46",
      "content": "<p>The same problem<br>\nMy increase/decrease a result on the leaderboard is pure random, independent of the results of cross validation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1058468,
      "author_name": "ajcostarino",
      "author_url": "",
      "post_date": "10/23/2020 18:42:04",
      "content": "<p>What are other peoples CV -&gt; LB scores. I'm probably overfit though. Early stopping conditions on LGBM probably not working as all the data comes from one volcano<br>\nMine is 2.5e6 -&gt; 6e6</p>",
      "votes": null,
      "replies": [
        {
          "id": 1058496,
          "author_name": "batyazhizni",
          "author_url": "",
          "post_date": "10/23/2020 19:23:49",
          "content": "<p>xgb 4.0e6 -&gt; 6.8e6<br>\nrf 3.8e6 -&gt; 6.5e6<br>\nkNN 2.6e6 -&gt; 6e6</p>\n<p>That would be ok, but my CV score changes are not linearly related to the LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1078007,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "11/14/2020 08:01:20",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/ajcostarino\" target=\"_blank\">@ajcostarino</a>,</p>\n<p>4.13 M -&gt; 6.64 M using the new experimental scikit-learn <a href=\"https://www.kaggle.com/carlmcbrideellis/histogram-gradient-boosting-regression-example\" target=\"_blank\">histogram gradient boosting regressor</a> with 10-fold cross-validation (run time: 70 seconds) applied to your <a href=\"https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline\" target=\"_blank\"><em>\"LGBM Baseline\"</em></a> data.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1061893,
      "author_name": "jace2005",
      "author_url": "",
      "post_date": "10/27/2020 12:12:28",
      "content": "<p>Hi,<br>\nI think that this problem of overfit has two reasons :</p>\n<ul>\n<li>The size of training set vs the size of the test set.</li>\n<li>but most important, the fact that the sumarized statistics on each segment are showing that most of 50 %  of these segments have no usefull information. <br>\nFor instance if you look the std() of each segment for each sensor ordered by time-to-eruption, you will notice that some segment far from the eruption has the same 'activity' than a segment close to the eruption. You can do the exercice on others variables, same results. For instance on std() of sensor data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ff07140b67f1afb66b24afa306fa90a06%2Fstd.png?generation=1603804594848525&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>In the actual state of art for this kind of analysis, the researchers are using time series approach, with arma/sarima,… calculations in order to take into account the underlying physical phenomena and accumulation of usefull events.</p>\n<p>Because of this behavior regarding the global statistics of segments, you will always have overfitting problem(i'm guessing that you have used only global stats on each segments). Looking only to the global segment stats will lead to a false path. The analysis of inside data of each segment is only hope to have interesting results, but due the inactivity of most of them, there is a good chance to finish also in a dead end. My vision is that the seismic activity regarding the time to eruption is a non linear problem, and cannot give good results without a time series analysis and at least a AR/MA implementation.</p>\n<p>Best regards</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1062536,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "10/28/2020 00:21:51",
      "content": "<p>My CV was headed the same way as my LB most of the time (large changes to CV).</p>\n<p>Only change I made so far which made CV go up but LB go down (so reduced gap) was dropping features where there are a lot more missing values in Test vs Train. I think Sensor 10 for example has a lot of missing values in Test. But I've not done a lot of depth so this is just an idea…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065118,
      "author_name": "ajcostarino",
      "author_url": "",
      "post_date": "10/30/2020 21:39:10",
      "content": "<p>Trying to add to this discussion with some Adversarial Validation -&gt; <a href=\"https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a></p>\n<p>Let me know what you all think</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1068700,
      "author_name": "sergeypm",
      "author_url": "",
      "post_date": "11/03/2020 16:49:30",
      "content": "<p>So big difference between CV and LB also could be explained by data splitting method to train and test datasets. If data were splitted randomly, then we should have much smaller difference between CV and LB. Probably data were splitted by eruptions between train and test dataset. Moreover test dataset also could be splitted by eruption for public and private leaderboard  and we don't know which part of them contains most broken sensors lol</p>",
      "votes": null,
      "replies": [
        {
          "id": 1068771,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/03/2020 18:16:10",
          "content": "<p>Hi,<br>\nI'm currently trying to understand these differences, but currently no suitable solutions found. For instance, I splitted the train set in two subsets ( 80%/20%) in order to simulate the problem with only the train data ( i tried to respect the repartition of segments regarding the time_to_eruption target). <br>\nThe results are quite good in predictions ( around 1.5 M ) for this \"Homemade test set\", without using hyperparameters procedure, but on the LB it's around 6M with the same model.</p>\n<p>If i'm using hyperparameters procedure recommanded parameters, i have quite stupid results:<br>\n [3545 rows x 750 columns]</p>\n<p>rmse:  2089.373865027154<br>\nrmse:  4580.725797320655<br>\nrmse:  7067.023060134235<br>\nrmse:  6092.631826368005<br>\nrmse:  2124.6503030565063<br>\nSimple LGB model rmse:  2 124<br>\nSimple test LGB model rmse:  8 969 509</p>\n<p>My feeling is that the test set is containing some eruption sequence unknown in the train set ( as demonstrated in brilliant notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a>), and that is the core of this competition: Find a common \"indicator\" to different **profiles **of eruption sequence whatever the circumstances. As mentionned by the organizator, the broken sensors problem is part of real conditions for collecting the data, and we have to deal with it. I didn't find for the moment also a suitable way for handling the problem. </p>\n<p>And as i wrote before, with this competition dataset, there is also periods of inactivity in the underlying phenomena showed by sensors, which will make overfitting inevitable with segment global statisitics approach. I will be very surprized and quite admirative with people getting a score &lt;2M, or even 3M, as the problem seems particular complicated to manage with global stats on segment.</p>\n<p>Update:<br>\nwith the following results:</p>\n<p>rmse:  1293109.1611915014<br>\nrmse:  1305693.3730671261<br>\nrmse:  1338845.0306131486<br>\nrmse:  1324890.647749942<br>\nrmse:  1283450.642537269<br>\nrmse:  1340519.5381940873<br>\nrmse:  1301327.6153771526<br>\nrmse:  1331438.680263962<br>\nrmse:  1248229.1251629512<br>\nrmse:  1174334.53324965<br>\nSimple LGB model rmse:  1174334.53324965<br>\nSimple test LGB model rmse:  6654154.783135601</p>\n<p>it gives the following results with homemade test set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ffa7d74cb6cd91694e262a34dc9df4563%2Fres.png?generation=1604432179916381&amp;alt=media\" alt=\"\"></p>\n<p>I have to test with different samples of the home made test set, but it looks like that some events after the eruption ( hypothesis that the segments in the dataset are continuous between eruption sequence) are troubleshooting the model. May be my hypothesis that there is not enough informations in some segments is false, and that the important point if the seismic event after eruption.</p>\n<p>BR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068851,
          "author_name": "ajcostarino",
          "author_url": "",
          "post_date": "11/03/2020 20:03:19",
          "content": "<p>Great chart! I'm not sure <a href=\"https://www.kaggle.com/jace2005\" target=\"_blank\">@jace2005</a> but I believe you posted about aftershocks, which could be related to the dip here?</p>\n<p>I believe that there is a sequence of multiple eruptions, at least that's my hypothesis, and therefore some measurements must be right after an eruption.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068959,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/03/2020 22:54:14",
          "content": "<p>Hi,<br>\nfor the moment my idea is to check the hypothesis that there is a confusion for the model between the sensor's datas just before the eruption and just after, and because it was just a single test, i will generate others subsets from the train set, randomly, but keeping in mind to respect the repartition of segment on the time_to_eruption value. <br>\nif the hypothesis that there are some aftershocks after the eruption is validated, i think i will try to isolate the just after/before population ( by using only the segments in a specific window) to check if there a way to distinguish them. <br>\nI think now it's the main issue, if the hypothesis is confirmed, that can explain the difference between training results and tests results. For this task, I will probably use TSFRESH, because it has a lot of advanced function that can help to do so.   <br>\nCurrently, for performance issues on my laptop, my model is quite basic: i'm using the [sensor value], [sensor value - sensor value t-1], abs[sensor value - sensor value t-1], rolling[sensor value], rolling[abs[sensor value - sensor value t-1]], declined by sum, mean, std,….</p>\n<p>There are also some peaks in the middle of the chart but my intuition is that they are related to the previous problem and missing data on sensors and for the moment they are less important than the first issue.</p>\n<p>BR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069749,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/04/2020 21:33:12",
          "content": "<p>After generating 60 randomly splitted train/test set from the original train set, and aggregating by the mean of prediction, i had the following results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F5a42e0c7a4076b5598bc91921c9a5e3b%2Ffinal.png?generation=1604523888385144&amp;alt=media\" alt=\"\"></p>\n<p>The aftershocks after eruption are clearly interpreted by my small model as events before eruption. It is also affecting the prediction segments closed to the eruption. In order to identify the kind of population each segment belongs, it worth trying to create indicators for instance to compare \"the activity\" ( to be defined) between the begin and the end of a segment &amp; removing features creating this gap if they exist.</p>\n<p>Any suggestion?</p>\n<p>BR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1070902,
          "author_name": "sergeypm",
          "author_url": "",
          "post_date": "11/06/2020 09:47:06",
          "content": "<p><a href=\"https://www.kaggle.com/jace2005\" target=\"_blank\">@jace2005</a> nice catch. For now I see only one way - to find good features to distinguish between normal data and data with aftershocks. I think it is good way to apply to it the same trick like here for finding difference between train and test dataset <a href=\"https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\" target=\"_blank\">https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1071764,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/07/2020 11:40:50",
          "content": "<p>Hi Sergey,<br>\nI did the a basic binary classification and the results were quite good:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F43fde33451b8ef6a740116f86b9df911%2Fmm.png?generation=1604748712994835&amp;alt=media\" alt=\"\"></p>\n<p>'1' means segments closed to the eruption ( I created the variable for segments having a time_to eruption less than 8688340). I applied the classification model with a new feature on the final Train/Test sets of the competition, but the results were worst than before. </p>\n<p>Two options at this point:</p>\n<ul>\n<li>The way i generated the classification is overfitting (again ;) ). May be i can check for a multiclass classification, balanced class, and shufftling the data for generating the classification .</li>\n<li>The way i use the new binary classification result in a feature is not good regarding the LGBM regression i'm doing.</li>\n</ul>\n<p>Keep searching….</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079119,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/15/2020 16:29:18",
          "content": "<p>Just an update.</p>\n<p>After reading the <a href=\"https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a> notebook, i decided to use the results datasets to check the behavior.</p>\n<p>Using the train/test sets created from the original train set, i had the following results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd404e8d0a808a6ac5c2aa8430abb97ec%2Ftft.png?generation=1605456510445095&amp;alt=media\" alt=\"\"></p>\n<p>Simple LGB model rmse:  1 698 735<br>\nSimple test LGB model rmse:  3 599 313</p>\n<p>The issue with the aftershocks disappeared which is a rather good news, but the predictions trend is more horizontal than the real eruption time. </p>\n<p>So i tried to add a binary classification ( segment time eruption&lt; 24 000 000 and &gt;24 000 000) and i add the following results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F56b3c37478f27907f15390e1a2420158%2Fbinary.png?generation=1605456764742946&amp;alt=media\" alt=\"\"><br>\nSimple LGB model rmse:  1 102 708<br>\nSimple test LGB model rmse:  2 787 922</p>\n<p>then i tried a multiclass classification ( creating 3 classes : 0-16M,16M-32M,32M-48M) and i obtained this result:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd1c8caff6d818ed507216f40201ca89a%2Fmuti.png?generation=1605456951654452&amp;alt=media\" alt=\"\"><br>\nSimple LGB model rmse:  954 478<br>\nSimple test LGB model rmse:  3 009 237</p>\n<p>That's annoying because the multiclass new variable allows in the model to generate better predictions for most of segments, but at the same time is creating unexpected peaks.</p>\n<p>The peaks are frequently close to the limit of each class. Does someone have any idea to fix the problem of peak with the introduction of a classification variable in the dataset ?</p>\n<p>Fyi the multi class results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fa7cfee3b9ceae3635ee743e25a3e68af%2Fclass.png?generation=1605457673012051&amp;alt=media\" alt=\"\"></p>\n<p>BR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079124,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "11/15/2020 16:37:21",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/jace2005\" target=\"_blank\">@jace2005</a>,</p>\n<p>I am sure you know this but just in case; the submissions in this particular competition are evaluated using the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.mean_absolute_error.html\" target=\"_blank\"><em>mean absolute error</em> (MAE)</a> between the predicted loss and the actual loss, not the root mean squared error (RMSE).</p>\n<p>All the best, <br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079168,
          "author_name": "jace2005",
          "author_url": "",
          "post_date": "11/15/2020 17:50:56",
          "content": "<p>Thanks for the remark.<br>\nIt's an error in my script…😔</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1057918": "I noticed that for many CV does not coincide with the LB (at least in public kernels).\n\nBesides, when tuning the model, as it turned out, the score in LB did not correlate at all(\n\nWho solved this problem how, if not secret?",
    "1057937": "Dear @mrmorj,\n\nAs you well know, the key to avoiding the shakeup in a competition is a good CV setup such that ones personal CV \\\\( \\approx \\\\) LB (if ones CV does not correspond to the LB one is usually [overfitting](https://www.kaggle.com/carlmcbrideellis/overfitting-and-underfitting-the-titanic)), and a good part of the problem is how one prepares the data to model, and what features are selected for modelling.\nHowever, in this case I must agree that this dataset is quite tricky to pin down. Indeed perhaps the famously difficult nature of this particular problem is  maybe the reason the scientists at the INGV have been offered the data to kaggle in the first place; to see what novel approaches people come up with.\n\nAll the best,\ncarl",
    "1057972": "The same problem\nMy increase/decrease a result on the leaderboard is pure random, independent of the results of cross validation",
    "1058468": "What are other peoples CV -> LB scores. I'm probably overfit though. Early stopping conditions on LGBM probably not working as all the data comes from one volcano\nMine is 2.5e6 -> 6e6",
    "1058496": "xgb 4.0e6 -> 6.8e6\nrf 3.8e6 -> 6.5e6\nkNN 2.6e6 -> 6e6\n\n\nThat would be ok, but my CV score changes are not linearly related to the LB",
    "1061893": "Hi,\nI think that this problem of overfit has two reasons :\n- The size of training set vs the size of the test set.\n- but most important, the fact that the sumarized statistics on each segment are showing that most of 50 %  of these segments have no usefull information. \nFor instance if you look the std() of each segment for each sensor ordered by time-to-eruption, you will notice that some segment far from the eruption has the same 'activity' than a segment close to the eruption. You can do the exercice on others variables, same results. For instance on std() of sensor data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ff07140b67f1afb66b24afa306fa90a06%2Fstd.png?generation=1603804594848525&alt=media)\n\nIn the actual state of art for this kind of analysis, the researchers are using time series approach, with arma/sarima,... calculations in order to take into account the underlying physical phenomena and accumulation of usefull events.\n\nBecause of this behavior regarding the global statistics of segments, you will always have overfitting problem(i'm guessing that you have used only global stats on each segments). Looking only to the global segment stats will lead to a false path. The analysis of inside data of each segment is only hope to have interesting results, but due the inactivity of most of them, there is a good chance to finish also in a dead end. My vision is that the seismic activity regarding the time to eruption is a non linear problem, and cannot give good results without a time series analysis and at least a AR/MA implementation.\n\nBest regards",
    "1062536": "My CV was headed the same way as my LB most of the time (large changes to CV).\n\nOnly change I made so far which made CV go up but LB go down (so reduced gap) was dropping features where there are a lot more missing values in Test vs Train. I think Sensor 10 for example has a lot of missing values in Test. But I've not done a lot of depth so this is just an idea...",
    "1065118": "Trying to add to this discussion with some Adversarial Validation -> https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences\n\nLet me know what you all think",
    "1068700": "So big difference between CV and LB also could be explained by data splitting method to train and test datasets. If data were splitted randomly, then we should have much smaller difference between CV and LB. Probably data were splitted by eruptions between train and test dataset. Moreover test dataset also could be splitted by eruption for public and private leaderboard  and we don't know which part of them contains most broken sensors lol",
    "1068771": "Hi,\nI'm currently trying to understand these differences, but currently no suitable solutions found. For instance, I splitted the train set in two subsets ( 80%/20%) in order to simulate the problem with only the train data ( i tried to respect the repartition of segments regarding the time_to_eruption target). \nThe results are quite good in predictions ( around 1.5 M ) for this \"Homemade test set\", without using hyperparameters procedure, but on the LB it's around 6M with the same model.\n\nIf i'm using hyperparameters procedure recommanded parameters, i have quite stupid results:\n [3545 rows x 750 columns]\n\nrmse:  2089.373865027154\nrmse:  4580.725797320655\nrmse:  7067.023060134235\nrmse:  6092.631826368005\nrmse:  2124.6503030565063\nSimple LGB model rmse:  2 124\nSimple test LGB model rmse:  8 969 509\n\nMy feeling is that the test set is containing some eruption sequence unknown in the train set ( as demonstrated in brilliant notebook [https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences](url)), and that is the core of this competition: Find a common \"indicator\" to different **profiles **of eruption sequence whatever the circumstances. As mentionned by the organizator, the broken sensors problem is part of real conditions for collecting the data, and we have to deal with it. I didn't find for the moment also a suitable way for handling the problem. \n\nAnd as i wrote before, with this competition dataset, there is also periods of inactivity in the underlying phenomena showed by sensors, which will make overfitting inevitable with segment global statisitics approach. I will be very surprized and quite admirative with people getting a score <2M, or even 3M, as the problem seems particular complicated to manage with global stats on segment.\n\nUpdate:\nwith the following results:\n\nrmse:  1293109.1611915014\nrmse:  1305693.3730671261\nrmse:  1338845.0306131486\nrmse:  1324890.647749942\nrmse:  1283450.642537269\nrmse:  1340519.5381940873\nrmse:  1301327.6153771526\nrmse:  1331438.680263962\nrmse:  1248229.1251629512\nrmse:  1174334.53324965\nSimple LGB model rmse:  1174334.53324965\nSimple test LGB model rmse:  6654154.783135601\n\nit gives the following results with homemade test set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Ffa7d74cb6cd91694e262a34dc9df4563%2Fres.png?generation=1604432179916381&alt=media)\n\nI have to test with different samples of the home made test set, but it looks like that some events after the eruption ( hypothesis that the segments in the dataset are continuous between eruption sequence) are troubleshooting the model. May be my hypothesis that there is not enough informations in some segments is false, and that the important point if the seismic event after eruption.\n\n\nBR",
    "1068851": "Great chart! I'm not sure @jace2005 but I believe you posted about aftershocks, which could be related to the dip here?\n\nI believe that there is a sequence of multiple eruptions, at least that's my hypothesis, and therefore some measurements must be right after an eruption.",
    "1068959": "Hi,\nfor the moment my idea is to check the hypothesis that there is a confusion for the model between the sensor's datas just before the eruption and just after, and because it was just a single test, i will generate others subsets from the train set, randomly, but keeping in mind to respect the repartition of segment on the time_to_eruption value. \nif the hypothesis that there are some aftershocks after the eruption is validated, i think i will try to isolate the just after/before population ( by using only the segments in a specific window) to check if there a way to distinguish them. \nI think now it's the main issue, if the hypothesis is confirmed, that can explain the difference between training results and tests results. For this task, I will probably use TSFRESH, because it has a lot of advanced function that can help to do so.   \nCurrently, for performance issues on my laptop, my model is quite basic: i'm using the [sensor value], [sensor value - sensor value t-1], abs[sensor value - sensor value t-1], rolling[sensor value], rolling[abs[sensor value - sensor value t-1]], declined by sum, mean, std,....\n\nThere are also some peaks in the middle of the chart but my intuition is that they are related to the previous problem and missing data on sensors and for the moment they are less important than the first issue.\n\nBR",
    "1069749": "After generating 60 randomly splitted train/test set from the original train set, and aggregating by the mean of prediction, i had the following results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F5a42e0c7a4076b5598bc91921c9a5e3b%2Ffinal.png?generation=1604523888385144&alt=media)\n\nThe aftershocks after eruption are clearly interpreted by my small model as events before eruption. It is also affecting the prediction segments closed to the eruption. In order to identify the kind of population each segment belongs, it worth trying to create indicators for instance to compare \"the activity\" ( to be defined) between the begin and the end of a segment & removing features creating this gap if they exist.\n\nAny suggestion?\n\nBR",
    "1070902": "jace2005 nice catch. For now I see only one way - to find good features to distinguish between normal data and data with aftershocks. I think it is good way to apply to it the same trick like here for finding difference between train and test dataset https://www.kaggle.com/ajcostarino/ignv-adversarial-validation-cv-lb-differences",
    "1071764": "Hi Sergey,\nI did the a basic binary classification and the results were quite good:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F43fde33451b8ef6a740116f86b9df911%2Fmm.png?generation=1604748712994835&alt=media)\n\n'1' means segments closed to the eruption ( I created the variable for segments having a time_to eruption less than 8688340). I applied the classification model with a new feature on the final Train/Test sets of the competition, but the results were worst than before. \n\nTwo options at this point:\n- The way i generated the classification is overfitting (again ;) ). May be i can check for a multiclass classification, balanced class, and shufftling the data for generating the classification .\n- The way i use the new binary classification result in a feature is not good regarding the LGBM regression i'm doing.\n\nKeep searching....",
    "1078007": "Dear @ajcostarino,\n\n4.13 M -> 6.64 M using the new experimental scikit-learn [histogram gradient boosting regressor](https://www.kaggle.com/carlmcbrideellis/histogram-gradient-boosting-regression-example) with 10-fold cross-validation (run time: 70 seconds) applied to your [*\"LGBM Baseline\"*](https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline) data.\n\nAll the best,\ncarl",
    "1079119": "Just an update.\n\nAfter reading the https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft notebook, i decided to use the results datasets to check the behavior.\n\nUsing the train/test sets created from the original train set, i had the following results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd404e8d0a808a6ac5c2aa8430abb97ec%2Ftft.png?generation=1605456510445095&alt=media)\n\nSimple LGB model rmse:  1 698 735\nSimple test LGB model rmse:  3 599 313\n\nThe issue with the aftershocks disappeared which is a rather good news, but the predictions trend is more horizontal than the real eruption time. \n\nSo i tried to add a binary classification ( segment time eruption< 24 000 000 and >24 000 000) and i add the following results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2F56b3c37478f27907f15390e1a2420158%2Fbinary.png?generation=1605456764742946&alt=media)\nSimple LGB model rmse:  1 102 708\nSimple test LGB model rmse:  2 787 922\n\nthen i tried a multiclass classification ( creating 3 classes : 0-16M,16M-32M,32M-48M) and i obtained this result:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fd1c8caff6d818ed507216f40201ca89a%2Fmuti.png?generation=1605456951654452&alt=media)\nSimple LGB model rmse:  954 478\nSimple test LGB model rmse:  3 009 237\n\nThat's annoying because the multiclass new variable allows in the model to generate better predictions for most of segments, but at the same time is creating unexpected peaks.\n\nThe peaks are frequently close to the limit of each class. Does someone have any idea to fix the problem of peak with the introduction of a classification variable in the dataset ?\n\nFyi the multi class results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1346183%2Fa7cfee3b9ceae3635ee743e25a3e68af%2Fclass.png?generation=1605457673012051&alt=media)\n\n\nBR",
    "1079124": "Dear @jace2005,\n\nI am sure you know this but just in case; the submissions in this particular competition are evaluated using the [*mean absolute error* (MAE)](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.mean_absolute_error.html) between the predicted loss and the actual loss, not the root mean squared error (RMSE).\n\nAll the best, \ncarl",
    "1079168": "Thanks for the remark.\nIt's an error in my script...😔"
  },
  "source": "meta"
}