{
  "id": 92679,
  "title": "No Magic",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92679",
  "author_name": "CPMP",
  "post_date": "2019-05-19T09:00:07.397000",
  "votes": 91,
  "comment_count": 96,
  "views": 0,
  "content": "<p>Some people pinged me about what could be my magic features, given I get good results with less than 15 features (single LGB at 1.288 as of now).  Well, removing my best feature according to lgb worsens my CV score by 0.023, which would most probably translate to same LB score evolution.  Therefore it is not a magic feature.  But it still is an interesting feature.  When we plot it on train data along target, we see that it has 16 peaks corresponding to the high acoustic peaks before EQ, but not much peaks for the mini EQ that are hard to ignore.  I re-scaled the feature value to make the picture nicer.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533443/13245/best.png\" alt=\"plot\">\nI use this kind of plot to select my features.</p>",
  "messages": [
    {
      "id": 533443,
      "postDate": "2019-05-19T09:00:07.397Z",
      "content": "<p>Some people pinged me about what could be my magic features, given I get good results with less than 15 features (single LGB at 1.288 as of now).  Well, removing my best feature according to lgb worsens my CV score by 0.023, which would most probably translate to same LB score evolution.  Therefore it is not a magic feature.  But it still is an interesting feature.  When we plot it on train data along target, we see that it has 16 peaks corresponding to the high acoustic peaks before EQ, but not much peaks for the mini EQ that are hard to ignore.  I re-scaled the feature value to make the picture nicer.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533443/13245/best.png\" alt=\"plot\">\nI use this kind of plot to select my features.</p>",
      "rawMarkdown": "Some people pinged me about what could be my magic features, given I get good results with less than 15 features (single LGB at 1.288 as of now).  Well, removing my best feature according to lgb worsens my CV score by 0.023, which would most probably translate to same LB score evolution.  Therefore it is not a magic feature.  But it still is an interesting feature.  When we plot it on train data along target, we see that it has 16 peaks corresponding to the high acoustic peaks before EQ, but not much peaks for the mini EQ that are hard to ignore.  I re-scaled the feature value to make the picture nicer.\n\n![plot](https://storage.googleapis.com/kaggle-forum-message-attachments/533443/13245/best.png)\nI use this kind of plot to select my features.\n\n",
      "votes": 91
    },
    {
      "id": 533623,
      "postDate": "2019-05-19T16:23:59.053Z",
      "content": "<p>Nice feature indeed. It's great to predict those middle peaks right of course, but the segments with minor quakes are quite few. The problem is with the systematic error they cause especially before the mini-quake. Here is a figure of my absolute error on the oof predictions to show the case.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533623/13252/ooferror.png\" alt=\"\"></p>",
      "rawMarkdown": "Nice feature indeed. It's great to predict those middle peaks right of course, but the segments with minor quakes are quite few. The problem is with the systematic error they cause especially before the mini-quake. Here is a figure of my absolute error on the oof predictions to show the case.\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/533623/13252/ooferror.png)",
      "votes": 5,
      "replies": [
        {
          "id": 534074,
          "postDate": "2019-05-20T15:51:19.693Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 534428,
          "postDate": "2019-05-21T08:34:18.447Z",
          "content": "<p>He has plot the error, not the predictions.</p>",
          "rawMarkdown": "He has plot the error, not the predictions.",
          "votes": 1
        },
        {
          "id": 534542,
          "postDate": "2019-05-21T12:35:36.170Z",
          "content": "<p>Since doing manual feature selection to improve random forest is a good idea, it looks like a good idea to find the week learners by hand/visually.</p>\n\n<p>Nice plot. Sorry for not being able to contribute.</p>",
          "rawMarkdown": "Since doing manual feature selection to improve random forest is a good idea, it looks like a good idea to find the week learners by hand/visually.\n\nNice plot. Sorry for not being able to contribute.",
          "votes": 1
        },
        {
          "id": 534560,
          "postDate": "2019-05-21T13:12:35.450Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 535796,
          "postDate": "2019-05-23T13:12:18.743Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 535799,
          "postDate": "2019-05-23T13:18:03.140Z",
          "content": "<p>Sorry I don't remember. Must be in the 1.90s.</p>",
          "rawMarkdown": "Sorry I don't remember. Must be in the 1.90s."
        },
        {
          "id": 535901,
          "postDate": "2019-05-23T15:40:20.813Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 533706,
      "postDate": "2019-05-19T19:16:47.257Z",
      "content": "<p>congratulations. Do you want to share this feature? i think my model can improve with it 😈</p>",
      "rawMarkdown": "congratulations. Do you want to share this feature? i think my model can improve with it 😈",
      "votes": 3,
      "replies": [
        {
          "id": 534032,
          "postDate": "2019-05-20T13:38:45.337Z",
          "content": "<p>How much are you ready to offer for it ? ;)</p>",
          "rawMarkdown": "How much are you ready to offer for it ? ;)",
          "votes": 3
        }
      ]
    },
    {
      "id": 533543,
      "postDate": "2019-05-19T13:12:24.543Z",
      "content": "<p>Also, this is also good according to my feeling...\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533543/13249/2.png\" alt=\"\"></p>",
      "rawMarkdown": "Also, this is also good according to my feeling...\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/533543/13249/2.png)",
      "votes": 3,
      "replies": [
        {
          "id": 533567,
          "postDate": "2019-05-19T14:16:47.950Z",
          "content": "<p>I have similar ones, but they have peaks in between EQ.</p>",
          "rawMarkdown": "I have similar ones, but they have peaks in between EQ.",
          "votes": 3
        }
      ]
    },
    {
      "id": 533537,
      "postDate": "2019-05-19T13:05:03.927Z",
      "content": "<p>This is my best feature.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533537/13248/1.png\" alt=\"\"></p>",
      "rawMarkdown": "This is my best feature.\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/533537/13248/1.png)",
      "votes": 3
    },
    {
      "id": 533812,
      "postDate": "2019-05-20T05:07:27.657Z",
      "content": "<p><img src=\"https://i.ytimg.com/vi/uh4dLo7T2Ow/maxresdefault.jpg\" alt=\"a kind of magic\"></p>",
      "rawMarkdown": "![a kind of magic](https://i.ytimg.com/vi/uh4dLo7T2Ow/maxresdefault.jpg)",
      "votes": 4
    },
    {
      "id": 534368,
      "postDate": "2019-05-21T06:15:43.453Z",
      "content": "<p>I've been looking at this picture all afternoon but I still can't think of it</p>",
      "rawMarkdown": "I've been looking at this picture all afternoon but I still can't think of it",
      "votes": 3
    },
    {
      "id": 534900,
      "postDate": "2019-05-22T02:05:39.313Z",
      "content": "<p>Thanks for sharing <a href=\"/cpmpml\">@cpmpml</a> . This made me to work again on this competition.</p>",
      "rawMarkdown": "Thanks for sharing @cpmpml . This made me to work again on this competition.",
      "votes": 1
    },
    {
      "id": 534713,
      "postDate": "2019-05-21T18:23:24.703Z",
      "content": "<p>It's awesome!</p>",
      "rawMarkdown": "It's awesome!",
      "votes": 1
    },
    {
      "id": 534451,
      "postDate": "2019-05-21T09:09:57.043Z",
      "content": "<p>Hi <a href=\"/cpmpml\">@cpmpml</a> . First of all, thanks for your insight. But what is correlation coefficient between this feature and target?</p>",
      "rawMarkdown": "Hi @cpmpml . First of all, thanks for your insight. But what is correlation coefficient between this feature and target?",
      "votes": 1,
      "replies": [
        {
          "id": 534453,
          "postDate": "2019-05-21T09:13:53.150Z",
          "content": "<p>I don't compute correlation between features and target.</p>",
          "rawMarkdown": "I don't compute correlation between features and target.",
          "votes": 3
        }
      ]
    },
    {
      "id": 534386,
      "postDate": "2019-05-21T06:54:11.897Z",
      "content": "<p>Wow its amazing</p>",
      "rawMarkdown": "Wow its amazing",
      "votes": 1
    },
    {
      "id": 534106,
      "postDate": "2019-05-20T17:35:15.763Z",
      "content": "<p>Can you explain how much these features extent on Leaderboard score? </p>",
      "rawMarkdown": "Can you explain how much these features extent on Leaderboard score? ",
      "votes": 1,
      "replies": [
        {
          "id": 534126,
          "postDate": "2019-05-20T18:12:11.287Z",
          "content": "<p>All I know is that removing this feature degrades my CV by about 0.02.  I think LB would degrade the same, but I will not burn a submission to check it.</p>",
          "rawMarkdown": "All I know is that removing this feature degrades my CV by about 0.02.  I think LB would degrade the same, but I will not burn a submission to check it."
        }
      ]
    },
    {
      "id": 533593,
      "postDate": "2019-05-19T15:08:41.863Z",
      "content": "<p>This is my first experience with ML, python and competitions. So i ask to more experienced programmers.\nHave I a feature similar to yours? In my opinion it give some informations around minor quakes.\nI have try to build a model around this feature, but for now i do not perform well in the leaderboard :-)</p>",
      "rawMarkdown": "This is my first experience with ML, python and competitions. So i ask to more experienced programmers.\nHave I a feature similar to yours? In my opinion it give some informations around minor quakes.\nI have try to build a model around this feature, but for now i do not perform well in the leaderboard :-)",
      "votes": 1
    },
    {
      "id": 539799,
      "postDate": "2019-05-30T14:39:42.430Z",
      "content": "<p>I've plotted 295 features here: <a href=\"https://www.kaggle.com/felipefonte99/plotting-295-features-vs-time-to-failure\">https://www.kaggle.com/felipefonte99/plotting-295-features-vs-time-to-failure</a></p>",
      "rawMarkdown": "I've plotted 295 features here: https://www.kaggle.com/felipefonte99/plotting-295-features-vs-time-to-failure",
      "votes": 2
    },
    {
      "id": 539346,
      "postDate": "2019-05-30T00:12:21.307Z",
      "content": "<p>Just for curiosity, how <code>plt.scatter(train.time_to_failure, train[feature], s=1)</code> would look like for this feature?</p>",
      "rawMarkdown": "Just for curiosity, how `plt.scatter(train.time_to_failure, train[feature], s=1)` would look like for this feature?",
      "votes": 2,
      "replies": [
        {
          "id": 539356,
          "postDate": "2019-05-30T00:45:46.920Z",
          "content": "<p>here it is:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/539356/13320/feat.png\" alt=\"plot\"></p>",
          "rawMarkdown": "here it is:\n\n![plot](https://storage.googleapis.com/kaggle-forum-message-attachments/539356/13320/feat.png)",
          "votes": 3
        },
        {
          "id": 539848,
          "postDate": "2019-05-30T15:48:17.530Z",
          "content": "<p>Thanks for this useful plot. The leftmost  side of this plot like most of my useful features is causing a great part of model errors.  Is it safe to assume ttf &lt; 0.23 are outliers?</p>",
          "rawMarkdown": "Thanks for this useful plot. The leftmost  side of this plot like most of my useful features is causing a great part of model errors.  Is it safe to assume ttf &lt; 0.23 are outliers?"
        },
        {
          "id": 540021,
          "postDate": "2019-05-30T21:47:17.647Z",
          "content": "<p>Most acoustic spikes are within ttf 0.310 - 0.315 and 150000 samples duration is around 0.035s. So I ignore ttf &lt; 0.275 to still capture those spikes. Otherwise model will always consider acoustic spike as the ones from false earthquakes, which are far from 0 ttf.</p>\n\n<p>I'm more conserned with right side of the plot, which still shows multiple parallel trends</p>",
          "rawMarkdown": "Most acoustic spikes are within ttf 0.310 - 0.315 and 150000 samples duration is around 0.035s. So I ignore ttf &lt; 0.275 to still capture those spikes. Otherwise model will always consider acoustic spike as the ones from false earthquakes, which are far from 0 ttf.\n\nI'm more conserned with right side of the plot, which still shows multiple parallel trends"
        },
        {
          "id": 540024,
          "postDate": "2019-05-30T21:57:04.563Z",
          "content": "<p>While that may seem like a good idea, we have no idea whether test samples include times &lt; 0.275 like you are suggesting. In fact we do not even know if test samples take the back end of an EQ with the front end of another EQ and use that as a random test sample.</p>",
          "rawMarkdown": "While that may seem like a good idea, we have no idea whether test samples include times &lt; 0.275 like you are suggesting. In fact we do not even know if test samples take the back end of an EQ with the front end of another EQ and use that as a random test sample."
        },
        {
          "id": 540052,
          "postDate": "2019-05-31T00:13:24.707Z",
          "content": "<p>Well we do know that test data has about 10 chunks with extremely hi acoustic spikes and even have spikes <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/82336#latest-515978\">higher than all training data</a> Which are all very likely to have nearly 0.315 ttf</p>",
          "rawMarkdown": "Well we do know that test data has about 10 chunks with extremely hi acoustic spikes and even have spikes [higher than all training data](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/82336#latest-515978) Which are all very likely to have nearly 0.315 ttf",
          "votes": 1
        }
      ]
    },
    {
      "id": 534611,
      "postDate": "2019-05-21T14:41:36.100Z",
      "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> this is very inspiring. Makes me want to join the competition :)</p>",
      "rawMarkdown": "Thanks @cpmpml this is very inspiring. Makes me want to join the competition :)",
      "votes": 2,
      "replies": [
        {
          "id": 534622,
          "postDate": "2019-05-21T14:59:04.223Z",
          "content": "<p>It's a little bit late, isn't it? Obviously, I don't doubt you'll end up high ;) What will you focus on for these  two weeks?</p>",
          "rawMarkdown": "It's a little bit late, isn't it? Obviously, I don't doubt you'll end up high ;) What will you focus on for these  two weeks?",
          "votes": 2
        },
        {
          "id": 534625,
          "postDate": "2019-05-21T15:02:38.277Z",
          "content": "<blockquote>\n  <p>It's a little bit late, isn't it? </p>\n</blockquote>\n\n<p>Don't underestimate Giba or other top kagglers, 13 days with a small dataset like this can be enough to get a good result.</p>",
          "rawMarkdown": "&gt; It's a little bit late, isn't it? \n\nDon't underestimate Giba or other top kagglers, 13 days with a small dataset like this can be enough to get a good result.",
          "votes": 6
        },
        {
          "id": 534686,
          "postDate": "2019-05-21T17:26:15.090Z",
          "content": "<p><a href=\"/titericz\">@titericz</a> my model would benefit from a professional touch if you want to join and have a head start </p>",
          "rawMarkdown": "@titericz my model would benefit from a professional touch if you want to join and have a head start ",
          "votes": 4
        },
        {
          "id": 534693,
          "postDate": "2019-05-21T17:37:33.013Z",
          "content": "<p>It seems easy to hire a grandmaster at 17th place 🥇😄. But what to do at the 100th place? 😞</p>",
          "rawMarkdown": "It seems easy to hire a grandmaster at 17th place 🥇😄. But what to do at the 100th place? 😞",
          "votes": 1
        },
        {
          "id": 534694,
          "postDate": "2019-05-21T17:41:27.670Z",
          "content": "<p>pray</p>",
          "rawMarkdown": "pray",
          "votes": 8
        },
        {
          "id": 534726,
          "postDate": "2019-05-21T18:33:14.680Z",
          "content": "<p><a href=\"/sggpls\">@sggpls</a>  I don't think it's easy :D</p>",
          "rawMarkdown": "@sggpls  I don't think it's easy :D",
          "votes": 2
        },
        {
          "id": 535247,
          "postDate": "2019-05-22T15:10:30.797Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 534043,
      "postDate": "2019-05-20T14:02:35.843Z",
      "content": "<p>This is the plot of my best feature with TTF, maybe I'd better do more feature selection and change my CV strategy.....<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/534043/13257/best.PNG\" alt=\"\"></p>",
      "rawMarkdown": "This is the plot of my best feature with TTF, maybe I'd better do more feature selection and change my CV strategy.....![](https://storage.googleapis.com/kaggle-forum-message-attachments/534043/13257/best.PNG)",
      "votes": 2,
      "replies": [
        {
          "id": 534082,
          "postDate": "2019-05-20T16:16:19.857Z",
          "content": "<p>looks interesting.</p>",
          "rawMarkdown": "looks interesting.",
          "votes": 1
        }
      ]
    },
    {
      "id": 533455,
      "postDate": "2019-05-19T09:31:53.313Z",
      "content": "<p>I wouldn't call such a feature \"interesting\". It is rather a super feature. However, I have found that in this competition some good features work much better in pairs, so I wish you to find a feature for the pair.</p>",
      "rawMarkdown": "I wouldn't call such a feature \"interesting\". It is rather a super feature. However, I have found that in this competition some good features work much better in pairs, so I wish you to find a feature for the pair.",
      "votes": 2,
      "replies": [
        {
          "id": 533528,
          "postDate": "2019-05-19T12:46:22.447Z",
          "content": "<p>It probably works well with other features I have indeed. I have not looked into feature interactions, that's not something I do but I know I probably should.</p>",
          "rawMarkdown": "It probably works well with other features I have indeed. I have not looked into feature interactions, that's not something I do but I know I probably should.",
          "votes": 1
        }
      ]
    },
    {
      "id": 533465,
      "postDate": "2019-05-19T09:49:59.163Z",
      "content": "<p>You have a typo here? ‘’worsens my CV score by 0.023, which would most probably translate to same CV score”</p>",
      "rawMarkdown": "You have a typo here? ‘’worsens my CV score by 0.023, which would most probably translate to same CV score”",
      "votes": 1,
      "replies": [
        {
          "id": 533527,
          "postDate": "2019-05-19T12:45:20.350Z",
          "content": "<p>Thanks, I fixed the typo.</p>",
          "rawMarkdown": "Thanks, I fixed the typo."
        }
      ]
    },
    {
      "id": 543075,
      "postDate": "2019-06-04T10:12:14.127Z",
      "content": "<p>Hi, <a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>First of all, I want to thank you. You have been an inspiration for me to learn machine learning.</p>\n\n<p>Here, is the plot of the mean feature of a signal.</p>\n\n<p><img src=\"https://drive.google.com/file/d/1PBWruz3bnpK8yXIQQrgP4jYY_tUmV4iZ/view\" alt=\"Graph\"></p>\n\n<p>Don't know, here the image is not displayed.</p>\n\n<p>Image Link :  \"&gt;https://drive.google.com/open?id=1uzwv9eTtjGNOyWPZWeGNgCIz2qaP3O_</p>\n\n<p>The intimation from the plot is that it's not a good feature for prediction.\nBut, This feature took 3rd place as in graph of feature importance of trained model.\nHow could we decide which feature is important for the model?\nAgain the correlation between this feature and target is quite low.</p>\n\n<p>Please take a look.\nCan you guide me to know how you recognize the important feature from plotting?</p>\n\n<p>Thanks a lot.</p>",
      "rawMarkdown": "Hi, @cpmpml \n\nFirst of all, I want to thank you. You have been an inspiration for me to learn machine learning.\n\nHere, is the plot of the mean feature of a signal.\n\n![Graph](https://drive.google.com/file/d/1PBWruz3bnpK8yXIQQrgP4jYY_tUmV4iZ/view)\n\nDon't know, here the image is not displayed.\n\nImage Link :  https://drive.google.com/open?id=1uzwv9eTtjGNOyWPZWeGNgCIz2qaP3O__\n\nThe intimation from the plot is that it's not a good feature for prediction.\nBut, This feature took 3rd place as in graph of feature importance of trained model.\nHow could we decide which feature is important for the model?\nAgain the correlation between this feature and target is quite low.\n\nPlease take a look.\nCan you guide me to know how you recognize the important feature from plotting?\n\nThanks a lot.",
      "replies": [
        {
          "id": 543109,
          "postDate": "2019-06-04T10:32:57.070Z",
          "content": "<p>I use plot to select features, but I further select by training models.  If models are better with it, then i keep the feature.  Mean is a special case here, it is highly discriminant between train and test and we decided not to use it for that reason.</p>",
          "rawMarkdown": "I use plot to select features, but I further select by training models.  If models are better with it, then i keep the feature.  Mean is a special case here, it is highly discriminant between train and test and we decided not to use it for that reason.",
          "votes": 1
        },
        {
          "id": 543125,
          "postDate": "2019-06-04T10:43:46.790Z",
          "content": "<p>Thanks for the reply..</p>",
          "rawMarkdown": "Thanks for the reply.."
        }
      ]
    },
    {
      "id": 540981,
      "postDate": "2019-06-01T13:06:09.207Z",
      "content": "<p>I tried to plot them, this may help\n<a href=\"https://www.kaggle.com/vinayaks/feature-selection-simplified\">https://www.kaggle.com/vinayaks/feature-selection-simplified</a></p>",
      "rawMarkdown": "I tried to plot them, this may help\nhttps://www.kaggle.com/vinayaks/feature-selection-simplified"
    },
    {
      "id": 538475,
      "postDate": "2019-05-28T15:59:21.383Z",
      "content": "<p>Great job.   What's your training and cross validation mae ?</p>",
      "rawMarkdown": "Great job.   What's your training and cross validation mae ?"
    },
    {
      "id": 538021,
      "postDate": "2019-05-28T02:31:55.243Z",
      "content": "<p>I got something alike.</p>",
      "rawMarkdown": "I got something alike."
    },
    {
      "id": 537801,
      "postDate": "2019-05-27T16:03:30.250Z",
      "content": "<p>Though my opinion may not matter as your team is already on top of LB which means you are definitely heading in right direction. But since you are one of my favorite GM I would like to ask something. Is it not obvious that earthquakes are not periodic? If you see the slope of decay of <code>time = t</code> the next EQ occurs, shouldn't it be decaying like a negative exponential?</p>\n\n<p><img src=\"https://www.physics.uoguelph.ca/tutorials/exp/graph5.gif\" alt=\"\"></p>\n\n<p>using image to show the decay it should be, its not based on dataset</p>",
      "rawMarkdown": "Though my opinion may not matter as your team is already on top of LB which means you are definitely heading in right direction. But since you are one of my favorite GM I would like to ask something. Is it not obvious that earthquakes are not periodic? If you see the slope of decay of `time = t` the next EQ occurs, shouldn't it be decaying like a negative exponential?\n\n\n![](https://www.physics.uoguelph.ca/tutorials/exp/graph5.gif)\n\nusing image to show the decay it should be, its not based on dataset",
      "replies": [
        {
          "id": 538221,
          "postDate": "2019-05-28T10:10:41.830Z",
          "content": "<blockquote>\n  <p>Is it not obvious that earthquakes are not periodic?</p>\n</blockquote>\n\n<p>I think organizers wrote in one of their paper or in the material available here that EQ are not periodic indeed.</p>",
          "rawMarkdown": "&gt;  Is it not obvious that earthquakes are not periodic?\n\nI think organizers wrote in one of their paper or in the material available here that EQ are not periodic indeed.",
          "votes": 1
        },
        {
          "id": 538389,
          "postDate": "2019-05-28T14:08:39.367Z",
          "content": "<p>From the introduction post ( <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525</a> ):</p>\n\n<blockquote>\n  <p>For this challenge we selected an experiment that exhibits a very aperiodic and more realistic behavior compared to the data we studied in our early work, with earthquakes occurring very irregularly. </p>\n</blockquote>",
          "rawMarkdown": "From the introduction post ( https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525 ):\n\n&gt; For this challenge we selected an experiment that exhibits a very aperiodic and more realistic behavior compared to the data we studied in our early work, with earthquakes occurring very irregularly. \n"
        }
      ]
    },
    {
      "id": 535983,
      "postDate": "2019-05-23T18:28:17.607Z",
      "content": "<p>I got a similar structure using Mean Squared Percentage Error: (A-X)/A\nBut I haven't the skill to use it!!\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/535983/13282/MSPA.png\" alt=\"MSPE\"></p>\n\n<p>So I am guessing it is some percentile feature - cannot wait to find out!</p>",
      "rawMarkdown": "I got a similar structure using Mean Squared Percentage Error: (A-X)/A\nBut I haven't the skill to use it!!\n![MSPE](https://storage.googleapis.com/kaggle-forum-message-attachments/535983/13282/MSPA.png)\n\nSo I am guessing it is some percentile feature - cannot wait to find out!"
    },
    {
      "id": 535329,
      "postDate": "2019-05-22T17:51:26.067Z",
      "content": "<p>Did you augment or balance your data?</p>",
      "rawMarkdown": "Did you augment or balance your data?",
      "replies": [
        {
          "id": 535339,
          "postDate": "2019-05-22T18:09:01.690Z",
          "content": "<p>I'll answer that after competition end ;)</p>",
          "rawMarkdown": "I'll answer that after competition end ;)",
          "votes": 2
        }
      ]
    },
    {
      "id": 535256,
      "postDate": "2019-05-22T15:20:52.120Z",
      "content": "<p>Thanks a lot, this is really insightful !\nI have a question about the difference distribution between train and test data. Did you adjust train data distribution to become similar to test's and then selected featres? Or your features are already general to detect same meaning in train and test? </p>",
      "rawMarkdown": "Thanks a lot, this is really insightful !\nI have a question about the difference distribution between train and test data. Did you adjust train data distribution to become similar to test's and then selected featres? Or your features are already general to detect same meaning in train and test? ",
      "replies": [
        {
          "id": 535260,
          "postDate": "2019-05-22T15:22:40.600Z",
          "content": "<p>I'll answer that after competition end ;)</p>",
          "rawMarkdown": "I'll answer that after competition end ;)",
          "votes": 2
        }
      ]
    },
    {
      "id": 533949,
      "postDate": "2019-05-20T10:09:00.643Z",
      "content": "<p>Sadly I already have this feature, but my LB score is stuck at 1.484. I am a total newbie though,(6th day into machine learning)  learning new things everyday. :)\nP.S: I am looking for a mentor :p</p>",
      "rawMarkdown": "Sadly I already have this feature, but my LB score is stuck at 1.484. I am a total newbie though,(6th day into machine learning)  learning new things everyday. :)\nP.S: I am looking for a mentor :p",
      "replies": [
        {
          "id": 534031,
          "postDate": "2019-05-20T13:38:24.690Z",
          "content": "<blockquote>\n  <p>Sadly I already have this feature</p>\n</blockquote>\n\n<p>How do you know it is the same feature?  I frankly doubt it is, given how I constructed it.  You may have a similar one, but the odds of using the exact same one are very little IMHO.</p>\n\n<p>Can you plot you feature the same way I did and share here?</p>",
          "rawMarkdown": "&gt; Sadly I already have this feature\n\nHow do you know it is the same feature?  I frankly doubt it is, given how I constructed it.  You may have a similar one, but the odds of using the exact same one are very little IMHO.\n\nCan you plot you feature the same way I did and share here?",
          "votes": 2
        },
        {
          "id": 534674,
          "postDate": "2019-05-21T16:51:04.790Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> you are right my feature is indeed quite different. Looked similar at first glance. I ll try to replicate your pattern you posted here. Thanks for the hint. </p>",
          "rawMarkdown": "@cpmpml you are right my feature is indeed quite different. Looked similar at first glance. I ll try to replicate your pattern you posted here. Thanks for the hint. "
        }
      ]
    },
    {
      "id": 533944,
      "postDate": "2019-05-20T09:41:00.887Z",
      "content": "<p>How about this feature? Is this good?</p>",
      "rawMarkdown": "How about this feature? Is this good?",
      "replies": [
        {
          "id": 533948,
          "postDate": "2019-05-20T10:03:21.747Z",
          "content": "<p>How would I know?  We can't see the values except for few peaks.</p>",
          "rawMarkdown": "How would I know?  We can't see the values except for few peaks."
        },
        {
          "id": 533952,
          "postDate": "2019-05-20T10:14:42.777Z",
          "content": "<p>I think the other values are 0.\nI mean feature that shows just the EQ could be useful?</p>",
          "rawMarkdown": "I think the other values are 0.\nI mean feature that shows just the EQ could be useful?"
        },
        {
          "id": 533957,
          "postDate": "2019-05-20T10:25:14.037Z",
          "content": "<p>just the peak is of not much value, it may help predict 16 points at most.</p>",
          "rawMarkdown": "just the peak is of not much value, it may help predict 16 points at most.",
          "votes": 1
        }
      ]
    },
    {
      "id": 533777,
      "postDate": "2019-05-20T02:22:26.687Z",
      "content": "<p>Can i ask you how do you generate those plots for the features? No for this one in particular, i mean in general. Thanks</p>",
      "rawMarkdown": "Can i ask you how do you generate those plots for the features? No for this one in particular, i mean in general. Thanks",
      "replies": [
        {
          "id": 533794,
          "postDate": "2019-05-20T04:08:43.407Z",
          "content": "<p>In a notebook:</p>\n\n<pre><code>from matplotlib import pyplot as plt\n%matploltlib inline\nfig, ax = plt.subplots(1, 1, figsize=(12, 6))\nax.plot(train.time_to_failure)\nax.plot(train[feature])\n</code></pre>",
          "rawMarkdown": "In a notebook:\n\n    from matplotlib import pyplot as plt\n    %matploltlib inline\n    fig, ax = plt.subplots(1, 1, figsize=(12, 6))\n    ax.plot(train.time_to_failure)\n    ax.plot(train[feature])",
          "votes": 6
        },
        {
          "id": 534034,
          "postDate": "2019-05-20T13:51:58.077Z",
          "content": "<p>Thank you, this will be helpful</p>",
          "rawMarkdown": "Thank you, this will be helpful"
        }
      ]
    },
    {
      "id": 533621,
      "postDate": "2019-05-19T16:12:40.353Z",
      "content": "<p>That is corref?</p>",
      "rawMarkdown": "That is corref?",
      "replies": [
        {
          "id": 534430,
          "postDate": "2019-05-21T08:35:30.680Z",
          "content": "<p>What is corref?</p>",
          "rawMarkdown": "What is corref?"
        },
        {
          "id": 534543,
          "postDate": "2019-05-21T12:36:33.707Z",
          "content": "<p>Sorry my bad. </p>",
          "rawMarkdown": "Sorry my bad. "
        }
      ]
    },
    {
      "id": 533512,
      "postDate": "2019-05-19T12:03:49.310Z",
      "content": "<p>Is your \"single LGBM\" the average of many folds?</p>",
      "rawMarkdown": "Is your \"single LGBM\" the average of many folds?",
      "replies": [
        {
          "id": 533530,
          "postDate": "2019-05-19T12:49:31.463Z",
          "content": "<p>You should look at what works best for you, averaging fold models predictions or retrain on full data.  It does not make much difference in general, but maybe it would in your case.  And to be clear, I would call 'averaging out of fold prediction' a 'single model' as long as same code and features are used to compute all oofs.</p>",
          "rawMarkdown": "You should look at what works best for you, averaging fold models predictions or retrain on full data.  It does not make much difference in general, but maybe it would in your case.  And to be clear, I would call 'averaging out of fold prediction' a 'single model' as long as same code and features are used to compute all oofs.",
          "votes": 6
        },
        {
          "id": 533909,
          "postDate": "2019-05-20T08:29:28.950Z",
          "content": "<p>Thanks, the part after 'to be clear' was interesting for me.  I had an impression some people call this averaging 'single model'.</p>",
          "rawMarkdown": "Thanks, the part after 'to be clear' was interesting for me.  I had an impression some people call this averaging 'single model'."
        },
        {
          "id": 534442,
          "postDate": "2019-05-21T08:50:17.733Z",
          "content": "<p>Some people even call an average of several runs with different random seeds 'a single model'.</p>",
          "rawMarkdown": "Some people even call an average of several runs with different random seeds 'a single model'.",
          "votes": 1
        },
        {
          "id": 535950,
          "postDate": "2019-05-23T17:16:55.097Z",
          "content": "<p>I didn't know that. It may be that I am using a single model oO</p>",
          "rawMarkdown": "I didn't know that. It may be that I am using a single model oO",
          "votes": 1
        }
      ]
    },
    {
      "id": 534267,
      "postDate": "2019-05-21T02:14:29.290Z",
      "rawMarkdown": "",
      "votes": 10,
      "isDeleted": true,
      "replies": [
        {
          "id": 534382,
          "postDate": "2019-05-21T06:46:19.600Z",
          "content": "<p>My pipeline is what you wrote.  Yes, in some cases adding more than one feature at once can be better.  It is all a matter of trial and error.  Machine learning is an experimental science.   You perform experiments (eg adding one or more features), and you learn form the result of the experience.</p>\n\n<p>For your 5 features I would apply the pipeline above ;)</p>",
          "rawMarkdown": "My pipeline is what you wrote.  Yes, in some cases adding more than one feature at once can be better.  It is all a matter of trial and error.  Machine learning is an experimental science.   You perform experiments (eg adding one or more features), and you learn form the result of the experience.\n\nFor your 5 features I would apply the pipeline above ;)",
          "votes": 9
        }
      ]
    },
    {
      "id": 534016,
      "postDate": "2019-05-20T12:58:12.033Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true,
      "replies": [
        {
          "id": 534017,
          "postDate": "2019-05-20T13:03:02.783Z",
          "content": "<ol>\n<li><p>I don't care about feature correlation at all.  All I care about is how they impact CV score.  I just tried a number of features.  When you have a small number of features then testing their effect on CV score takes little time.</p></li>\n<li><p>I do tweak features to improve them.  I know hat some kagglers get new feature ideas by looking at feature interactions, but I'm not doing this, yet.  Maybe I should.</p></li>\n<li><p>I could not find additional features that helps, which is why I stick to 14 or less for now.  I welcome ideas for new features ;)</p></li>\n</ol>",
          "rawMarkdown": "1.  I don't care about feature correlation at all.  All I care about is how they impact CV score.  I just tried a number of features.  When you have a small number of features then testing their effect on CV score takes little time.\n\n2. I do tweak features to improve them.  I know hat some kagglers get new feature ideas by looking at feature interactions, but I'm not doing this, yet.  Maybe I should.\n\n3. I could not find additional features that helps, which is why I stick to 14 or less for now.  I welcome ideas for new features ;)\n",
          "votes": 8
        },
        {
          "id": 534024,
          "postDate": "2019-05-20T13:22:28.083Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 534030,
          "postDate": "2019-05-20T13:37:02.727Z",
          "content": "<blockquote>\n  <p>Will you re-tune hyperparamaters of the model with each additional feature? </p>\n</blockquote>\n\n<p>Absolutely not, this is bad for many reasons.</p>\n\n<ol>\n<li>It prevents fair comparison with previous runs</li>\n<li>Parameter tuning is the best way i know to overfit.  The less you do it, the better IMHO.  I usually do it only twice during a competition, first after a couple of weeks, and then few weeks before end when I have a fairly stable set of features.</li>\n</ol>\n\n<p>I am not sure I understand your second question correctly.  To me, the definition of good features is that adding them lowers CV score.</p>",
          "rawMarkdown": "&gt; Will you re-tune hyperparamaters of the model with each additional feature? \n\nAbsolutely not, this is bad for many reasons.\n\n1. It prevents fair comparison with previous runs\n2. Parameter tuning is the best way i know to overfit.  The less you do it, the better IMHO.  I usually do it only twice during a competition, first after a couple of weeks, and then few weeks before end when I have a fairly stable set of features.\n\nI am not sure I understand your second question correctly.  To me, the definition of good features is that adding them lowers CV score.",
          "votes": 8
        },
        {
          "id": 534049,
          "postDate": "2019-05-20T14:25:15.580Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 534055,
          "postDate": "2019-05-20T14:39:42.137Z",
          "content": "<p>I wrote you what I do ;)  I can only tell you to try and see what works best for you.  I don't own the definitive truth here ;)</p>",
          "rawMarkdown": "I wrote you what I do ;)  I can only tell you to try and see what works best for you.  I don't own the definitive truth here ;)",
          "votes": 2
        },
        {
          "id": 534130,
          "postDate": "2019-05-20T18:15:42.870Z",
          "content": "<p>My answer was too cryptic maybe.  If you need to tune parameters to make new features look good then I think they aren't that good.</p>",
          "rawMarkdown": "My answer was too cryptic maybe.  If you need to tune parameters to make new features look good then I think they aren't that good.",
          "votes": 2
        },
        {
          "id": 534165,
          "postDate": "2019-05-20T20:21:09.270Z",
          "content": "<p>Keep downvoting and you won’t need to read me.</p>",
          "rawMarkdown": "Keep downvoting and you won’t need to read me.",
          "votes": 3
        },
        {
          "id": 534205,
          "postDate": "2019-05-20T23:17:04.537Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Blast! You have foiled my dastardly scheme to ruin other competitors' submissions by providing them with an easy way to tune their parameters! I appear to have largely avoided my own trap since I generally use the default params for feature selection. Perhaps I should add a caveat to my kernel advising against its overuse. </p>",
          "rawMarkdown": "@cpmpml Blast! You have foiled my dastardly scheme to ruin other competitors' submissions by providing them with an easy way to tune their parameters! I appear to have largely avoided my own trap since I generally use the default params for feature selection. Perhaps I should add a caveat to my kernel advising against its overuse. ",
          "votes": 3
        },
        {
          "id": 535129,
          "postDate": "2019-05-22T10:55:37.063Z",
          "content": "<p>Feature impact on CV depends on CV method. Some particular feature can improve CV score on shuffled k-fold but make it worse on eq-wise CV. How do you deal with that?</p>",
          "rawMarkdown": "Feature impact on CV depends on CV method. Some particular feature can improve CV score on shuffled k-fold but make it worse on eq-wise CV. How do you deal with that?",
          "votes": 1
        },
        {
          "id": 535145,
          "postDate": "2019-05-22T11:24:27.417Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> \nYou've said, you always(?) add a single feature and test its impact on CV score. I would be anxious to miss <strong>interactions of two features that can explain a lot, although both on their own explain next to nothing</strong>. (E.g. think of two features X1 and X2, both with values uniformly distributed in [-1,1] and their impact on the target being X1*X2.)</p>\n\n<p>What's your opinion on adding a lot of features, if not all you can think of, to the model, fit it, and then look at feature importance (or probably better at <a href=\"https://github.com/slundberg/shap/blob/master/README.md\">SHAP values</a>). Only include those features that the model considers important. Or do you think, that the risk of overfitting and, hence, selecting the wrong features is too high - even if you apply regularisation?</p>",
          "rawMarkdown": "@cpmpml \nYou've said, you always(?) add a single feature and test its impact on CV score. I would be anxious to miss **interactions of two features that can explain a lot, although both on their own explain next to nothing**. (E.g. think of two features X1 and X2, both with values uniformly distributed in [-1,1] and their impact on the target being X1*X2.)\n\nWhat's your opinion on adding a lot of features, if not all you can think of, to the model, fit it, and then look at feature importance (or probably better at [SHAP values](https://github.com/slundberg/shap/blob/master/README.md)). Only include those features that the model considers important. Or do you think, that the risk of overfitting and, hence, selecting the wrong features is too high - even if you apply regularisation?"
        },
        {
          "id": 535196,
          "postDate": "2019-05-22T13:19:25.517Z",
          "content": "<blockquote>\n  <p>What's your opinion on adding a lot of features</p>\n</blockquote>\n\n<p>Given the small number of samples I will never add lots of features ;)</p>\n\n<p>In other, large, datasets, I can easily use hundreds of features.  And in that case, I can evaluate them group by group indeed. I never used SHAP so far but it is on my general todo list.  Here I don't think I'll need it.</p>",
          "rawMarkdown": "&gt; What's your opinion on adding a lot of features\n\nGiven the small number of samples I will never add lots of features ;)\n\nIn other, large, datasets, I can easily use hundreds of features.  And in that case, I can evaluate them group by group indeed. I never used SHAP so far but it is on my general todo list.  Here I don't think I'll need it.",
          "votes": 1
        },
        {
          "id": 535237,
          "postDate": "2019-05-22T14:52:53.170Z",
          "content": "<blockquote>\n  <p>Or do you think, that the risk of overfitting and, hence, selecting the wrong features is too high - even if you apply regularisation?</p>\n</blockquote>\n\n<p>So, I guess, your answer is \"Yes, it's a too high risk, also with regularisation (for our small dataset).\" And you rather risk to not include such features I described above. Or in other words: The risk of overfitting is higher than that of missing such an interaction.</p>",
          "rawMarkdown": "&gt; Or do you think, that the risk of overfitting and, hence, selecting the wrong features is too high - even if you apply regularisation?\n\nSo, I guess, your answer is \"Yes, it's a too high risk, also with regularisation (for our small dataset).\" And you rather risk to not include such features I described above. Or in other words: The risk of overfitting is higher than that of missing such an interaction."
        },
        {
          "id": 535250,
          "postDate": "2019-05-22T15:16:57.197Z",
          "content": "<blockquote>\n  <p>Feature impact on CV depends on CV method. Some particular feature can improve CV score on shuffled k-fold but make it worse on eq-wise CV. How do you deal with that?</p>\n</blockquote>\n\n<p>I use my CV strategy, I spent a full week on designing it when I entered the competition.  I cannot comment on other CV strategy.  </p>",
          "rawMarkdown": "&gt; Feature impact on CV depends on CV method. Some particular feature can improve CV score on shuffled k-fold but make it worse on eq-wise CV. How do you deal with that?\n\nI use my CV strategy, I spent a full week on designing it when I entered the competition.  I cannot comment on other CV strategy.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 533626,
      "postDate": "2019-05-19T16:32:48.863Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 534534,
          "postDate": "2019-05-21T12:18:26.580Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 534539,
          "postDate": "2019-05-21T12:29:13.890Z",
          "content": "<p>Nice plot. Congratulations.</p>",
          "rawMarkdown": "Nice plot. Congratulations."
        }
      ]
    },
    {
      "id": 534207,
      "postDate": "2019-05-20T23:24:10.050Z",
      "content": "<p>Thank you for sharing too much. :D</p>",
      "rawMarkdown": "Thank you for sharing too much. :D",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 533623,
      "author_name": "Amjad",
      "author_url": "",
      "post_date": "2019-05-19T16:23:59.053000",
      "content": "<p>Nice feature indeed. It's great to predict those middle peaks right of course, but the segments with minor quakes are quite few. The problem is with the systematic error they cause especially before the mini-quake. Here is a figure of my absolute error on the oof predictions to show the case.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533623/13252/ooferror.png\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 534074,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T15:51:19.693000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534428,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-21T08:34:18.447000",
          "content": "<p>He has plot the error, not the predictions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 534542,
          "author_name": "Alex V B",
          "author_url": "",
          "post_date": "2019-05-21T12:35:36.170000",
          "content": "<p>Since doing manual feature selection to improve random forest is a good idea, it looks like a good idea to find the week learners by hand/visually.</p>\n\n<p>Nice plot. Sorry for not being able to contribute.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 534560,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T13:12:35.450000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535796,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-23T13:12:18.743000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535799,
          "author_name": "Amjad",
          "author_url": "",
          "post_date": "2019-05-23T13:18:03.140000",
          "content": "<p>Sorry I don't remember. Must be in the 1.90s.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535901,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-23T15:40:20.813000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533706,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2019-05-19T19:16:47.257000",
      "content": "<p>congratulations. Do you want to share this feature? i think my model can improve with it 😈</p>",
      "votes": 3,
      "replies": [
        {
          "id": 534032,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T13:38:45.337000",
          "content": "<p>How much are you ready to offer for it ? ;)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 533543,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-05-19T13:12:24.543000",
      "content": "<p>Also, this is also good according to my feeling...\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533543/13249/2.png\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 533567,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-19T14:16:47.950000",
          "content": "<p>I have similar ones, but they have peaks in between EQ.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 533537,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-05-19T13:05:03.927000",
      "content": "<p>This is my best feature.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/533537/13248/1.png\" alt=\"\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 533812,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2019-05-20T05:07:27.657000",
      "content": "<p><img src=\"https://i.ytimg.com/vi/uh4dLo7T2Ow/maxresdefault.jpg\" alt=\"a kind of magic\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 534368,
      "author_name": "安静",
      "author_url": "",
      "post_date": "2019-05-21T06:15:43.453000",
      "content": "<p>I've been looking at this picture all afternoon but I still can't think of it</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 534900,
      "author_name": "Karan Jakhar",
      "author_url": "",
      "post_date": "2019-05-22T02:05:39.313000",
      "content": "<p>Thanks for sharing <a href=\"/cpmpml\">@cpmpml</a> . This made me to work again on this competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 534713,
      "author_name": "Rashidul H",
      "author_url": "",
      "post_date": "2019-05-21T18:23:24.703000",
      "content": "<p>It's awesome!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 534451,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2019-05-21T09:09:57.043000",
      "content": "<p>Hi <a href=\"/cpmpml\">@cpmpml</a> . First of all, thanks for your insight. But what is correlation coefficient between this feature and target?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 534453,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-21T09:13:53.150000",
          "content": "<p>I don't compute correlation between features and target.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 534386,
      "author_name": "jeong jongmin",
      "author_url": "",
      "post_date": "2019-05-21T06:54:11.897000",
      "content": "<p>Wow its amazing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 534106,
      "author_name": "Himanshu Soni",
      "author_url": "",
      "post_date": "2019-05-20T17:35:15.763000",
      "content": "<p>Can you explain how much these features extent on Leaderboard score? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 534126,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T18:12:11.287000",
          "content": "<p>All I know is that removing this feature degrades my CV by about 0.02.  I think LB would degrade the same, but I will not burn a submission to check it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533593,
      "author_name": "Bellò Filippo",
      "author_url": "",
      "post_date": "2019-05-19T15:08:41.863000",
      "content": "<p>This is my first experience with ML, python and competitions. So i ask to more experienced programmers.\nHave I a feature similar to yours? In my opinion it give some informations around minor quakes.\nI have try to build a model around this feature, but for now i do not perform well in the leaderboard :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 539799,
      "author_name": "Felipe Loque",
      "author_url": "",
      "post_date": "2019-05-30T14:39:42.430000",
      "content": "<p>I've plotted 295 features here: <a href=\"https://www.kaggle.com/felipefonte99/plotting-295-features-vs-time-to-failure\">https://www.kaggle.com/felipefonte99/plotting-295-features-vs-time-to-failure</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 539346,
      "author_name": "Serhii Hrynko",
      "author_url": "",
      "post_date": "2019-05-30T00:12:21.307000",
      "content": "<p>Just for curiosity, how <code>plt.scatter(train.time_to_failure, train[feature], s=1)</code> would look like for this feature?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 539356,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-30T00:45:46.920000",
          "content": "<p>here it is:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/539356/13320/feat.png\" alt=\"plot\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 539848,
          "author_name": "Sia",
          "author_url": "",
          "post_date": "2019-05-30T15:48:17.530000",
          "content": "<p>Thanks for this useful plot. The leftmost  side of this plot like most of my useful features is causing a great part of model errors.  Is it safe to assume ttf &lt; 0.23 are outliers?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 540021,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-05-30T21:47:17.647000",
          "content": "<p>Most acoustic spikes are within ttf 0.310 - 0.315 and 150000 samples duration is around 0.035s. So I ignore ttf &lt; 0.275 to still capture those spikes. Otherwise model will always consider acoustic spike as the ones from false earthquakes, which are far from 0 ttf.</p>\n\n<p>I'm more conserned with right side of the plot, which still shows multiple parallel trends</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 540024,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-05-30T21:57:04.563000",
          "content": "<p>While that may seem like a good idea, we have no idea whether test samples include times &lt; 0.275 like you are suggesting. In fact we do not even know if test samples take the back end of an EQ with the front end of another EQ and use that as a random test sample.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 540052,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-05-31T00:13:24.707000",
          "content": "<p>Well we do know that test data has about 10 chunks with extremely hi acoustic spikes and even have spikes <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/82336#latest-515978\">higher than all training data</a> Which are all very likely to have nearly 0.315 ttf</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 534611,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-05-21T14:41:36.100000",
      "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> this is very inspiring. Makes me want to join the competition :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 534622,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-05-21T14:59:04.223000",
          "content": "<p>It's a little bit late, isn't it? Obviously, I don't doubt you'll end up high ;) What will you focus on for these  two weeks?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 534625,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-21T15:02:38.277000",
          "content": "<blockquote>\n  <p>It's a little bit late, isn't it? </p>\n</blockquote>\n\n<p>Don't underestimate Giba or other top kagglers, 13 days with a small dataset like this can be enough to get a good result.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 534686,
          "author_name": "Amjad",
          "author_url": "",
          "post_date": "2019-05-21T17:26:15.090000",
          "content": "<p><a href=\"/titericz\">@titericz</a> my model would benefit from a professional touch if you want to join and have a head start </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 534693,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2019-05-21T17:37:33.013000",
          "content": "<p>It seems easy to hire a grandmaster at 17th place 🥇😄. But what to do at the 100th place? 😞</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 534694,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-05-21T17:41:27.670000",
          "content": "<p>pray</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 534726,
          "author_name": "Amjad",
          "author_url": "",
          "post_date": "2019-05-21T18:33:14.680000",
          "content": "<p><a href=\"/sggpls\">@sggpls</a>  I don't think it's easy :D</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 535247,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-22T15:10:30.797000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 534043,
      "author_name": "Shinsei66",
      "author_url": "",
      "post_date": "2019-05-20T14:02:35.843000",
      "content": "<p>This is the plot of my best feature with TTF, maybe I'd better do more feature selection and change my CV strategy.....<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/534043/13257/best.PNG\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 534082,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T16:16:19.857000",
          "content": "<p>looks interesting.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 533455,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2019-05-19T09:31:53.313000",
      "content": "<p>I wouldn't call such a feature \"interesting\". It is rather a super feature. However, I have found that in this competition some good features work much better in pairs, so I wish you to find a feature for the pair.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 533528,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-19T12:46:22.447000",
          "content": "<p>It probably works well with other features I have indeed. I have not looked into feature interactions, that's not something I do but I know I probably should.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 533465,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-05-19T09:49:59.163000",
      "content": "<p>You have a typo here? ‘’worsens my CV score by 0.023, which would most probably translate to same CV score”</p>",
      "votes": 1,
      "replies": [
        {
          "id": 533527,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-19T12:45:20.350000",
          "content": "<p>Thanks, I fixed the typo.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543075,
      "author_name": "Bhavika",
      "author_url": "",
      "post_date": "2019-06-04T10:12:14.127000",
      "content": "<p>Hi, <a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>First of all, I want to thank you. You have been an inspiration for me to learn machine learning.</p>\n\n<p>Here, is the plot of the mean feature of a signal.</p>\n\n<p><img src=\"https://drive.google.com/file/d/1PBWruz3bnpK8yXIQQrgP4jYY_tUmV4iZ/view\" alt=\"Graph\"></p>\n\n<p>Don't know, here the image is not displayed.</p>\n\n<p>Image Link :  \"&gt;https://drive.google.com/open?id=1uzwv9eTtjGNOyWPZWeGNgCIz2qaP3O_</p>\n\n<p>The intimation from the plot is that it's not a good feature for prediction.\nBut, This feature took 3rd place as in graph of feature importance of trained model.\nHow could we decide which feature is important for the model?\nAgain the correlation between this feature and target is quite low.</p>\n\n<p>Please take a look.\nCan you guide me to know how you recognize the important feature from plotting?</p>\n\n<p>Thanks a lot.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 543109,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T10:32:57.070000",
          "content": "<p>I use plot to select features, but I further select by training models.  If models are better with it, then i keep the feature.  Mean is a special case here, it is highly discriminant between train and test and we decided not to use it for that reason.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543125,
          "author_name": "Bhavika",
          "author_url": "",
          "post_date": "2019-06-04T10:43:46.790000",
          "content": "<p>Thanks for the reply..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 540981,
      "author_name": "Vinayak Sharma",
      "author_url": "",
      "post_date": "2019-06-01T13:06:09.207000",
      "content": "<p>I tried to plot them, this may help\n<a href=\"https://www.kaggle.com/vinayaks/feature-selection-simplified\">https://www.kaggle.com/vinayaks/feature-selection-simplified</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 538475,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2019-05-28T15:59:21.383000",
      "content": "<p>Great job.   What's your training and cross validation mae ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 538021,
      "author_name": "Peiyi",
      "author_url": "",
      "post_date": "2019-05-28T02:31:55.243000",
      "content": "<p>I got something alike.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 537801,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "2019-05-27T16:03:30.250000",
      "content": "<p>Though my opinion may not matter as your team is already on top of LB which means you are definitely heading in right direction. But since you are one of my favorite GM I would like to ask something. Is it not obvious that earthquakes are not periodic? If you see the slope of decay of <code>time = t</code> the next EQ occurs, shouldn't it be decaying like a negative exponential?</p>\n\n<p><img src=\"https://www.physics.uoguelph.ca/tutorials/exp/graph5.gif\" alt=\"\"></p>\n\n<p>using image to show the decay it should be, its not based on dataset</p>",
      "votes": 0,
      "replies": [
        {
          "id": 538221,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-28T10:10:41.830000",
          "content": "<blockquote>\n  <p>Is it not obvious that earthquakes are not periodic?</p>\n</blockquote>\n\n<p>I think organizers wrote in one of their paper or in the material available here that EQ are not periodic indeed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 538389,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-28T14:08:39.367000",
          "content": "<p>From the introduction post ( <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525</a> ):</p>\n\n<blockquote>\n  <p>For this challenge we selected an experiment that exhibits a very aperiodic and more realistic behavior compared to the data we studied in our early work, with earthquakes occurring very irregularly. </p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 535983,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-05-23T18:28:17.607000",
      "content": "<p>I got a similar structure using Mean Squared Percentage Error: (A-X)/A\nBut I haven't the skill to use it!!\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/535983/13282/MSPA.png\" alt=\"MSPE\"></p>\n\n<p>So I am guessing it is some percentile feature - cannot wait to find out!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 535329,
      "author_name": "SUPERLUMINAL",
      "author_url": "",
      "post_date": "2019-05-22T17:51:26.067000",
      "content": "<p>Did you augment or balance your data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 535339,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-22T18:09:01.690000",
          "content": "<p>I'll answer that after competition end ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 535256,
      "author_name": "aaaa",
      "author_url": "",
      "post_date": "2019-05-22T15:20:52.120000",
      "content": "<p>Thanks a lot, this is really insightful !\nI have a question about the difference distribution between train and test data. Did you adjust train data distribution to become similar to test's and then selected featres? Or your features are already general to detect same meaning in train and test? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 535260,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-22T15:22:40.600000",
          "content": "<p>I'll answer that after competition end ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 533949,
      "author_name": "Ankur Mittal",
      "author_url": "",
      "post_date": "2019-05-20T10:09:00.643000",
      "content": "<p>Sadly I already have this feature, but my LB score is stuck at 1.484. I am a total newbie though,(6th day into machine learning)  learning new things everyday. :)\nP.S: I am looking for a mentor :p</p>",
      "votes": 0,
      "replies": [
        {
          "id": 534031,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T13:38:24.690000",
          "content": "<blockquote>\n  <p>Sadly I already have this feature</p>\n</blockquote>\n\n<p>How do you know it is the same feature?  I frankly doubt it is, given how I constructed it.  You may have a similar one, but the odds of using the exact same one are very little IMHO.</p>\n\n<p>Can you plot you feature the same way I did and share here?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 534674,
          "author_name": "Ankur Mittal",
          "author_url": "",
          "post_date": "2019-05-21T16:51:04.790000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> you are right my feature is indeed quite different. Looked similar at first glance. I ll try to replicate your pattern you posted here. Thanks for the hint. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533944,
      "author_name": "Mohsen Yazdinejad",
      "author_url": "",
      "post_date": "2019-05-20T09:41:00.887000",
      "content": "<p>How about this feature? Is this good?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 533948,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T10:03:21.747000",
          "content": "<p>How would I know?  We can't see the values except for few peaks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 533952,
          "author_name": "Mohsen Yazdinejad",
          "author_url": "",
          "post_date": "2019-05-20T10:14:42.777000",
          "content": "<p>I think the other values are 0.\nI mean feature that shows just the EQ could be useful?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 533957,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T10:25:14.037000",
          "content": "<p>just the peak is of not much value, it may help predict 16 points at most.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 533777,
      "author_name": "Alexis Moraga",
      "author_url": "",
      "post_date": "2019-05-20T02:22:26.687000",
      "content": "<p>Can i ask you how do you generate those plots for the features? No for this one in particular, i mean in general. Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 533794,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-20T04:08:43.407000",
          "content": "<p>In a notebook:</p>\n\n<pre><code>from matplotlib import pyplot as plt\n%matploltlib inline\nfig, ax = plt.subplots(1, 1, figsize=(12, 6))\nax.plot(train.time_to_failure)\nax.plot(train[feature])\n</code></pre>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 534034,
          "author_name": "Alexis Moraga",
          "author_url": "",
          "post_date": "2019-05-20T13:51:58.077000",
          "content": "<p>Thank you, this will be helpful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533621,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-19T16:12:40.353000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 534430,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T08:35:30.680000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534543,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T12:36:33.707000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533512,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-19T12:03:49.310000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 533530,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-19T12:49:31.463000",
          "content": "",
          "votes": 6,
          "replies": []
        },
        {
          "id": 533909,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T08:29:28.950000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534442,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T08:50:17.733000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535950,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-23T17:16:55.097000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 534267,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-21T02:14:29.290000",
      "content": "",
      "votes": 10,
      "replies": [
        {
          "id": 534382,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T06:46:19.600000",
          "content": "",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 534016,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-20T12:58:12.033000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 534017,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T13:03:02.783000",
          "content": "",
          "votes": 8,
          "replies": []
        },
        {
          "id": 534024,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T13:22:28.083000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 534030,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T13:37:02.727000",
          "content": "",
          "votes": 8,
          "replies": []
        },
        {
          "id": 534049,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T14:25:15.580000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 534055,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T14:39:42.137000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 534130,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T18:15:42.870000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 534165,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T20:21:09.270000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 534205,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T23:17:04.537000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 535129,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-22T10:55:37.063000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535145,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-22T11:24:27.417000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535196,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-22T13:19:25.517000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535237,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-22T14:52:53.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535250,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-22T15:16:57.197000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 533626,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-19T16:32:48.863000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 534534,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T12:18:26.580000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534539,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-21T12:29:13.890000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 534207,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-20T23:24:10.050000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "533443": "Some people pinged me about what could be my magic features, given I get good results with less than 15 features (single LGB at 1.288 as of now).  Well, removing my best feature according to lgb worsens my CV score by 0.023, which would most probably translate to same LB score evolution.  Therefore it is not a magic feature.  But it still is an interesting feature.  When we plot it on train data along target, we see that it has 16 peaks corresponding to the high acoustic peaks before EQ, but not much peaks for the mini EQ that are hard to ignore.  I re-scaled the feature value to make the picture nicer.\n\n![plot](https://storage.googleapis.com/kaggle-forum-message-attachments/533443/13245/best.png)\nI use this kind of plot to select my features.\n\n",
    "533623": "Nice feature indeed. It's great to predict those middle peaks right of course, but the segments with minor quakes are quite few. The problem is with the systematic error they cause especially before the mini-quake. Here is a figure of my absolute error on the oof predictions to show the case.\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/533623/13252/ooferror.png)",
    "533706": "congratulations. Do you want to share this feature? i think my model can improve with it 😈",
    "533543": "Also, this is also good according to my feeling...\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/533543/13249/2.png)",
    "533537": "This is my best feature.\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/533537/13248/1.png)",
    "533812": "![a kind of magic](https://i.ytimg.com/vi/uh4dLo7T2Ow/maxresdefault.jpg)",
    "534368": "I've been looking at this picture all afternoon but I still can't think of it",
    "534900": "Thanks for sharing @cpmpml . This made me to work again on this competition.",
    "534713": "It's awesome!",
    "534451": "Hi @cpmpml . First of all, thanks for your insight. But what is correlation coefficient between this feature and target?",
    "534386": "Wow its amazing",
    "534106": "Can you explain how much these features extent on Leaderboard score? ",
    "533593": "This is my first experience with ML, python and competitions. So i ask to more experienced programmers.\nHave I a feature similar to yours? In my opinion it give some informations around minor quakes.\nI have try to build a model around this feature, but for now i do not perform well in the leaderboard :-)",
    "539799": "I've plotted 295 features here: https://www.kaggle.com/felipefonte99/plotting-295-features-vs-time-to-failure",
    "539346": "Just for curiosity, how `plt.scatter(train.time_to_failure, train[feature], s=1)` would look like for this feature?",
    "534611": "Thanks @cpmpml this is very inspiring. Makes me want to join the competition :)",
    "534043": "This is the plot of my best feature with TTF, maybe I'd better do more feature selection and change my CV strategy.....![](https://storage.googleapis.com/kaggle-forum-message-attachments/534043/13257/best.PNG)",
    "533455": "I wouldn't call such a feature \"interesting\". It is rather a super feature. However, I have found that in this competition some good features work much better in pairs, so I wish you to find a feature for the pair.",
    "533465": "You have a typo here? ‘’worsens my CV score by 0.023, which would most probably translate to same CV score”",
    "543075": "Hi, @cpmpml \n\nFirst of all, I want to thank you. You have been an inspiration for me to learn machine learning.\n\nHere, is the plot of the mean feature of a signal.\n\n![Graph](https://drive.google.com/file/d/1PBWruz3bnpK8yXIQQrgP4jYY_tUmV4iZ/view)\n\nDon't know, here the image is not displayed.\n\nImage Link :  https://drive.google.com/open?id=1uzwv9eTtjGNOyWPZWeGNgCIz2qaP3O__\n\nThe intimation from the plot is that it's not a good feature for prediction.\nBut, This feature took 3rd place as in graph of feature importance of trained model.\nHow could we decide which feature is important for the model?\nAgain the correlation between this feature and target is quite low.\n\nPlease take a look.\nCan you guide me to know how you recognize the important feature from plotting?\n\nThanks a lot.",
    "540981": "I tried to plot them, this may help\nhttps://www.kaggle.com/vinayaks/feature-selection-simplified",
    "538475": "Great job.   What's your training and cross validation mae ?",
    "538021": "I got something alike.",
    "537801": "Though my opinion may not matter as your team is already on top of LB which means you are definitely heading in right direction. But since you are one of my favorite GM I would like to ask something. Is it not obvious that earthquakes are not periodic? If you see the slope of decay of `time = t` the next EQ occurs, shouldn't it be decaying like a negative exponential?\n\n\n![](https://www.physics.uoguelph.ca/tutorials/exp/graph5.gif)\n\nusing image to show the decay it should be, its not based on dataset",
    "535983": "I got a similar structure using Mean Squared Percentage Error: (A-X)/A\nBut I haven't the skill to use it!!\n![MSPE](https://storage.googleapis.com/kaggle-forum-message-attachments/535983/13282/MSPA.png)\n\nSo I am guessing it is some percentile feature - cannot wait to find out!",
    "535329": "Did you augment or balance your data?",
    "535256": "Thanks a lot, this is really insightful !\nI have a question about the difference distribution between train and test data. Did you adjust train data distribution to become similar to test's and then selected featres? Or your features are already general to detect same meaning in train and test? ",
    "533949": "Sadly I already have this feature, but my LB score is stuck at 1.484. I am a total newbie though,(6th day into machine learning)  learning new things everyday. :)\nP.S: I am looking for a mentor :p",
    "533944": "How about this feature? Is this good?",
    "533777": "Can i ask you how do you generate those plots for the features? No for this one in particular, i mean in general. Thanks",
    "533621": "That is corref?",
    "533512": "Is your \"single LGBM\" the average of many folds?",
    "534267": "",
    "534016": "",
    "533626": "",
    "534207": "Thank you for sharing too much. :D"
  }
}