{
  "id": 94903,
  "title": "22nd place solution described (molaee)",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94903",
  "author_name": "Behnam Molaee",
  "post_date": "2019-06-07T22:53:50.551000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Its my pleasure to share the ad-hoc method I used in this competition with you. I am not that expert in ML, however I hope that my method opens some fruitful discussions for all of us.</p>\n\n<p>My kernel was forked initially from here <a href=\"https://www.kaggle.com/artgor/even-more-features\">here</a> (Thank you <a href=\"/artgor\">@artgor</a> )</p>\n\n<p><strong>Feature extraction:</strong>\nAfter running the kernel, I realized that the TTF is not estimated very well in the vicinity of failure times (it becomes saturated). I worked a lot on feature extraction to overcome this issue. My focus was mostly on time-frequency and MFCC related features and could increase the Pearson correlation from 0.65 (in the original kernel) to 0.68 (when measured on the entire data). The pearson correlation increased from ~0.25 to 0.38 when measured on the data in the vicinity of the failure times (&lt;3.2 s).\nMy best feature: The area under 0.2 quantile of the cumulative STFT  (presented in log scale) of the audio signal (NFFT=128, signal de-trended).\nAfter a couple of days, I had about 1300 similar features!! -&gt; LB public score improved 0.1 from ~1.5 to 1.4</p>\n\n<p><strong>Benefit from the leak:</strong>\nI became aware of the leak only 2 days before the submission deadline. I did not carefully analyzed the leak data, nor did match the mean of predicted TTF with the one presented in the paper. I just, removed the 5th and 6th EQ data from the training set. Visually inspection of the test and train data, I had the feeling that after removing these two EQs, test and train data become very similar to each other.</p>\n\n<p><strong>Model:</strong>\nI got the lgb model presented in the forked Kernel. Then created 3 categories of data set: (A) entire data excluding the 2 above mentioned EQs, (B) data very close to failures (3.2 s and less) , (C) data far from failures (6 s and more)\nFor each data set, a separate lgb model was trained. The model corresponding to (A) had a mastering role. If its prediction was below a threshold (3 s) or above a threshold (7 s) then its output was being merged with that of (B) or (C) in a linear manner.</p>\n\n<p><code>\nfor ind, row in X_test.iterrows():\n    if prediction_lgb_m[ind] &amp;lt; thr_l:\n        alpha = (thr_l - prediction_lgb_m[ind])/thr_l\n        prediction_lgb[ind] = (alpha)*prediction_lgb_short[ind] + (1-alpha)*prediction_lgb_m[ind]\n    elif prediction_lgb_m[ind] &amp;gt; thr_h:\n        prediction_lgb[ind] = prediction_lgb_long[ind]\n    else:\n        prediction_lgb[ind] = prediction_lgb_m[ind]\nplt.hist([prediction_lgb,prediction_lgb_m],50)\n</code>\nTTF histogram over test data using model (A) [orange] and Combined modes (A,B,C) [Blue]:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/547591/13437/Capture.PNG\" alt=\"\"></p>\n\n<p><strong>Topics for discussions:</strong></p>\n\n<p>1) Does it make sense to have a cascade of lgb models as I described? In my opinion, this method should add another high-level branch to the tree of the original model (A). But I could not get the same results when I increased the number of layers or leafs in (A) !!!\nIndeed, when I used models (B) and (C) I could get pretty lower/higher TTF values for the training data close/far to the failure time. The model (A) could not provide such results.</p>\n\n<p>2) After one month working on feature extraction, I had about 2300 similar and correlated features. Compared to my original 1300 features, I could obtain a bit better CV values (2.01 --&gt; 1.99) and I could see less fluctuations over (y - ypred) training data set. However, LB public score became a bit worse (1.399 --&gt; 1.43) !! As if I had over fitted. Right? What is the general procedure to reduce the number of features from 2300 to a lower value?</p>\n\n<p>3) I tried to reduce the number of features by removing highly correlated features (0.95 and more) from the feature list. This reduced the number of features to about 800. However, CV increased (+0.01), as well as LB public (+0.03).</p>\n\n<p>Finally, I decided to use my entire 2300 features, 3 LGBs, and the removed EQs 5 and 6 from the training set.</p>\n\n<p><strong>What did not work:</strong>\n1) PCA over the features increased the LB public score. So I did not follow it.\n2) Combination of LGB with other methods was not really a big success. I preferred not to step in that way.</p>",
  "messages": [
    {
      "id": 547591,
      "postDate": "2019-06-07T22:53:50.550Z",
      "content": "<p>Its my pleasure to share the ad-hoc method I used in this competition with you. I am not that expert in ML, however I hope that my method opens some fruitful discussions for all of us.</p>\n\n<p>My kernel was forked initially from here <a href=\"https://www.kaggle.com/artgor/even-more-features\">here</a> (Thank you <a href=\"/artgor\">@artgor</a> )</p>\n\n<p><strong>Feature extraction:</strong>\nAfter running the kernel, I realized that the TTF is not estimated very well in the vicinity of failure times (it becomes saturated). I worked a lot on feature extraction to overcome this issue. My focus was mostly on time-frequency and MFCC related features and could increase the Pearson correlation from 0.65 (in the original kernel) to 0.68 (when measured on the entire data). The pearson correlation increased from ~0.25 to 0.38 when measured on the data in the vicinity of the failure times (&lt;3.2 s).\nMy best feature: The area under 0.2 quantile of the cumulative STFT  (presented in log scale) of the audio signal (NFFT=128, signal de-trended).\nAfter a couple of days, I had about 1300 similar features!! -&gt; LB public score improved 0.1 from ~1.5 to 1.4</p>\n\n<p><strong>Benefit from the leak:</strong>\nI became aware of the leak only 2 days before the submission deadline. I did not carefully analyzed the leak data, nor did match the mean of predicted TTF with the one presented in the paper. I just, removed the 5th and 6th EQ data from the training set. Visually inspection of the test and train data, I had the feeling that after removing these two EQs, test and train data become very similar to each other.</p>\n\n<p><strong>Model:</strong>\nI got the lgb model presented in the forked Kernel. Then created 3 categories of data set: (A) entire data excluding the 2 above mentioned EQs, (B) data very close to failures (3.2 s and less) , (C) data far from failures (6 s and more)\nFor each data set, a separate lgb model was trained. The model corresponding to (A) had a mastering role. If its prediction was below a threshold (3 s) or above a threshold (7 s) then its output was being merged with that of (B) or (C) in a linear manner.</p>\n\n<p><code>\nfor ind, row in X_test.iterrows():\n    if prediction_lgb_m[ind] &amp;lt; thr_l:\n        alpha = (thr_l - prediction_lgb_m[ind])/thr_l\n        prediction_lgb[ind] = (alpha)*prediction_lgb_short[ind] + (1-alpha)*prediction_lgb_m[ind]\n    elif prediction_lgb_m[ind] &amp;gt; thr_h:\n        prediction_lgb[ind] = prediction_lgb_long[ind]\n    else:\n        prediction_lgb[ind] = prediction_lgb_m[ind]\nplt.hist([prediction_lgb,prediction_lgb_m],50)\n</code>\nTTF histogram over test data using model (A) [orange] and Combined modes (A,B,C) [Blue]:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/547591/13437/Capture.PNG\" alt=\"\"></p>\n\n<p><strong>Topics for discussions:</strong></p>\n\n<p>1) Does it make sense to have a cascade of lgb models as I described? In my opinion, this method should add another high-level branch to the tree of the original model (A). But I could not get the same results when I increased the number of layers or leafs in (A) !!!\nIndeed, when I used models (B) and (C) I could get pretty lower/higher TTF values for the training data close/far to the failure time. The model (A) could not provide such results.</p>\n\n<p>2) After one month working on feature extraction, I had about 2300 similar and correlated features. Compared to my original 1300 features, I could obtain a bit better CV values (2.01 --&gt; 1.99) and I could see less fluctuations over (y - ypred) training data set. However, LB public score became a bit worse (1.399 --&gt; 1.43) !! As if I had over fitted. Right? What is the general procedure to reduce the number of features from 2300 to a lower value?</p>\n\n<p>3) I tried to reduce the number of features by removing highly correlated features (0.95 and more) from the feature list. This reduced the number of features to about 800. However, CV increased (+0.01), as well as LB public (+0.03).</p>\n\n<p>Finally, I decided to use my entire 2300 features, 3 LGBs, and the removed EQs 5 and 6 from the training set.</p>\n\n<p><strong>What did not work:</strong>\n1) PCA over the features increased the LB public score. So I did not follow it.\n2) Combination of LGB with other methods was not really a big success. I preferred not to step in that way.</p>",
      "rawMarkdown": "Its my pleasure to share the ad-hoc method I used in this competition with you. I am not that expert in ML, however I hope that my method opens some fruitful discussions for all of us.\n\nMy kernel was forked initially from here [here](https://www.kaggle.com/artgor/even-more-features) (Thank you @artgor )\n\n**Feature extraction:**\nAfter running the kernel, I realized that the TTF is not estimated very well in the vicinity of failure times (it becomes saturated). I worked a lot on feature extraction to overcome this issue. My focus was mostly on time-frequency and MFCC related features and could increase the Pearson correlation from 0.65 (in the original kernel) to 0.68 (when measured on the entire data). The pearson correlation increased from ~0.25 to 0.38 when measured on the data in the vicinity of the failure times (&lt;3.2 s).\nMy best feature: The area under 0.2 quantile of the cumulative STFT  (presented in log scale) of the audio signal (NFFT=128, signal de-trended).\nAfter a couple of days, I had about 1300 similar features!! -&gt; LB public score improved 0.1 from ~1.5 to 1.4\n\n**Benefit from the leak:**\nI became aware of the leak only 2 days before the submission deadline. I did not carefully analyzed the leak data, nor did match the mean of predicted TTF with the one presented in the paper. I just, removed the 5th and 6th EQ data from the training set. Visually inspection of the test and train data, I had the feeling that after removing these two EQs, test and train data become very similar to each other.\n\n**Model:**\nI got the lgb model presented in the forked Kernel. Then created 3 categories of data set: (A) entire data excluding the 2 above mentioned EQs, (B) data very close to failures (3.2 s and less) , (C) data far from failures (6 s and more)\nFor each data set, a separate lgb model was trained. The model corresponding to (A) had a mastering role. If its prediction was below a threshold (3 s) or above a threshold (7 s) then its output was being merged with that of (B) or (C) in a linear manner.\n\n`\nfor ind, row in X_test.iterrows():\n    if prediction_lgb_m[ind] &lt; thr_l:\n        alpha = (thr_l - prediction_lgb_m[ind])/thr_l\n        prediction_lgb[ind] = (alpha)*prediction_lgb_short[ind] + (1-alpha)*prediction_lgb_m[ind]\n    elif prediction_lgb_m[ind] &gt; thr_h:\n        prediction_lgb[ind] = prediction_lgb_long[ind]\n    else:\n        prediction_lgb[ind] = prediction_lgb_m[ind]\nplt.hist([prediction_lgb,prediction_lgb_m],50)\n`\nTTF histogram over test data using model (A) [orange] and Combined modes (A,B,C) [Blue]:\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/547591/13437/Capture.PNG)\n\n**Topics for discussions:**\n\n1) Does it make sense to have a cascade of lgb models as I described? In my opinion, this method should add another high-level branch to the tree of the original model (A). But I could not get the same results when I increased the number of layers or leafs in (A) !!!\nIndeed, when I used models (B) and (C) I could get pretty lower/higher TTF values for the training data close/far to the failure time. The model (A) could not provide such results.\n\n2) After one month working on feature extraction, I had about 2300 similar and correlated features. Compared to my original 1300 features, I could obtain a bit better CV values (2.01 --&gt; 1.99) and I could see less fluctuations over (y - ypred) training data set. However, LB public score became a bit worse (1.399 --&gt; 1.43) !! As if I had over fitted. Right? What is the general procedure to reduce the number of features from 2300 to a lower value?\n\n3) I tried to reduce the number of features by removing highly correlated features (0.95 and more) from the feature list. This reduced the number of features to about 800. However, CV increased (+0.01), as well as LB public (+0.03).\n\nFinally, I decided to use my entire 2300 features, 3 LGBs, and the removed EQs 5 and 6 from the training set.\n\n**What did not work:**\n1) PCA over the features increased the LB public score. So I did not follow it.\n2) Combination of LGB with other methods was not really a big success. I preferred not to step in that way.\n \n\n",
      "votes": 4
    },
    {
      "id": 547809,
      "postDate": "2019-06-08T10:06:44.363Z",
      "content": "<p>FYI: matching the mean of TTF of the test data to the one presented in the paper, could result 2.32 (LB private) </p>",
      "rawMarkdown": "FYI: matching the mean of TTF of the test data to the one presented in the paper, could result 2.32 (LB private) "
    },
    {
      "id": 547593,
      "postDate": "2019-06-07T22:56:59.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 547809,
      "author_name": "Behnam Molaee",
      "author_url": "",
      "post_date": "2019-06-08T10:06:44.363000",
      "content": "<p>FYI: matching the mean of TTF of the test data to the one presented in the paper, could result 2.32 (LB private) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 547593,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-07T22:56:59.150000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "547591": "Its my pleasure to share the ad-hoc method I used in this competition with you. I am not that expert in ML, however I hope that my method opens some fruitful discussions for all of us.\n\nMy kernel was forked initially from here [here](https://www.kaggle.com/artgor/even-more-features) (Thank you @artgor )\n\n**Feature extraction:**\nAfter running the kernel, I realized that the TTF is not estimated very well in the vicinity of failure times (it becomes saturated). I worked a lot on feature extraction to overcome this issue. My focus was mostly on time-frequency and MFCC related features and could increase the Pearson correlation from 0.65 (in the original kernel) to 0.68 (when measured on the entire data). The pearson correlation increased from ~0.25 to 0.38 when measured on the data in the vicinity of the failure times (&lt;3.2 s).\nMy best feature: The area under 0.2 quantile of the cumulative STFT  (presented in log scale) of the audio signal (NFFT=128, signal de-trended).\nAfter a couple of days, I had about 1300 similar features!! -&gt; LB public score improved 0.1 from ~1.5 to 1.4\n\n**Benefit from the leak:**\nI became aware of the leak only 2 days before the submission deadline. I did not carefully analyzed the leak data, nor did match the mean of predicted TTF with the one presented in the paper. I just, removed the 5th and 6th EQ data from the training set. Visually inspection of the test and train data, I had the feeling that after removing these two EQs, test and train data become very similar to each other.\n\n**Model:**\nI got the lgb model presented in the forked Kernel. Then created 3 categories of data set: (A) entire data excluding the 2 above mentioned EQs, (B) data very close to failures (3.2 s and less) , (C) data far from failures (6 s and more)\nFor each data set, a separate lgb model was trained. The model corresponding to (A) had a mastering role. If its prediction was below a threshold (3 s) or above a threshold (7 s) then its output was being merged with that of (B) or (C) in a linear manner.\n\n`\nfor ind, row in X_test.iterrows():\n    if prediction_lgb_m[ind] &lt; thr_l:\n        alpha = (thr_l - prediction_lgb_m[ind])/thr_l\n        prediction_lgb[ind] = (alpha)*prediction_lgb_short[ind] + (1-alpha)*prediction_lgb_m[ind]\n    elif prediction_lgb_m[ind] &gt; thr_h:\n        prediction_lgb[ind] = prediction_lgb_long[ind]\n    else:\n        prediction_lgb[ind] = prediction_lgb_m[ind]\nplt.hist([prediction_lgb,prediction_lgb_m],50)\n`\nTTF histogram over test data using model (A) [orange] and Combined modes (A,B,C) [Blue]:\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/547591/13437/Capture.PNG)\n\n**Topics for discussions:**\n\n1) Does it make sense to have a cascade of lgb models as I described? In my opinion, this method should add another high-level branch to the tree of the original model (A). But I could not get the same results when I increased the number of layers or leafs in (A) !!!\nIndeed, when I used models (B) and (C) I could get pretty lower/higher TTF values for the training data close/far to the failure time. The model (A) could not provide such results.\n\n2) After one month working on feature extraction, I had about 2300 similar and correlated features. Compared to my original 1300 features, I could obtain a bit better CV values (2.01 --&gt; 1.99) and I could see less fluctuations over (y - ypred) training data set. However, LB public score became a bit worse (1.399 --&gt; 1.43) !! As if I had over fitted. Right? What is the general procedure to reduce the number of features from 2300 to a lower value?\n\n3) I tried to reduce the number of features by removing highly correlated features (0.95 and more) from the feature list. This reduced the number of features to about 800. However, CV increased (+0.01), as well as LB public (+0.03).\n\nFinally, I decided to use my entire 2300 features, 3 LGBs, and the removed EQs 5 and 6 from the training set.\n\n**What did not work:**\n1) PCA over the features increased the LB public score. So I did not follow it.\n2) Combination of LGB with other methods was not really a big success. I preferred not to step in that way.\n \n\n",
    "547809": "FYI: matching the mean of TTF of the test data to the one presented in the paper, could result 2.32 (LB private) ",
    "547593": ""
  }
}