{
  "id": 94408,
  "title": "My bit of the 7th place solution (team ABC)",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94408",
  "author_name": "",
  "post_date": "2019-06-04T10:14:20.223241Z",
  "votes": 25,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, let me congratulate with my teammates: \n<a href=\"/cpmpml\">@cpmpml</a>  : thanks for accepting my invitation and for teaching us so much in the last few weeks!\n<a href=\"/areveillon\">@areveillon</a> : thanks for your original ideas and your perseverance. And most important: congratulations on becoming competition master!</p>\n\n<p>I am going to describe mostly what has been for me the pre-CPMP era of the competition. After merging, we adopted most of the brilliant ideas that CPMP will for sure describe in his post. I let him also describe ensembling and the like.</p>\n\n<p><em>CV</em>\n16 EQ-wise CV scheme. </p>\n\n<p><em>Data augmentation</em>\n150k segments with 50k overlap, i.e. 3x augmentation.</p>\n\n<p><em>Pre-processing</em>\nSubtracting mean of each segment.</p>\n\n<p><em>Feature engineering</em>\nWhat works: <br>\n- simple features like percentage close to median, smart combination of median, mean, std, min, max, number of peaks with support N, autocorrelation, rolling quantiles, etc. <br>\n- same as above, but applying various band/high/low pass filters . \n- MFCC coefficients inspired by public <a href=\"https://www.kaggle.com/taqanori/trying-mfcc-mel-frequency-cepstral-coefficients\">public kernel</a> . \n- STFT features inspired by <a href=\"https://www.kaggle.com/tsilveira/time-frequency-analysis-with-stft\">public kernel</a></p>\n\n<p><em>Feature selection</em>\nIterative Adversarial Validation to remove top train/test discriminating features until AUC &lt;= 0.7. <br>\nGenetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.</p>\n\n<p><em>Target transformation</em>\nSmall improvement from applying sqrt(target).</p>\n\n<p><em>Modelling</em>\n1 Lightgbm + 1 MLP both using mean_squared_error loss</p>\n\n<p><em>Post-processing</em>\nApplying heuristics to adjust predictions for segments with large acoustic events at around TTF ~ 0.3. Small effect as the quality of predictions increased towards the end of the competition.</p>",
  "messages": [
    {
      "id": "543079",
      "postDate": "06/04/2019 10:14:20",
      "content": "<p>First of all, let me congratulate with my teammates: \n<a href=\"/cpmpml\">@cpmpml</a>  : thanks for accepting my invitation and for teaching us so much in the last few weeks!\n<a href=\"/areveillon\">@areveillon</a> : thanks for your original ideas and your perseverance. And most important: congratulations on becoming competition master!</p>\n\n<p>I am going to describe mostly what has been for me the pre-CPMP era of the competition. After merging, we adopted most of the brilliant ideas that CPMP will for sure describe in his post. I let him also describe ensembling and the like.</p>\n\n<p><em>CV</em>\n16 EQ-wise CV scheme. </p>\n\n<p><em>Data augmentation</em>\n150k segments with 50k overlap, i.e. 3x augmentation.</p>\n\n<p><em>Pre-processing</em>\nSubtracting mean of each segment.</p>\n\n<p><em>Feature engineering</em>\nWhat works: <br>\n- simple features like percentage close to median, smart combination of median, mean, std, min, max, number of peaks with support N, autocorrelation, rolling quantiles, etc. <br>\n- same as above, but applying various band/high/low pass filters . \n- MFCC coefficients inspired by public <a href=\"https://www.kaggle.com/taqanori/trying-mfcc-mel-frequency-cepstral-coefficients\">public kernel</a> . \n- STFT features inspired by <a href=\"https://www.kaggle.com/tsilveira/time-frequency-analysis-with-stft\">public kernel</a></p>\n\n<p><em>Feature selection</em>\nIterative Adversarial Validation to remove top train/test discriminating features until AUC &lt;= 0.7. <br>\nGenetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.</p>\n\n<p><em>Target transformation</em>\nSmall improvement from applying sqrt(target).</p>\n\n<p><em>Modelling</em>\n1 Lightgbm + 1 MLP both using mean_squared_error loss</p>\n\n<p><em>Post-processing</em>\nApplying heuristics to adjust predictions for segments with large acoustic events at around TTF ~ 0.3. Small effect as the quality of predictions increased towards the end of the competition.</p>",
      "rawMarkdown": "First of all, let me congratulate with my teammates: \n@cpmpml  : thanks for accepting my invitation and for teaching us so much in the last few weeks!\n@areveillon : thanks for your original ideas and your perseverance. And most important: congratulations on becoming competition master!\n\nI am going to describe mostly what has been for me the pre-CPMP era of the competition. After merging, we adopted most of the brilliant ideas that CPMP will for sure describe in his post. I let him also describe ensembling and the like.\n\n*CV*\n16 EQ-wise CV scheme. \n\n*Data augmentation*\n150k segments with 50k overlap, i.e. 3x augmentation.\n\n*Pre-processing*\nSubtracting mean of each segment.\n\n*Feature engineering*\nWhat works:   \n- simple features like percentage close to median, smart combination of median, mean, std, min, max, number of peaks with support N, autocorrelation, rolling quantiles, etc.  \n- same as above, but applying various band/high/low pass filters . \n- MFCC coefficients inspired by public [public kernel](https://www.kaggle.com/taqanori/trying-mfcc-mel-frequency-cepstral-coefficients) . \n- STFT features inspired by [public kernel](https://www.kaggle.com/tsilveira/time-frequency-analysis-with-stft)\n\n*Feature selection*\nIterative Adversarial Validation to remove top train/test discriminating features until AUC &lt;= 0.7.  \nGenetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.\n\n*Target transformation*\nSmall improvement from applying sqrt(target).\n\n*Modelling*\n1 Lightgbm + 1 MLP both using mean_squared_error loss\n\n*Post-processing*\nApplying heuristics to adjust predictions for segments with large acoustic events at around TTF ~ 0.3. Small effect as the quality of predictions increased towards the end of the competition.",
      "votes": null
    },
    {
      "id": "543116",
      "postDate": "06/04/2019 10:38:47",
      "content": "<p>It was a pleasure to team with you!</p>",
      "rawMarkdown": "It was a pleasure to team with you!",
      "votes": null
    },
    {
      "id": "543174",
      "postDate": "06/04/2019 11:13:01",
      "content": "<p>Thanks for sharing. I was trying all this stuff but I was not confident enough. </p>",
      "rawMarkdown": "Thanks for sharing. I was trying all this stuff but I was not confident enough.",
      "votes": null
    },
    {
      "id": "543184",
      "postDate": "06/04/2019 11:17:32",
      "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> I have learned a lot from you. I am still in beginning stage of learning. In this competition I didn't go systematic. And I have marked your words <a href=\"/cpmpml\">@cpmpml</a> \"first set a strong validation technique \" and after that we should go systematic. I am happy 😊 that I got to learn this much.</p>",
      "rawMarkdown": "Thanks @cpmpml I have learned a lot from you. I am still in beginning stage of learning. In this competition I didn't go systematic. And I have marked your words @cpmpml \"first set a strong validation technique \" and after that we should go systematic. I am happy 😊 that I got to learn this much.",
      "votes": null
    },
    {
      "id": "543206",
      "postDate": "06/04/2019 11:28:43",
      "content": "<p>Congrats! </p>\n\n<blockquote>\n  <p>Genetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.</p>\n</blockquote>\n\n<p>Could you elaborate more on Genetic Process, it's my first time learn about it. Maybe a reference link? What does it mean \"ExtraTree is used for individual\"? </p>\n\n<p>What is the reason behind sqrt(target)?</p>\n\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Congrats! \n\n&gt; Genetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.\n\nCould you elaborate more on Genetic Process, it's my first time learn about it. Maybe a reference link? What does it mean \"ExtraTree is used for individual\"? \n\nWhat is the reason behind sqrt(target)?\n\nThanks a lot!",
      "votes": null
    },
    {
      "id": "543212",
      "postDate": "06/04/2019 11:32:01",
      "content": "<p>As I understand, STFT is a two-dimensional data, how to you use it as features? Serialize it or use some specific spectrum range, or through aggregation? Thanks!</p>",
      "rawMarkdown": "As I understand, STFT is a two-dimensional data, how to you use it as features? Serialize it or use some specific spectrum range, or through aggregation? Thanks!",
      "votes": null
    },
    {
      "id": "543341",
      "postDate": "06/04/2019 13:35:25",
      "content": "<p>Great work Teammate !</p>",
      "rawMarkdown": "Great work Teammate !",
      "votes": null
    },
    {
      "id": "543348",
      "postDate": "06/04/2019 13:40:39",
      "content": "<p>Congrats and thanks for sharing. Interesting how optimizing MSE loss improved MAE metric. Maybe its an effect of the target transformation.</p>",
      "rawMarkdown": "Congrats and thanks for sharing. Interesting how optimizing MSE loss improved MAE metric. Maybe its an effect of the target transformation.",
      "votes": null
    },
    {
      "id": "543352",
      "postDate": "06/04/2019 13:42:38",
      "content": "<p>Yes, <a href=\"/stecasasso\">@stecasasso</a> made us revisit mse as objective.  He found it was working great in stacking.</p>",
      "rawMarkdown": "Yes, @stecasasso made us revisit mse as objective.  He found it was working great in stacking.",
      "votes": null
    },
    {
      "id": "543407",
      "postDate": "06/04/2019 14:14:31",
      "content": "<p>To be more precise, MSE worked better on our weighted version of the CV. On the unweighted CV (i.e. train set) huber/gamma were working slightly better. Lightgbm MAE loss was slightly worse than the others.  </p>",
      "rawMarkdown": "To be more precise, MSE worked better on our weighted version of the CV. On the unweighted CV (i.e. train set) huber/gamma were working slightly better. Lightgbm MAE loss was slightly worse than the others.",
      "votes": null
    },
    {
      "id": "543719",
      "postDate": "06/04/2019 18:07:06",
      "content": "<p>Thanks!</p>\n\n<p>GP\nI recommend this reading: \n<a href=\"https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html\">https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html</a>\nI used ExtraTreeRegressor to evaluate feature sets (individuals)</p>\n\n<p>sqrt(target)\nIt encourages predictions at high(er) TTF as it penalizes less the tails at training time.</p>\n\n<p>STFT\nCorrect, the data is 2D in (freq, time domain).\nI binned the frequency domain and derived statistics in the time domain such as trends, std, min, max, etc. I did this separately for imag, real and abs.</p>",
      "rawMarkdown": "Thanks!\n\nGP\nI recommend this reading: \n[https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html](https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html)\nI used ExtraTreeRegressor to evaluate feature sets (individuals)\n\nsqrt(target)\nIt encourages predictions at high(er) TTF as it penalizes less the tails at training time.\n\nSTFT\nCorrect, the data is 2D in (freq, time domain).\nI binned the frequency domain and derived statistics in the time domain such as trends, std, min, max, etc. I did this separately for imag, real and abs.",
      "votes": null
    },
    {
      "id": "543893",
      "postDate": "06/04/2019 23:14:52",
      "content": "<p>Congrats, thank you for sharing!</p>",
      "rawMarkdown": "Congrats, thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543116,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 10:38:47",
      "content": "<p>It was a pleasure to team with you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543184,
          "author_name": "karanjakhar",
          "author_url": "",
          "post_date": "06/04/2019 11:17:32",
          "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> I have learned a lot from you. I am still in beginning stage of learning. In this competition I didn't go systematic. And I have marked your words <a href=\"/cpmpml\">@cpmpml</a> \"first set a strong validation technique \" and after that we should go systematic. I am happy 😊 that I got to learn this much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543174,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/04/2019 11:13:01",
      "content": "<p>Thanks for sharing. I was trying all this stuff but I was not confident enough. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543206,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "06/04/2019 11:28:43",
      "content": "<p>Congrats! </p>\n\n<blockquote>\n  <p>Genetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.</p>\n</blockquote>\n\n<p>Could you elaborate more on Genetic Process, it's my first time learn about it. Maybe a reference link? What does it mean \"ExtraTree is used for individual\"? </p>\n\n<p>What is the reason behind sqrt(target)?</p>\n\n<p>Thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543212,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "06/04/2019 11:32:01",
          "content": "<p>As I understand, STFT is a two-dimensional data, how to you use it as features? Serialize it or use some specific spectrum range, or through aggregation? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543719,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "06/04/2019 18:07:06",
          "content": "<p>Thanks!</p>\n\n<p>GP\nI recommend this reading: \n<a href=\"https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html\">https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html</a>\nI used ExtraTreeRegressor to evaluate feature sets (individuals)</p>\n\n<p>sqrt(target)\nIt encourages predictions at high(er) TTF as it penalizes less the tails at training time.</p>\n\n<p>STFT\nCorrect, the data is 2D in (freq, time domain).\nI binned the frequency domain and derived statistics in the time domain such as trends, std, min, max, etc. I did this separately for imag, real and abs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543341,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "06/04/2019 13:35:25",
      "content": "<p>Great work Teammate !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543348,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 13:40:39",
      "content": "<p>Congrats and thanks for sharing. Interesting how optimizing MSE loss improved MAE metric. Maybe its an effect of the target transformation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543352,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2019 13:42:38",
          "content": "<p>Yes, <a href=\"/stecasasso\">@stecasasso</a> made us revisit mse as objective.  He found it was working great in stacking.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543407,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "06/04/2019 14:14:31",
          "content": "<p>To be more precise, MSE worked better on our weighted version of the CV. On the unweighted CV (i.e. train set) huber/gamma were working slightly better. Lightgbm MAE loss was slightly worse than the others.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543893,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "06/04/2019 23:14:52",
      "content": "<p>Congrats, thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543079": "First of all, let me congratulate with my teammates: \n@cpmpml  : thanks for accepting my invitation and for teaching us so much in the last few weeks!\n@areveillon : thanks for your original ideas and your perseverance. And most important: congratulations on becoming competition master!\n\nI am going to describe mostly what has been for me the pre-CPMP era of the competition. After merging, we adopted most of the brilliant ideas that CPMP will for sure describe in his post. I let him also describe ensembling and the like.\n\n*CV*\n16 EQ-wise CV scheme. \n\n*Data augmentation*\n150k segments with 50k overlap, i.e. 3x augmentation.\n\n*Pre-processing*\nSubtracting mean of each segment.\n\n*Feature engineering*\nWhat works:   \n- simple features like percentage close to median, smart combination of median, mean, std, min, max, number of peaks with support N, autocorrelation, rolling quantiles, etc.  \n- same as above, but applying various band/high/low pass filters . \n- MFCC coefficients inspired by public [public kernel](https://www.kaggle.com/taqanori/trying-mfcc-mel-frequency-cepstral-coefficients) . \n- STFT features inspired by [public kernel](https://www.kaggle.com/tsilveira/time-frequency-analysis-with-stft)\n\n*Feature selection*\nIterative Adversarial Validation to remove top train/test discriminating features until AUC &lt;= 0.7.  \nGenetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.\n\n*Target transformation*\nSmall improvement from applying sqrt(target).\n\n*Modelling*\n1 Lightgbm + 1 MLP both using mean_squared_error loss\n\n*Post-processing*\nApplying heuristics to adjust predictions for segments with large acoustic events at around TTF ~ 0.3. Small effect as the quality of predictions increased towards the end of the competition.",
    "543116": "It was a pleasure to team with you!",
    "543174": "Thanks for sharing. I was trying all this stuff but I was not confident enough.",
    "543184": "Thanks @cpmpml I have learned a lot from you. I am still in beginning stage of learning. In this competition I didn't go systematic. And I have marked your words @cpmpml \"first set a strong validation technique \" and after that we should go systematic. I am happy 😊 that I got to learn this much.",
    "543206": "Congrats! \n\n&gt; Genetic Process to select ~50 features, using ExtraTrees for individuals and CV MAE for fitness.\n\nCould you elaborate more on Genetic Process, it's my first time learn about it. Maybe a reference link? What does it mean \"ExtraTree is used for individual\"? \n\nWhat is the reason behind sqrt(target)?\n\nThanks a lot!",
    "543212": "As I understand, STFT is a two-dimensional data, how to you use it as features? Serialize it or use some specific spectrum range, or through aggregation? Thanks!",
    "543341": "Great work Teammate !",
    "543348": "Congrats and thanks for sharing. Interesting how optimizing MSE loss improved MAE metric. Maybe its an effect of the target transformation.",
    "543352": "Yes, @stecasasso made us revisit mse as objective.  He found it was working great in stacking.",
    "543407": "To be more precise, MSE worked better on our weighted version of the CV. On the unweighted CV (i.e. train set) huber/gamma were working slightly better. Lightgbm MAE loss was slightly worse than the others.",
    "543719": "Thanks!\n\nGP\nI recommend this reading: \n[https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html](https://www.kdnuggets.com/2017/11/rapidminer-evolutionary-algorithms-feature-selection.html)\nI used ExtraTreeRegressor to evaluate feature sets (individuals)\n\nsqrt(target)\nIt encourages predictions at high(er) TTF as it penalizes less the tails at training time.\n\nSTFT\nCorrect, the data is 2D in (freq, time domain).\nI binned the frequency domain and derived statistics in the time domain such as trends, std, min, max, etc. I did this separately for imag, real and abs.",
    "543893": "Congrats, thank you for sharing!"
  },
  "source": "meta"
}