{
  "id": 94450,
  "title": "19-th place write up",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94450",
  "author_name": "",
  "post_date": "2019-06-04T15:23:52.142550200Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you for organizers and congratulations to all the participants, it was very difficult competition with unstable public LB. Also thank you to team mates for tackling this difficult problem together.</p>\n\n<p>We have chosen 2 final submissions as </p>\n\n<p>(A). Scale ttf value based on the test dataset ttf distribution prediction given in the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">discussion</a>\n  - original prediction is done by GBDT models (XGBoost, LightGBM, CatBoost)</p>\n\n<p>(B). Ensemble of 5 models, without ttf rescaling.\n  - GBDT models (XGBoost, LightGBM, CatBoost) + NN models (1D CNN + 2D CNN)</p>\n\n<p>Scores were (A): 2.401 and (B): 2.599 respectively.</p>\n\n<p>I was mainly working on NN part in the team, so I will write up mainly 2nd submission. (As a result, 2nd submission score was quite low compared to 1st submission, so this is just as a write up, NOT the solution (A) that wins 19-th place).\nSolution (A) will be posted by other team mates later!\n[EDIT] Please refer <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94594\">19th place solution (GBDT + post-processing)</a></p>\n\n<h2>CV strategy</h2>\n\n<p>We were validating using Group K-fold where group id is assigned to each quake.\nIt returns quite unstable result in each fold, because the validation dataset quake ttf distribution differs a lot.\nLater in the competition we splitted quake id manually to distribute ttf as even as possible to make training a little bit more stable.</p>\n\n<h2>NN model</h2>\n\n<p>We used 2 models.\n - 1D-CNN: directly handle 150000 data is difficult in terms of computation power. So the feature is calculated by sliding window to make small 1d series data to process on 1d cnn.\n - 2D-CNN: 150000 point data is converted to 2d array using STFT, and process it on 2d cnn.</p>\n\n<h2>Data processing</h2>\n\n<ul>\n<li>high pass filter --&gt; wavelet denoising</li>\n<li>clip acoustic data value to (-20, 20): surprisingly, it seems that even we clip the big value, the loss still decreases and clipping value makes the feature looks more stable.</li>\n</ul>\n\n<h3>artifact classification</h3>\n\n<p>We created a classifier which predicts 4096(4095) record jump by simple 2-layer MLP model given window width=20.\nWe could get high accuracy to predict record point, this is used to determine start point of STFT preprocessing for 2D-CNN.</p>\n\n<h3>Remove after quake segment</h3>\n\n<p>Even after quake happens (big acoustic_data shake ends) some segments (around 3 segment) have ttf = 0.\nIt may be outlier to make training difficult, and I removed these data from training. </p>\n\n<h2>Prediction target</h2>\n\n<p>For 2D-CNN, we trained following 3 labels at the same time.\n - ttf: time to failure\n - tsf: time since failure\n - tqt: total time quake (ttf value at the beginning of quake)</p>\n\n<p>tsf and tqt feature is calculated from ttf information.\nI thought learning multiple label works as regularizing effect. But it was not so significant to the performance.</p>\n\n<h2>Data augmentation</h2>\n\n<p>I thought it is quite important to reduce overfitting in this competition, so tried many kinds of data augmentation methods.</p>\n\n<h3>We tried and seems worked a bit (adopted)</h3>\n\n<ul>\n<li>adding noise</li>\n<li>flip (flip along mean)</li>\n<li>time flip</li>\n<li>cutout on 1d raw wave data</li>\n<li>cutout on STFT time domain</li>\n</ul>\n\n<h3>We tried but not worked (not adopted)</h3>\n\n<ul>\n<li>cutout on STFT frequency domain</li>\n<li>wave shift (enlarge/shrink on time domain)</li>\n</ul>\n\n<h2>Model pretraining with p4581</h2>\n\n<p>We wanted to utilize p4581 data. After I checked the data, it contains too long quake or too short quake. So I removed quake data with quake time &lt;6sec and &gt;18sec.\n2D-CNN model is trained with this data, and its weight is used to fine-tune with competition dataset.\nActually its effect to performance was not so much.</p>\n\n<h2>Hyper parameter tuning</h2>\n\n<p>We used <a href=\"https://github.com/pfnet/optuna\">optuna</a> for hyper parameter tuning.</p>\n\n<h2>Training performance</h2>\n\n<p>We trained model with 5-group K fold CV with different 5 seeds (total 25 models).</p>\n\n<ul>\n<li>1D CNN model: CV oof MAE around 1.93</li>\n<li>2D CNN model: CV oof MAE around 2.00</li>\n</ul>\n\n<h2>Ensemble/Stacking</h2>\n\n<p>We tried 2 ways for ensembling model's prediction\n - simply taking mean (submitted (B)): private LB 2.522\n - stacking using <code>BayesianRegression</code> (not submitted): private LB 2.599</p>",
  "messages": [
    {
      "id": "543533",
      "postDate": "06/04/2019 15:23:52",
      "content": "<p>Thank you for organizers and congratulations to all the participants, it was very difficult competition with unstable public LB. Also thank you to team mates for tackling this difficult problem together.</p>\n\n<p>We have chosen 2 final submissions as </p>\n\n<p>(A). Scale ttf value based on the test dataset ttf distribution prediction given in the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">discussion</a>\n  - original prediction is done by GBDT models (XGBoost, LightGBM, CatBoost)</p>\n\n<p>(B). Ensemble of 5 models, without ttf rescaling.\n  - GBDT models (XGBoost, LightGBM, CatBoost) + NN models (1D CNN + 2D CNN)</p>\n\n<p>Scores were (A): 2.401 and (B): 2.599 respectively.</p>\n\n<p>I was mainly working on NN part in the team, so I will write up mainly 2nd submission. (As a result, 2nd submission score was quite low compared to 1st submission, so this is just as a write up, NOT the solution (A) that wins 19-th place).\nSolution (A) will be posted by other team mates later!\n[EDIT] Please refer <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94594\">19th place solution (GBDT + post-processing)</a></p>\n\n<h2>CV strategy</h2>\n\n<p>We were validating using Group K-fold where group id is assigned to each quake.\nIt returns quite unstable result in each fold, because the validation dataset quake ttf distribution differs a lot.\nLater in the competition we splitted quake id manually to distribute ttf as even as possible to make training a little bit more stable.</p>\n\n<h2>NN model</h2>\n\n<p>We used 2 models.\n - 1D-CNN: directly handle 150000 data is difficult in terms of computation power. So the feature is calculated by sliding window to make small 1d series data to process on 1d cnn.\n - 2D-CNN: 150000 point data is converted to 2d array using STFT, and process it on 2d cnn.</p>\n\n<h2>Data processing</h2>\n\n<ul>\n<li>high pass filter --&gt; wavelet denoising</li>\n<li>clip acoustic data value to (-20, 20): surprisingly, it seems that even we clip the big value, the loss still decreases and clipping value makes the feature looks more stable.</li>\n</ul>\n\n<h3>artifact classification</h3>\n\n<p>We created a classifier which predicts 4096(4095) record jump by simple 2-layer MLP model given window width=20.\nWe could get high accuracy to predict record point, this is used to determine start point of STFT preprocessing for 2D-CNN.</p>\n\n<h3>Remove after quake segment</h3>\n\n<p>Even after quake happens (big acoustic_data shake ends) some segments (around 3 segment) have ttf = 0.\nIt may be outlier to make training difficult, and I removed these data from training. </p>\n\n<h2>Prediction target</h2>\n\n<p>For 2D-CNN, we trained following 3 labels at the same time.\n - ttf: time to failure\n - tsf: time since failure\n - tqt: total time quake (ttf value at the beginning of quake)</p>\n\n<p>tsf and tqt feature is calculated from ttf information.\nI thought learning multiple label works as regularizing effect. But it was not so significant to the performance.</p>\n\n<h2>Data augmentation</h2>\n\n<p>I thought it is quite important to reduce overfitting in this competition, so tried many kinds of data augmentation methods.</p>\n\n<h3>We tried and seems worked a bit (adopted)</h3>\n\n<ul>\n<li>adding noise</li>\n<li>flip (flip along mean)</li>\n<li>time flip</li>\n<li>cutout on 1d raw wave data</li>\n<li>cutout on STFT time domain</li>\n</ul>\n\n<h3>We tried but not worked (not adopted)</h3>\n\n<ul>\n<li>cutout on STFT frequency domain</li>\n<li>wave shift (enlarge/shrink on time domain)</li>\n</ul>\n\n<h2>Model pretraining with p4581</h2>\n\n<p>We wanted to utilize p4581 data. After I checked the data, it contains too long quake or too short quake. So I removed quake data with quake time &lt;6sec and &gt;18sec.\n2D-CNN model is trained with this data, and its weight is used to fine-tune with competition dataset.\nActually its effect to performance was not so much.</p>\n\n<h2>Hyper parameter tuning</h2>\n\n<p>We used <a href=\"https://github.com/pfnet/optuna\">optuna</a> for hyper parameter tuning.</p>\n\n<h2>Training performance</h2>\n\n<p>We trained model with 5-group K fold CV with different 5 seeds (total 25 models).</p>\n\n<ul>\n<li>1D CNN model: CV oof MAE around 1.93</li>\n<li>2D CNN model: CV oof MAE around 2.00</li>\n</ul>\n\n<h2>Ensemble/Stacking</h2>\n\n<p>We tried 2 ways for ensembling model's prediction\n - simply taking mean (submitted (B)): private LB 2.522\n - stacking using <code>BayesianRegression</code> (not submitted): private LB 2.599</p>",
      "rawMarkdown": "Thank you for organizers and congratulations to all the participants, it was very difficult competition with unstable public LB. Also thank you to team mates for tackling this difficult problem together.\n\nWe have chosen 2 final submissions as \n\n(A). Scale ttf value based on the test dataset ttf distribution prediction given in the [discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844)\n  - original prediction is done by GBDT models (XGBoost, LightGBM, CatBoost)\n\n(B). Ensemble of 5 models, without ttf rescaling.\n  - GBDT models (XGBoost, LightGBM, CatBoost) + NN models (1D CNN + 2D CNN)\n\nScores were (A): 2.401 and (B): 2.599 respectively.\n\nI was mainly working on NN part in the team, so I will write up mainly 2nd submission. (As a result, 2nd submission score was quite low compared to 1st submission, so this is just as a write up, NOT the solution (A) that wins 19-th place).\nSolution (A) will be posted by other team mates later!\n[EDIT] Please refer [19th place solution (GBDT + post-processing)](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94594)\n\n## CV strategy\nWe were validating using Group K-fold where group id is assigned to each quake.\nIt returns quite unstable result in each fold, because the validation dataset quake ttf distribution differs a lot.\nLater in the competition we splitted quake id manually to distribute ttf as even as possible to make training a little bit more stable.\n\n## NN model\nWe used 2 models.\n - 1D-CNN: directly handle 150000 data is difficult in terms of computation power. So the feature is calculated by sliding window to make small 1d series data to process on 1d cnn.\n - 2D-CNN: 150000 point data is converted to 2d array using STFT, and process it on 2d cnn.\n\n\n## Data processing\n - high pass filter --&gt; wavelet denoising\n - clip acoustic data value to (-20, 20): surprisingly, it seems that even we clip the big value, the loss still decreases and clipping value makes the feature looks more stable.\n\n### artifact classification\nWe created a classifier which predicts 4096(4095) record jump by simple 2-layer MLP model given window width=20.\nWe could get high accuracy to predict record point, this is used to determine start point of STFT preprocessing for 2D-CNN.\n\n### Remove after quake segment\nEven after quake happens (big acoustic_data shake ends) some segments (around 3 segment) have ttf = 0.\nIt may be outlier to make training difficult, and I removed these data from training. \n\n## Prediction target\nFor 2D-CNN, we trained following 3 labels at the same time.\n - ttf: time to failure\n - tsf: time since failure\n - tqt: total time quake (ttf value at the beginning of quake)\n\ntsf and tqt feature is calculated from ttf information.\nI thought learning multiple label works as regularizing effect. But it was not so significant to the performance.\n\n## Data augmentation\nI thought it is quite important to reduce overfitting in this competition, so tried many kinds of data augmentation methods.\n\n### We tried and seems worked a bit (adopted)\n - adding noise\n - flip (flip along mean)\n - time flip\n - cutout on 1d raw wave data\n - cutout on STFT time domain\n\n### We tried but not worked (not adopted)\n - cutout on STFT frequency domain\n - wave shift (enlarge/shrink on time domain)\n\n## Model pretraining with p4581\nWe wanted to utilize p4581 data. After I checked the data, it contains too long quake or too short quake. So I removed quake data with quake time &lt;6sec and &gt;18sec.\n2D-CNN model is trained with this data, and its weight is used to fine-tune with competition dataset.\nActually its effect to performance was not so much.\n\n## Hyper parameter tuning\nWe used [optuna](https://github.com/pfnet/optuna) for hyper parameter tuning.\n\n## Training performance\nWe trained model with 5-group K fold CV with different 5 seeds (total 25 models).\n\n - 1D CNN model: CV oof MAE around 1.93\n - 2D CNN model: CV oof MAE around 2.00\n\n## Ensemble/Stacking\nWe tried 2 ways for ensembling model's prediction\n - simply taking mean (submitted (B)): private LB 2.522\n - stacking using `BayesianRegression` (not submitted): private LB 2.599",
      "votes": null
    },
    {
      "id": "543549",
      "postDate": "06/04/2019 15:32:18",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": null
    },
    {
      "id": "543550",
      "postDate": "06/04/2019 15:33:37",
      "content": "<p>Thank you and Congrats to you too!\nI would like to know your approach as well :)</p>",
      "rawMarkdown": "Thank you and Congrats to you too!\nI would like to know your approach as well :)",
      "votes": null
    },
    {
      "id": "543563",
      "postDate": "06/04/2019 15:42:10",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!",
      "votes": null
    },
    {
      "id": "543865",
      "postDate": "06/04/2019 22:32:22",
      "content": "<p>Thank you for sharing. I found you created 2 final submission with great diversity. I suppose this led to the good results.  Well done!!</p>",
      "rawMarkdown": "Thank you for sharing. I found you created 2 final submission with great diversity. I suppose this led to the good results.  Well done!!",
      "votes": null
    },
    {
      "id": "543867",
      "postDate": "06/04/2019 22:36:01",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "544035",
      "postDate": "06/05/2019 03:59:36",
      "content": "<p>Congratulations! Thanks for sharing :) </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543549,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:32:18",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543550,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "06/04/2019 15:33:37",
          "content": "<p>Thank you and Congrats to you too!\nI would like to know your approach as well :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543563,
      "author_name": "a45632",
      "author_url": "",
      "post_date": "06/04/2019 15:42:10",
      "content": "<p>Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543865,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/04/2019 22:32:22",
      "content": "<p>Thank you for sharing. I found you created 2 final submission with great diversity. I suppose this led to the good results.  Well done!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543867,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "06/04/2019 22:36:01",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544035,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "06/05/2019 03:59:36",
      "content": "<p>Congratulations! Thanks for sharing :) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543533": "Thank you for organizers and congratulations to all the participants, it was very difficult competition with unstable public LB. Also thank you to team mates for tackling this difficult problem together.\n\nWe have chosen 2 final submissions as \n\n(A). Scale ttf value based on the test dataset ttf distribution prediction given in the [discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844)\n  - original prediction is done by GBDT models (XGBoost, LightGBM, CatBoost)\n\n(B). Ensemble of 5 models, without ttf rescaling.\n  - GBDT models (XGBoost, LightGBM, CatBoost) + NN models (1D CNN + 2D CNN)\n\nScores were (A): 2.401 and (B): 2.599 respectively.\n\nI was mainly working on NN part in the team, so I will write up mainly 2nd submission. (As a result, 2nd submission score was quite low compared to 1st submission, so this is just as a write up, NOT the solution (A) that wins 19-th place).\nSolution (A) will be posted by other team mates later!\n[EDIT] Please refer [19th place solution (GBDT + post-processing)](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94594)\n\n## CV strategy\nWe were validating using Group K-fold where group id is assigned to each quake.\nIt returns quite unstable result in each fold, because the validation dataset quake ttf distribution differs a lot.\nLater in the competition we splitted quake id manually to distribute ttf as even as possible to make training a little bit more stable.\n\n## NN model\nWe used 2 models.\n - 1D-CNN: directly handle 150000 data is difficult in terms of computation power. So the feature is calculated by sliding window to make small 1d series data to process on 1d cnn.\n - 2D-CNN: 150000 point data is converted to 2d array using STFT, and process it on 2d cnn.\n\n\n## Data processing\n - high pass filter --&gt; wavelet denoising\n - clip acoustic data value to (-20, 20): surprisingly, it seems that even we clip the big value, the loss still decreases and clipping value makes the feature looks more stable.\n\n### artifact classification\nWe created a classifier which predicts 4096(4095) record jump by simple 2-layer MLP model given window width=20.\nWe could get high accuracy to predict record point, this is used to determine start point of STFT preprocessing for 2D-CNN.\n\n### Remove after quake segment\nEven after quake happens (big acoustic_data shake ends) some segments (around 3 segment) have ttf = 0.\nIt may be outlier to make training difficult, and I removed these data from training. \n\n## Prediction target\nFor 2D-CNN, we trained following 3 labels at the same time.\n - ttf: time to failure\n - tsf: time since failure\n - tqt: total time quake (ttf value at the beginning of quake)\n\ntsf and tqt feature is calculated from ttf information.\nI thought learning multiple label works as regularizing effect. But it was not so significant to the performance.\n\n## Data augmentation\nI thought it is quite important to reduce overfitting in this competition, so tried many kinds of data augmentation methods.\n\n### We tried and seems worked a bit (adopted)\n - adding noise\n - flip (flip along mean)\n - time flip\n - cutout on 1d raw wave data\n - cutout on STFT time domain\n\n### We tried but not worked (not adopted)\n - cutout on STFT frequency domain\n - wave shift (enlarge/shrink on time domain)\n\n## Model pretraining with p4581\nWe wanted to utilize p4581 data. After I checked the data, it contains too long quake or too short quake. So I removed quake data with quake time &lt;6sec and &gt;18sec.\n2D-CNN model is trained with this data, and its weight is used to fine-tune with competition dataset.\nActually its effect to performance was not so much.\n\n## Hyper parameter tuning\nWe used [optuna](https://github.com/pfnet/optuna) for hyper parameter tuning.\n\n## Training performance\nWe trained model with 5-group K fold CV with different 5 seeds (total 25 models).\n\n - 1D CNN model: CV oof MAE around 1.93\n - 2D CNN model: CV oof MAE around 2.00\n\n## Ensemble/Stacking\nWe tried 2 ways for ensembling model's prediction\n - simply taking mean (submitted (B)): private LB 2.522\n - stacking using `BayesianRegression` (not submitted): private LB 2.599",
    "543549": "Congrats and thanks for sharing!",
    "543550": "Thank you and Congrats to you too!\nI would like to know your approach as well :)",
    "543563": "Great job!",
    "543865": "Thank you for sharing. I found you created 2 final submission with great diversity. I suppose this led to the good results.  Well done!!",
    "543867": "Thank you :)",
    "544035": "Congratulations! Thanks for sharing :)"
  },
  "source": "meta"
}