{
  "id": 94718,
  "title": "8th solution. XGBoost and Dataset balancing",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/ivan-bagmut-8th-solution-xgboost-and-dataset-balan",
  "author_name": "",
  "post_date": "2019-06-06T12:23:10.924726800Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Many thanks organizers for such interesting competition.</p>\n\n<p>And thanks authors for their very useful kernels:\n<a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a>\n<a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction</a>\n<a href=\"https://www.kaggle.com/allunia/shaking-earth\">https://www.kaggle.com/allunia/shaking-earth</a>\n<a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392</a></p>\n\n<p>I try different models: \"time series models\"- WaveNet, LSTM, ResNet-1d, FCN-1d.\n\"Features models\" - XGBRegressor, DNN, SVR.\nAnd XGBoost model was the best.\nThis is hyper parameters:\nmodel = xgb.XGBRegressor(booster='dart',\n                         tree_method='hist',\n                         n_estimators=100000,\n                         learning_rate=0.01,\n                         max_depth=3,\n                         subsample=0.9,\n                         colsample_bytree=0.5,\n                         reg_lambda=1,\n                         gamma = 1)</p>\n\n<p>And I noticed that the original training data set is unbalanced - there is few data with a time to failure more than 8 seconds.\nTherefore, when creating my own dataset for thrain the model, I used data from 2, 7, 14 \"long\" quakes \"more\" than others, and used 4th quake for validation. \nThis improved the prediction for longer times to failure. </p>",
  "messages": [
    {
      "id": "546275",
      "postDate": "06/06/2019 12:23:10",
      "content": "<p>Many thanks organizers for such interesting competition.</p>\n\n<p>And thanks authors for their very useful kernels:\n<a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a>\n<a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction</a>\n<a href=\"https://www.kaggle.com/allunia/shaking-earth\">https://www.kaggle.com/allunia/shaking-earth</a>\n<a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392</a></p>\n\n<p>I try different models: \"time series models\"- WaveNet, LSTM, ResNet-1d, FCN-1d.\n\"Features models\" - XGBRegressor, DNN, SVR.\nAnd XGBoost model was the best.\nThis is hyper parameters:\nmodel = xgb.XGBRegressor(booster='dart',\n                         tree_method='hist',\n                         n_estimators=100000,\n                         learning_rate=0.01,\n                         max_depth=3,\n                         subsample=0.9,\n                         colsample_bytree=0.5,\n                         reg_lambda=1,\n                         gamma = 1)</p>\n\n<p>And I noticed that the original training data set is unbalanced - there is few data with a time to failure more than 8 seconds.\nTherefore, when creating my own dataset for thrain the model, I used data from 2, 7, 14 \"long\" quakes \"more\" than others, and used 4th quake for validation. \nThis improved the prediction for longer times to failure. </p>",
      "rawMarkdown": "Many thanks organizers for such interesting competition.\n\nAnd thanks authors for their very useful kernels:\nhttps://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\nhttps://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\nhttps://www.kaggle.com/allunia/shaking-earth\nhttps://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\n\nI try different models: \"time series models\"- WaveNet, LSTM, ResNet-1d, FCN-1d.\n\"Features models\" - XGBRegressor, DNN, SVR.\nAnd XGBoost model was the best.\nThis is hyper parameters:\nmodel = xgb.XGBRegressor(booster='dart',\n                         tree_method='hist',\n                         n_estimators=100000,\n                         learning_rate=0.01,\n                         max_depth=3,\n                         subsample=0.9,\n                         colsample_bytree=0.5,\n                         reg_lambda=1,\n                         gamma = 1)\n\nAnd I noticed that the original training data set is unbalanced - there is few data with a time to failure more than 8 seconds.\nTherefore, when creating my own dataset for thrain the model, I used data from 2, 7, 14 \"long\" quakes \"more\" than others, and used 4th quake for validation. \nThis improved the prediction for longer times to failure.",
      "votes": null
    },
    {
      "id": "546324",
      "postDate": "06/06/2019 13:21:35",
      "content": "<p>Thanks for sharing and congrats on result!</p>",
      "rawMarkdown": "Thanks for sharing and congrats on result!",
      "votes": null
    },
    {
      "id": "546332",
      "postDate": "06/06/2019 13:33:20",
      "content": "<p>Thank you too, congratulations! I think you have the most robust model of all the participants, \nbecause in the public and private leaderboard you are in the top. ))</p>",
      "rawMarkdown": "Thank you too, congratulations! I think you have the most robust model of all the participants, \nbecause in the public and private leaderboard you are in the top. ))",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 546324,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/06/2019 13:21:35",
      "content": "<p>Thanks for sharing and congrats on result!</p>",
      "votes": null,
      "replies": [
        {
          "id": 546332,
          "author_name": "ivanbagmut",
          "author_url": "",
          "post_date": "06/06/2019 13:33:20",
          "content": "<p>Thank you too, congratulations! I think you have the most robust model of all the participants, \nbecause in the public and private leaderboard you are in the top. ))</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "546275": "Many thanks organizers for such interesting competition.\n\nAnd thanks authors for their very useful kernels:\nhttps://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\nhttps://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\nhttps://www.kaggle.com/allunia/shaking-earth\nhttps://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\n\nI try different models: \"time series models\"- WaveNet, LSTM, ResNet-1d, FCN-1d.\n\"Features models\" - XGBRegressor, DNN, SVR.\nAnd XGBoost model was the best.\nThis is hyper parameters:\nmodel = xgb.XGBRegressor(booster='dart',\n                         tree_method='hist',\n                         n_estimators=100000,\n                         learning_rate=0.01,\n                         max_depth=3,\n                         subsample=0.9,\n                         colsample_bytree=0.5,\n                         reg_lambda=1,\n                         gamma = 1)\n\nAnd I noticed that the original training data set is unbalanced - there is few data with a time to failure more than 8 seconds.\nTherefore, when creating my own dataset for thrain the model, I used data from 2, 7, 14 \"long\" quakes \"more\" than others, and used 4th quake for validation. \nThis improved the prediction for longer times to failure.",
    "546324": "Thanks for sharing and congrats on result!",
    "546332": "Thank you too, congratulations! I think you have the most robust model of all the participants, \nbecause in the public and private leaderboard you are in the top. ))"
  },
  "source": "meta"
}