{
  "id": 94352,
  "title": "Short summary for our approach (public 5th -> private 212th)",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/kaggler-ja-shake-it-up-short-summary-for-our-appro",
  "author_name": "",
  "post_date": "2019-06-04T03:31:33.256707800Z",
  "votes": 28,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Congrats to all winners. Of course we are a little disappointed at the result, but we would like to share our approach. We believe our experience can be lessons for us, and also for some other people.</p>\n\n<h1>Summary</h1>\n\n<p>Our best public LB score 1.259 is provided by ensemble of 6 LightGBM models.</p>\n\n<ul>\n<li>Use features of “<a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">Masters Final Project</a>” (Its size is over 800)</li>\n<li>Reduced the size up to 449 by:\n<ul><li>Kolmogorov–Smirnov test</li>\n<li>scipy.stats.pearsonr</li>\n<li>adversarial validation</li></ul></li>\n<li>Train LightGBM models with:\n<ul><li>gamma regression</li>\n<li>rounds different from 10000 to 20000 (separated by 2000)</li></ul></li>\n</ul>\n\n<h1>CV Strategy</h1>\n\n<ul>\n<li>KFold</li>\n<li>8 fold</li>\n<li>shuffle=True</li>\n</ul>\n\n<p>When shuffle=False, cv score in each fold fluctuates a lot <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366\">as olivier says</a>. We felt it’s not appropriate for good cv, so I decided to use shuffle=True.</p>\n\n<h1>Models</h1>\n\n<p>We tried {LightGBM, XGBoost, Catboost, MLP, RNN}. And LightGBM is the best model for us, the others couldn’t be better.</p>\n\n<h2>Hyperparams</h2>\n\n<p>Gamma regression is suitable for this competition. By just changing objective from <code>regression</code> to <code>gamma</code>, we can get great gain in public LB score. </p>\n\n<h1>Ensemble</h1>\n\n<p>We utilize “rounds averaging” for ensemble.</p>\n\n<p>Since we use the augmented features, early stopping doesn’t work properly. We can’t see what is the best round for LightGBM, so we take the average of 6 predictions generated from different rounds models (from 10000 to 20000 separately).</p>\n\n<h1>Submission Strategy</h1>\n\n<p>We prepared the following 2 submission:</p>\n\n<ul>\n<li>Almost the same model described above\n<ul><li>Only one difference is the number of rounds averaging. We use 11 model (from 10000 to 20000 separately)</li>\n<li>public: 1.260, private 2.49606</li></ul></li>\n<li>Feature dropped version\n<ul><li>We removed all features named ‘mean’ so that we can avoid overfit</li>\n<li>public: 1.273, private 2.49699</li></ul></li>\n</ul>\n\n<h1>What didn’t work for us</h1>\n\n<ul>\n<li>Convert targets to sqrt(targets) in order that targets follow normal distribution</li>\n<li>Remove some train data by adversarial validation</li>\n<li>Use meta feature (prediction value as a feature)</li>\n<li>Use dart as LightGBM hyperparams</li>\n<li>Stacking</li>\n<li>Seed averaging</li>\n</ul>\n\n<p>Anyway, I would like to thank every participants especially for my teammates. We would like to learn more from winner solutions!</p>",
  "messages": [
    {
      "id": "542687",
      "postDate": "06/04/2019 03:31:33",
      "content": "<p>Congrats to all winners. Of course we are a little disappointed at the result, but we would like to share our approach. We believe our experience can be lessons for us, and also for some other people.</p>\n\n<h1>Summary</h1>\n\n<p>Our best public LB score 1.259 is provided by ensemble of 6 LightGBM models.</p>\n\n<ul>\n<li>Use features of “<a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">Masters Final Project</a>” (Its size is over 800)</li>\n<li>Reduced the size up to 449 by:\n<ul><li>Kolmogorov–Smirnov test</li>\n<li>scipy.stats.pearsonr</li>\n<li>adversarial validation</li></ul></li>\n<li>Train LightGBM models with:\n<ul><li>gamma regression</li>\n<li>rounds different from 10000 to 20000 (separated by 2000)</li></ul></li>\n</ul>\n\n<h1>CV Strategy</h1>\n\n<ul>\n<li>KFold</li>\n<li>8 fold</li>\n<li>shuffle=True</li>\n</ul>\n\n<p>When shuffle=False, cv score in each fold fluctuates a lot <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366\">as olivier says</a>. We felt it’s not appropriate for good cv, so I decided to use shuffle=True.</p>\n\n<h1>Models</h1>\n\n<p>We tried {LightGBM, XGBoost, Catboost, MLP, RNN}. And LightGBM is the best model for us, the others couldn’t be better.</p>\n\n<h2>Hyperparams</h2>\n\n<p>Gamma regression is suitable for this competition. By just changing objective from <code>regression</code> to <code>gamma</code>, we can get great gain in public LB score. </p>\n\n<h1>Ensemble</h1>\n\n<p>We utilize “rounds averaging” for ensemble.</p>\n\n<p>Since we use the augmented features, early stopping doesn’t work properly. We can’t see what is the best round for LightGBM, so we take the average of 6 predictions generated from different rounds models (from 10000 to 20000 separately).</p>\n\n<h1>Submission Strategy</h1>\n\n<p>We prepared the following 2 submission:</p>\n\n<ul>\n<li>Almost the same model described above\n<ul><li>Only one difference is the number of rounds averaging. We use 11 model (from 10000 to 20000 separately)</li>\n<li>public: 1.260, private 2.49606</li></ul></li>\n<li>Feature dropped version\n<ul><li>We removed all features named ‘mean’ so that we can avoid overfit</li>\n<li>public: 1.273, private 2.49699</li></ul></li>\n</ul>\n\n<h1>What didn’t work for us</h1>\n\n<ul>\n<li>Convert targets to sqrt(targets) in order that targets follow normal distribution</li>\n<li>Remove some train data by adversarial validation</li>\n<li>Use meta feature (prediction value as a feature)</li>\n<li>Use dart as LightGBM hyperparams</li>\n<li>Stacking</li>\n<li>Seed averaging</li>\n</ul>\n\n<p>Anyway, I would like to thank every participants especially for my teammates. We would like to learn more from winner solutions!</p>",
      "rawMarkdown": "Congrats to all winners. Of course we are a little disappointed at the result, but we would like to share our approach. We believe our experience can be lessons for us, and also for some other people.\n\n# Summary\n\nOur best public LB score 1.259 is provided by ensemble of 6 LightGBM models.\n\n- Use features of “[Masters Final Project](https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392)” (Its size is over 800)\n- Reduced the size up to 449 by:\n  - Kolmogorov–Smirnov test\n  - scipy.stats.pearsonr\n  - adversarial validation\n- Train LightGBM models with:\n  - gamma regression\n  - rounds different from 10000 to 20000 (separated by 2000)\n\n# CV Strategy\n\n- KFold\n- 8 fold\n- shuffle=True\n\nWhen shuffle=False, cv score in each fold fluctuates a lot [as olivier says](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366). We felt it’s not appropriate for good cv, so I decided to use shuffle=True.\n\n# Models\n\nWe tried {LightGBM, XGBoost, Catboost, MLP, RNN}. And LightGBM is the best model for us, the others couldn’t be better.\n\n## Hyperparams\n\nGamma regression is suitable for this competition. By just changing objective from `regression` to `gamma`, we can get great gain in public LB score. \n\n# Ensemble\n\nWe utilize “rounds averaging” for ensemble.\n\nSince we use the augmented features, early stopping doesn’t work properly. We can’t see what is the best round for LightGBM, so we take the average of 6 predictions generated from different rounds models (from 10000 to 20000 separately).\n\n# Submission Strategy\n\nWe prepared the following 2 submission:\n\n- Almost the same model described above\n  - Only one difference is the number of rounds averaging. We use 11 model (from 10000 to 20000 separately)\n  - public: 1.260, private 2.49606\n- Feature dropped version\n  - We removed all features named ‘mean’ so that we can avoid overfit\n  - public: 1.273, private 2.49699\n\n# What didn’t work for us\n\n- Convert targets to sqrt(targets) in order that targets follow normal distribution\n- Remove some train data by adversarial validation\n- Use meta feature (prediction value as a feature)\n- Use dart as LightGBM hyperparams\n- Stacking\n- Seed averaging\n\nAnyway, I would like to thank every participants especially for my teammates. We would like to learn more from winner solutions!",
      "votes": null
    },
    {
      "id": "542699",
      "postDate": "06/04/2019 03:39:03",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": null
    },
    {
      "id": "542706",
      "postDate": "06/04/2019 03:46:37",
      "content": "<p>Many thanks for the sharing, btw, could you explain a bit regarding \"rounds different from 10000 to 20000 (separated by 2000)\". I am sorry as i just a novice who eager to learn.  by the way when you remove the \"mean\" related features, maybe you might have better private test result if you do post processing by multiplying factor that consider the private test mean.</p>",
      "rawMarkdown": "Many thanks for the sharing, btw, could you explain a bit regarding \"rounds different from 10000 to 20000 (separated by 2000)\". I am sorry as i just a novice who eager to learn.  by the way when you remove the \"mean\" related features, maybe you might have better private test result if you do post processing by multiplying factor that consider the private test mean.",
      "votes": null
    },
    {
      "id": "542716",
      "postDate": "06/04/2019 03:57:22",
      "content": "<p>I'm sorry for the lack of explanations.</p>\n\n<p>I mean we created the following 6 model:</p>\n\n<ul>\n<li>lgbm training rounds = 10000</li>\n<li>lgbm training rounds = 12000</li>\n<li>lgbm training rounds = 14000</li>\n<li>lgbm training rounds = 16000</li>\n<li>lgbm training rounds = 18000</li>\n<li>lgbm training rounds = 20000</li>\n</ul>\n\n<p>And finally we averaged predictions of 6 models. </p>\n\n<p>Single \"lgbm training rounds = 10000\" scored 1.263 in public LB, and \"lgbm training rounds = 20000\" scored 1.262. And averaged submission gave us 1.259.</p>",
      "rawMarkdown": "I'm sorry for the lack of explanations.\n\nI mean we created the following 6 model:\n\n- lgbm training rounds = 10000\n- lgbm training rounds = 12000\n- lgbm training rounds = 14000\n- lgbm training rounds = 16000\n- lgbm training rounds = 18000\n- lgbm training rounds = 20000\n\nAnd finally we averaged predictions of 6 models. \n\nSingle \"lgbm training rounds = 10000\" scored 1.263 in public LB, and \"lgbm training rounds = 20000\" scored 1.262. And averaged submission gave us 1.259.",
      "votes": null
    },
    {
      "id": "542725",
      "postDate": "06/04/2019 04:06:34",
      "content": "<p>I learn nicely from your approach. many thanks :)</p>",
      "rawMarkdown": "I learn nicely from your approach. many thanks :)",
      "votes": null
    },
    {
      "id": "542775",
      "postDate": "06/04/2019 04:52:00",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "542983",
      "postDate": "06/04/2019 08:51:38",
      "content": "<p>Overfiting ?\nI also used features from “Masters Final Project”  ,but I selected features very simple,only depend on the importance of lgb,and I also add features about the signal frequent.\nThe most important thing is the imbalance of the data,which means the data greater than 10 is lacking. So I uesd oversampling.(If i don't use oversampling, the prediction-valid can't reach 10.)\nResult in my public LB 1.47 and my private 191th。\nMaybe you can try oversampling.</p>",
      "rawMarkdown": "Overfiting ?\nI also used features from “Masters Final Project”  ,but I selected features very simple,only depend on the importance of lgb,and I also add features about the signal frequent.\nThe most important thing is the imbalance of the data,which means the data greater than 10 is lacking. So I uesd oversampling.(If i don't use oversampling, the prediction-valid can't reach 10.)\nResult in my public LB 1.47 and my private 191th。\nMaybe you can try oversampling.",
      "votes": null
    },
    {
      "id": "543000",
      "postDate": "06/04/2019 09:01:41",
      "content": "<p>Thank you for your advice! I'll try and do late submissions.</p>",
      "rawMarkdown": "Thank you for your advice! I'll try and do late submissions.",
      "votes": null
    },
    {
      "id": "543082",
      "postDate": "06/04/2019 10:16:42",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": null
    },
    {
      "id": "543090",
      "postDate": "06/04/2019 10:21:18",
      "content": "<p>Thanks for sharing！！</p>",
      "rawMarkdown": "Thanks for sharing！！",
      "votes": null
    },
    {
      "id": "543098",
      "postDate": "06/04/2019 10:26:28",
      "content": "<p>Thanks for sharing. We learn from everything.</p>",
      "rawMarkdown": "Thanks for sharing. We learn from everything.",
      "votes": null
    },
    {
      "id": "543276",
      "postDate": "06/04/2019 12:36:03",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "543616",
      "postDate": "06/04/2019 16:26:20",
      "content": "<p>Talking about overfiting, I had the same feeling .... I proposed several submissions the more accurate I was (better results with more features !!!) the less I scored .... at the end of the day ... I selected the less accurate (and also the one with less features) and I had a good surprise when I discovered the private score. I jumped more than 1900 ranks :-)\nBut hey you still did a really great job  ...  :-)\nHope to reach your level of mastering soon ;-)</p>",
      "rawMarkdown": "Talking about overfiting, I had the same feeling .... I proposed several submissions the more accurate I was (better results with more features !!!) the less I scored .... at the end of the day ... I selected the less accurate (and also the one with less features) and I had a good surprise when I discovered the private score. I jumped more than 1900 ranks :-)\nBut hey you still did a really great job  ...  :-)\nHope to reach your level of mastering soon ;-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542699,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "06/04/2019 03:39:03",
      "content": "<p>Thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542706,
      "author_name": "arisukma",
      "author_url": "",
      "post_date": "06/04/2019 03:46:37",
      "content": "<p>Many thanks for the sharing, btw, could you explain a bit regarding \"rounds different from 10000 to 20000 (separated by 2000)\". I am sorry as i just a novice who eager to learn.  by the way when you remove the \"mean\" related features, maybe you might have better private test result if you do post processing by multiplying factor that consider the private test mean.</p>",
      "votes": null,
      "replies": [
        {
          "id": 542716,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "06/04/2019 03:57:22",
          "content": "<p>I'm sorry for the lack of explanations.</p>\n\n<p>I mean we created the following 6 model:</p>\n\n<ul>\n<li>lgbm training rounds = 10000</li>\n<li>lgbm training rounds = 12000</li>\n<li>lgbm training rounds = 14000</li>\n<li>lgbm training rounds = 16000</li>\n<li>lgbm training rounds = 18000</li>\n<li>lgbm training rounds = 20000</li>\n</ul>\n\n<p>And finally we averaged predictions of 6 models. </p>\n\n<p>Single \"lgbm training rounds = 10000\" scored 1.263 in public LB, and \"lgbm training rounds = 20000\" scored 1.262. And averaged submission gave us 1.259.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 542725,
          "author_name": "arisukma",
          "author_url": "",
          "post_date": "06/04/2019 04:06:34",
          "content": "<p>I learn nicely from your approach. many thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542775,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 04:52:00",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542983,
      "author_name": "hanakana",
      "author_url": "",
      "post_date": "06/04/2019 08:51:38",
      "content": "<p>Overfiting ?\nI also used features from “Masters Final Project”  ,but I selected features very simple,only depend on the importance of lgb,and I also add features about the signal frequent.\nThe most important thing is the imbalance of the data,which means the data greater than 10 is lacking. So I uesd oversampling.(If i don't use oversampling, the prediction-valid can't reach 10.)\nResult in my public LB 1.47 and my private 191th。\nMaybe you can try oversampling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543000,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "06/04/2019 09:01:41",
          "content": "<p>Thank you for your advice! I'll try and do late submissions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543616,
          "author_name": "bigfay",
          "author_url": "",
          "post_date": "06/04/2019 16:26:20",
          "content": "<p>Talking about overfiting, I had the same feeling .... I proposed several submissions the more accurate I was (better results with more features !!!) the less I scored .... at the end of the day ... I selected the less accurate (and also the one with less features) and I had a good surprise when I discovered the private score. I jumped more than 1900 ranks :-)\nBut hey you still did a really great job  ...  :-)\nHope to reach your level of mastering soon ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543082,
      "author_name": "shwetagoyal4",
      "author_url": "",
      "post_date": "06/04/2019 10:16:42",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543090,
      "author_name": "xxbxyae",
      "author_url": "",
      "post_date": "06/04/2019 10:21:18",
      "content": "<p>Thanks for sharing！！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543098,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/04/2019 10:26:28",
      "content": "<p>Thanks for sharing. We learn from everything.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543276,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "06/04/2019 12:36:03",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542687": "Congrats to all winners. Of course we are a little disappointed at the result, but we would like to share our approach. We believe our experience can be lessons for us, and also for some other people.\n\n# Summary\n\nOur best public LB score 1.259 is provided by ensemble of 6 LightGBM models.\n\n- Use features of “[Masters Final Project](https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392)” (Its size is over 800)\n- Reduced the size up to 449 by:\n  - Kolmogorov–Smirnov test\n  - scipy.stats.pearsonr\n  - adversarial validation\n- Train LightGBM models with:\n  - gamma regression\n  - rounds different from 10000 to 20000 (separated by 2000)\n\n# CV Strategy\n\n- KFold\n- 8 fold\n- shuffle=True\n\nWhen shuffle=False, cv score in each fold fluctuates a lot [as olivier says](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366). We felt it’s not appropriate for good cv, so I decided to use shuffle=True.\n\n# Models\n\nWe tried {LightGBM, XGBoost, Catboost, MLP, RNN}. And LightGBM is the best model for us, the others couldn’t be better.\n\n## Hyperparams\n\nGamma regression is suitable for this competition. By just changing objective from `regression` to `gamma`, we can get great gain in public LB score. \n\n# Ensemble\n\nWe utilize “rounds averaging” for ensemble.\n\nSince we use the augmented features, early stopping doesn’t work properly. We can’t see what is the best round for LightGBM, so we take the average of 6 predictions generated from different rounds models (from 10000 to 20000 separately).\n\n# Submission Strategy\n\nWe prepared the following 2 submission:\n\n- Almost the same model described above\n  - Only one difference is the number of rounds averaging. We use 11 model (from 10000 to 20000 separately)\n  - public: 1.260, private 2.49606\n- Feature dropped version\n  - We removed all features named ‘mean’ so that we can avoid overfit\n  - public: 1.273, private 2.49699\n\n# What didn’t work for us\n\n- Convert targets to sqrt(targets) in order that targets follow normal distribution\n- Remove some train data by adversarial validation\n- Use meta feature (prediction value as a feature)\n- Use dart as LightGBM hyperparams\n- Stacking\n- Seed averaging\n\nAnyway, I would like to thank every participants especially for my teammates. We would like to learn more from winner solutions!",
    "542699": "Thanks for sharing !",
    "542706": "Many thanks for the sharing, btw, could you explain a bit regarding \"rounds different from 10000 to 20000 (separated by 2000)\". I am sorry as i just a novice who eager to learn.  by the way when you remove the \"mean\" related features, maybe you might have better private test result if you do post processing by multiplying factor that consider the private test mean.",
    "542716": "I'm sorry for the lack of explanations.\n\nI mean we created the following 6 model:\n\n- lgbm training rounds = 10000\n- lgbm training rounds = 12000\n- lgbm training rounds = 14000\n- lgbm training rounds = 16000\n- lgbm training rounds = 18000\n- lgbm training rounds = 20000\n\nAnd finally we averaged predictions of 6 models. \n\nSingle \"lgbm training rounds = 10000\" scored 1.263 in public LB, and \"lgbm training rounds = 20000\" scored 1.262. And averaged submission gave us 1.259.",
    "542725": "I learn nicely from your approach. many thanks :)",
    "542775": "Thanks for sharing",
    "542983": "Overfiting ?\nI also used features from “Masters Final Project”  ,but I selected features very simple,only depend on the importance of lgb,and I also add features about the signal frequent.\nThe most important thing is the imbalance of the data,which means the data greater than 10 is lacking. So I uesd oversampling.(If i don't use oversampling, the prediction-valid can't reach 10.)\nResult in my public LB 1.47 and my private 191th。\nMaybe you can try oversampling.",
    "543000": "Thank you for your advice! I'll try and do late submissions.",
    "543082": "Thanks for sharing!!",
    "543090": "Thanks for sharing！！",
    "543098": "Thanks for sharing. We learn from everything.",
    "543276": "Thanks for sharing!",
    "543616": "Talking about overfiting, I had the same feeling .... I proposed several submissions the more accurate I was (better results with more features !!!) the less I scored .... at the end of the day ... I selected the less accurate (and also the one with less features) and I had a good surprise when I discovered the private score. I jumped more than 1900 ranks :-)\nBut hey you still did a really great job  ...  :-)\nHope to reach your level of mastering soon ;-)"
  },
  "source": "meta"
}