{
  "id": 94458,
  "title": "first competition thoughts",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94458",
  "author_name": "",
  "post_date": "2019-06-04T16:14:34.601622100Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This is actually my first real Kaggle competition. Although I started very late (2 weeks ago), I got pretty good results in the end (public 79th, private 40th). I cannot make it wothout the Kaggle community sharing lots of codes and ideas. Here is what I learned from this competition:\n- <strong>feature is much more important than model</strong> Plenty of people used 4194 non-overlapping samples as training data. I used those as well at the beginning. But I found that was not a good choice because cropping the training data in non-overlapping chunk is not enough. No matter how I tune the models, the ranking does not imporve much. So I looked at other kernels and found <a href=\"https://www.kaggle.com/tocha4/lanl-master-s-approach\">https://www.kaggle.com/tocha4/lanl-master-s-approach</a>. I used their features as training data and it improves A LOT.\n- <strong>A good computer is necessary and start early</strong> the features i extracted is around 30000 by 1000. It usually takes overnight to train the models and grid search the best hyper-parameters. I also moved the computation to GPU if it has GPU support. I have a very good computer (i9-9900k/64g/2080ti) so that I can do lots of computational expensive training. But due to the complex model I used, I only trained my final model once. I guess i could get better score if I have more time.</p>\n\n<p>The model I used is:\nrandom forest/knn/extra tree/ada boost/Nu SVR/light gbm/xgboost/cat boost - 8 models\nlight gbm/xgboost/cat boost - to do stacking on above 8 models\nlight gbm stacking on above 3 models</p>\n\n<p>It looks like there are some testing sets leakage. I should look at them if I started early.</p>",
  "messages": [
    {
      "id": "543605",
      "postDate": "06/04/2019 16:14:34",
      "content": "<p>This is actually my first real Kaggle competition. Although I started very late (2 weeks ago), I got pretty good results in the end (public 79th, private 40th). I cannot make it wothout the Kaggle community sharing lots of codes and ideas. Here is what I learned from this competition:\n- <strong>feature is much more important than model</strong> Plenty of people used 4194 non-overlapping samples as training data. I used those as well at the beginning. But I found that was not a good choice because cropping the training data in non-overlapping chunk is not enough. No matter how I tune the models, the ranking does not imporve much. So I looked at other kernels and found <a href=\"https://www.kaggle.com/tocha4/lanl-master-s-approach\">https://www.kaggle.com/tocha4/lanl-master-s-approach</a>. I used their features as training data and it improves A LOT.\n- <strong>A good computer is necessary and start early</strong> the features i extracted is around 30000 by 1000. It usually takes overnight to train the models and grid search the best hyper-parameters. I also moved the computation to GPU if it has GPU support. I have a very good computer (i9-9900k/64g/2080ti) so that I can do lots of computational expensive training. But due to the complex model I used, I only trained my final model once. I guess i could get better score if I have more time.</p>\n\n<p>The model I used is:\nrandom forest/knn/extra tree/ada boost/Nu SVR/light gbm/xgboost/cat boost - 8 models\nlight gbm/xgboost/cat boost - to do stacking on above 8 models\nlight gbm stacking on above 3 models</p>\n\n<p>It looks like there are some testing sets leakage. I should look at them if I started early.</p>",
      "rawMarkdown": "This is actually my first real Kaggle competition. Although I started very late (2 weeks ago), I got pretty good results in the end (public 79th, private 40th). I cannot make it wothout the Kaggle community sharing lots of codes and ideas. Here is what I learned from this competition:\n- **feature is much more important than model** Plenty of people used 4194 non-overlapping samples as training data. I used those as well at the beginning. But I found that was not a good choice because cropping the training data in non-overlapping chunk is not enough. No matter how I tune the models, the ranking does not imporve much. So I looked at other kernels and found https://www.kaggle.com/tocha4/lanl-master-s-approach. I used their features as training data and it improves A LOT.\n- **A good computer is necessary and start early** the features i extracted is around 30000 by 1000. It usually takes overnight to train the models and grid search the best hyper-parameters. I also moved the computation to GPU if it has GPU support. I have a very good computer (i9-9900k/64g/2080ti) so that I can do lots of computational expensive training. But due to the complex model I used, I only trained my final model once. I guess i could get better score if I have more time.\n\nThe model I used is:\nrandom forest/knn/extra tree/ada boost/Nu SVR/light gbm/xgboost/cat boost - 8 models\nlight gbm/xgboost/cat boost - to do stacking on above 8 models\nlight gbm stacking on above 3 models\n\nIt looks like there are some testing sets leakage. I should look at them if I started early.",
      "votes": null
    },
    {
      "id": "543622",
      "postDate": "06/04/2019 16:30:51",
      "content": "<p>Interesting that you used a 3 level stack.  Have you checked that it really improves over your second level models?  </p>",
      "rawMarkdown": "Interesting that you used a 3 level stack.  Have you checked that it really improves over your second level models?",
      "votes": null
    },
    {
      "id": "543627",
      "postDate": "06/04/2019 16:36:33",
      "content": "<p>Always having a good computer is an advantage. Specially with multiple GPUs for image classification challenges. But only having RAM are necessary in this case.</p>",
      "rawMarkdown": "Always having a good computer is an advantage. Specially with multiple GPUs for image classification challenges. But only having RAM are necessary in this case.",
      "votes": null
    },
    {
      "id": "543642",
      "postDate": "06/04/2019 16:50:19",
      "content": "<p>I tried two submission, one is 3 level stacking, the other one is 2 level stacking with smallest CV error(catboost). THe first one is better: public 1.34808 vs 1.35378; private 2.43624 vs 2.49996</p>",
      "rawMarkdown": "I tried two submission, one is 3 level stacking, the other one is 2 level stacking with smallest CV error(catboost). THe first one is better: public 1.34808 vs 1.35378; private 2.43624 vs 2.49996",
      "votes": null
    },
    {
      "id": "543654",
      "postDate": "06/04/2019 17:11:32",
      "content": "<p>training xgboost and catboost with GPU is much faster than CPU.</p>",
      "rawMarkdown": "training xgboost and catboost with GPU is much faster than CPU.",
      "votes": null
    },
    {
      "id": "543840",
      "postDate": "06/04/2019 21:53:50",
      "content": "<p>Thanks for letting me know, I will try Catboost on GPU. </p>\n\n<p>But in this competition the dataset is pretty small and even on CPU train very fast. </p>",
      "rawMarkdown": "Thanks for letting me know, I will try Catboost on GPU. \n\nBut in this competition the dataset is pretty small and even on CPU train very fast.",
      "votes": null
    },
    {
      "id": "543909",
      "postDate": "06/04/2019 23:37:25",
      "content": "<p>it also depends on how many features and samples used. I used around 30000x1000 training set. With 2080ti, catboost takes 18min for each CV folder. if I use cpu, it would be hours, I guess.</p>",
      "rawMarkdown": "it also depends on how many features and samples used. I used around 30000x1000 training set. With 2080ti, catboost takes 18min for each CV folder. if I use cpu, it would be hours, I guess.",
      "votes": null
    },
    {
      "id": "545979",
      "postDate": "06/06/2019 05:26:58",
      "content": "<p>My only problem with the computer was RAM.  I have 32 GB RAM and 30GB swap disc space, and I have 1080ti.  My computer run out of memory a few times.  </p>",
      "rawMarkdown": "My only problem with the computer was RAM.  I have 32 GB RAM and 30GB swap disc space, and I have 1080ti.  My computer run out of memory a few times.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543622,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 16:30:51",
      "content": "<p>Interesting that you used a 3 level stack.  Have you checked that it really improves over your second level models?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 543642,
          "author_name": "waylongo",
          "author_url": "",
          "post_date": "06/04/2019 16:50:19",
          "content": "<p>I tried two submission, one is 3 level stacking, the other one is 2 level stacking with smallest CV error(catboost). THe first one is better: public 1.34808 vs 1.35378; private 2.43624 vs 2.49996</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543627,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 16:36:33",
      "content": "<p>Always having a good computer is an advantage. Specially with multiple GPUs for image classification challenges. But only having RAM are necessary in this case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543654,
          "author_name": "waylongo",
          "author_url": "",
          "post_date": "06/04/2019 17:11:32",
          "content": "<p>training xgboost and catboost with GPU is much faster than CPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543840,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "06/04/2019 21:53:50",
          "content": "<p>Thanks for letting me know, I will try Catboost on GPU. </p>\n\n<p>But in this competition the dataset is pretty small and even on CPU train very fast. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543909,
          "author_name": "waylongo",
          "author_url": "",
          "post_date": "06/04/2019 23:37:25",
          "content": "<p>it also depends on how many features and samples used. I used around 30000x1000 training set. With 2080ti, catboost takes 18min for each CV folder. if I use cpu, it would be hours, I guess.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 545979,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "06/06/2019 05:26:58",
      "content": "<p>My only problem with the computer was RAM.  I have 32 GB RAM and 30GB swap disc space, and I have 1080ti.  My computer run out of memory a few times.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543605": "This is actually my first real Kaggle competition. Although I started very late (2 weeks ago), I got pretty good results in the end (public 79th, private 40th). I cannot make it wothout the Kaggle community sharing lots of codes and ideas. Here is what I learned from this competition:\n- **feature is much more important than model** Plenty of people used 4194 non-overlapping samples as training data. I used those as well at the beginning. But I found that was not a good choice because cropping the training data in non-overlapping chunk is not enough. No matter how I tune the models, the ranking does not imporve much. So I looked at other kernels and found https://www.kaggle.com/tocha4/lanl-master-s-approach. I used their features as training data and it improves A LOT.\n- **A good computer is necessary and start early** the features i extracted is around 30000 by 1000. It usually takes overnight to train the models and grid search the best hyper-parameters. I also moved the computation to GPU if it has GPU support. I have a very good computer (i9-9900k/64g/2080ti) so that I can do lots of computational expensive training. But due to the complex model I used, I only trained my final model once. I guess i could get better score if I have more time.\n\nThe model I used is:\nrandom forest/knn/extra tree/ada boost/Nu SVR/light gbm/xgboost/cat boost - 8 models\nlight gbm/xgboost/cat boost - to do stacking on above 8 models\nlight gbm stacking on above 3 models\n\nIt looks like there are some testing sets leakage. I should look at them if I started early.",
    "543622": "Interesting that you used a 3 level stack.  Have you checked that it really improves over your second level models?",
    "543627": "Always having a good computer is an advantage. Specially with multiple GPUs for image classification challenges. But only having RAM are necessary in this case.",
    "543642": "I tried two submission, one is 3 level stacking, the other one is 2 level stacking with smallest CV error(catboost). THe first one is better: public 1.34808 vs 1.35378; private 2.43624 vs 2.49996",
    "543654": "training xgboost and catboost with GPU is much faster than CPU.",
    "543840": "Thanks for letting me know, I will try Catboost on GPU. \n\nBut in this competition the dataset is pretty small and even on CPU train very fast.",
    "543909": "it also depends on how many features and samples used. I used around 30000x1000 training set. With 2080ti, catboost takes 18min for each CV folder. if I use cpu, it would be hours, I guess.",
    "545979": "My only problem with the computer was RAM.  I have 32 GB RAM and 30GB swap disc space, and I have 1080ti.  My computer run out of memory a few times."
  },
  "source": "meta"
}