{
  "id": 94321,
  "title": "My one trick for feature selection/data augmentation (183rd place)",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/da-li-my-one-trick-for-feature-selection-data-augm",
  "author_name": "",
  "post_date": "2019-06-04T01:09:35.356079200Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>With such a small amount of data (4194 segments), it is way too easy to overfit the cross validation and public LB. What we did to try to compensate for the small data size was to keep the segment length at 150,000, but train the models on 140,000 data points instead. For each segment a starting offset of 0, 2000, 4000, 6000, 8000, and 10000 was used to produce 6-times the training data to feed into the models, which also removes any spurious features that might happen to have some importance by chance. On the test data the 6 outputs are then averaged. </p>\n\n<p>This setup only marginally improves CV score by about 0.003, but it makes the feature importance for a whole lot of features zero or close to zero. Our best submission was actually a single lgbm model trained with this scheme, although we would have never picked that over other stacked or averaged results.</p>",
  "messages": [
    {
      "id": "542541",
      "postDate": "06/04/2019 01:09:35",
      "content": "<p>With such a small amount of data (4194 segments), it is way too easy to overfit the cross validation and public LB. What we did to try to compensate for the small data size was to keep the segment length at 150,000, but train the models on 140,000 data points instead. For each segment a starting offset of 0, 2000, 4000, 6000, 8000, and 10000 was used to produce 6-times the training data to feed into the models, which also removes any spurious features that might happen to have some importance by chance. On the test data the 6 outputs are then averaged. </p>\n\n<p>This setup only marginally improves CV score by about 0.003, but it makes the feature importance for a whole lot of features zero or close to zero. Our best submission was actually a single lgbm model trained with this scheme, although we would have never picked that over other stacked or averaged results.</p>",
      "rawMarkdown": "With such a small amount of data (4194 segments), it is way too easy to overfit the cross validation and public LB. What we did to try to compensate for the small data size was to keep the segment length at 150,000, but train the models on 140,000 data points instead. For each segment a starting offset of 0, 2000, 4000, 6000, 8000, and 10000 was used to produce 6-times the training data to feed into the models, which also removes any spurious features that might happen to have some importance by chance. On the test data the 6 outputs are then averaged. \n\nThis setup only marginally improves CV score by about 0.003, but it makes the feature importance for a whole lot of features zero or close to zero. Our best submission was actually a single lgbm model trained with this scheme, although we would have never picked that over other stacked or averaged results.",
      "votes": null
    },
    {
      "id": "542601",
      "postDate": "06/04/2019 01:58:51",
      "content": "<p>The train is not 4194 segments, actually train is composed by only 16 quakes. Its much worse. Imagine training a model with a dataset of 16 rows. </p>",
      "rawMarkdown": "The train is not 4194 segments, actually train is composed by only 16 quakes. Its much worse. Imagine training a model with a dataset of 16 rows.",
      "votes": null
    },
    {
      "id": "542902",
      "postDate": "06/04/2019 07:12:52",
      "content": "<p>It is indeed a nice trick, in the first half of the competition I also tried something similar with a RNN, before realizing RNN was not going to work and starting from scratch. </p>\n\n<p>But also yes, as @Giba explains, the big(er) problem here was extremely small data what limits efective data augmentation. </p>",
      "rawMarkdown": "It is indeed a nice trick, in the first half of the competition I also tried something similar with a RNN, before realizing RNN was not going to work and starting from scratch. \n\nBut also yes, as @Giba explains, the big(er) problem here was extremely small data what limits efective data augmentation.",
      "votes": null
    },
    {
      "id": "542906",
      "postDate": "06/04/2019 07:15:02",
      "content": "<p>We don't need to imagine ;) I just did!</p>",
      "rawMarkdown": "We don't need to imagine ;) I just did!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542601,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 01:58:51",
      "content": "<p>The train is not 4194 segments, actually train is composed by only 16 quakes. Its much worse. Imagine training a model with a dataset of 16 rows. </p>",
      "votes": null,
      "replies": [
        {
          "id": 542906,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "06/04/2019 07:15:02",
          "content": "<p>We don't need to imagine ;) I just did!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542902,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "06/04/2019 07:12:52",
      "content": "<p>It is indeed a nice trick, in the first half of the competition I also tried something similar with a RNN, before realizing RNN was not going to work and starting from scratch. </p>\n\n<p>But also yes, as @Giba explains, the big(er) problem here was extremely small data what limits efective data augmentation. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542541": "With such a small amount of data (4194 segments), it is way too easy to overfit the cross validation and public LB. What we did to try to compensate for the small data size was to keep the segment length at 150,000, but train the models on 140,000 data points instead. For each segment a starting offset of 0, 2000, 4000, 6000, 8000, and 10000 was used to produce 6-times the training data to feed into the models, which also removes any spurious features that might happen to have some importance by chance. On the test data the 6 outputs are then averaged. \n\nThis setup only marginally improves CV score by about 0.003, but it makes the feature importance for a whole lot of features zero or close to zero. Our best submission was actually a single lgbm model trained with this scheme, although we would have never picked that over other stacked or averaged results.",
    "542601": "The train is not 4194 segments, actually train is composed by only 16 quakes. Its much worse. Imagine training a model with a dataset of 16 rows.",
    "542902": "It is indeed a nice trick, in the first half of the competition I also tried something similar with a RNN, before realizing RNN was not going to work and starting from scratch. \n\nBut also yes, as @Giba explains, the big(er) problem here was extremely small data what limits efective data augmentation.",
    "542906": "We don't need to imagine ;) I just did!"
  },
  "source": "meta"
}