{
  "id": 77546,
  "title": "How many features are you currently using?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77546",
  "author_name": "Abhishek Thakur",
  "post_date": "2019-01-14T07:04:41.210000",
  "votes": 11,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Since this is a feature building competition, it would be a good idea to know how many features is everyone using :)</p>\n\n<p>Approx. 60 features / LB 1.505</p>",
  "messages": [
    {
      "id": 455565,
      "postDate": "2019-01-14T07:04:41.210Z",
      "content": "<p>Since this is a feature building competition, it would be a good idea to know how many features is everyone using :)</p>\n\n<p>Approx. 60 features / LB 1.505</p>",
      "rawMarkdown": "Since this is a feature building competition, it would be a good idea to know how many features is everyone using :)\n\n\nApprox. 60 features / LB 1.505\n\n",
      "votes": 10
    },
    {
      "id": 456160,
      "postDate": "2019-01-15T08:52:10.460Z",
      "content": "<p>currently, I am using just few aggregated features, as following:\n- mean, std, var, min, max, abs_max / segment;\n- A0-A10: being the FFT module values for CC and first 10 harmonics; R1-R5: being the real components of 1st 5 harmonics;\nTotally: 21 features / LB: 1.884 (very poor).</p>\n\n<p>I am also calculated 10 times more segments than  the 4194 resulting by simply dividing the train data in 150k segments (I simply shift segments with a shift = 1/10 * segment, a technique for data multiplication we used to use in the past: we call it 'shifting aperture').</p>\n\n<p>It looks that my validation error is quite close to the LB score, but the train error is much smaller so I do have a major problem with my model, a lot to improve to it.</p>",
      "rawMarkdown": "currently, I am using just few aggregated features, as following:\n- mean, std, var, min, max, abs_max / segment;\n- A0-A10: being the FFT module values for CC and first 10 harmonics; R1-R5: being the real components of 1st 5 harmonics;\nTotally: 21 features / LB: 1.884 (very poor).\n\nI am also calculated 10 times more segments than  the 4194 resulting by simply dividing the train data in 150k segments (I simply shift segments with a shift = 1/10 * segment, a technique for data multiplication we used to use in the past: we call it 'shifting aperture').\n\nIt looks that my validation error is quite close to the LB score, but the train error is much smaller so I do have a major problem with my model, a lot to improve to it.",
      "votes": 5,
      "replies": [
        {
          "id": 456329,
          "postDate": "2019-01-15T15:26:47.050Z",
          "content": "<p>Seems like you may not be treating you train / validation sets in the same way, or your validation set does not resemble the test set.</p>",
          "rawMarkdown": "Seems like you may not be treating you train / validation sets in the same way, or your validation set does not resemble the test set."
        },
        {
          "id": 458202,
          "postDate": "2019-01-19T03:59:42.723Z",
          "content": "<p>Hi! Gabriel, I been using shifting aperture, but the more shifts I add my LB Score gets worst, I'm trying to find the sweet spot to calibrate my CV with LB based on the number of shifts</p>",
          "rawMarkdown": "Hi! Gabriel, I been using shifting aperture, but the more shifts I add my LB Score gets worst, I'm trying to find the sweet spot to calibrate my CV with LB based on the number of shifts"
        },
        {
          "id": 460285,
          "postDate": "2019-01-23T11:11:54.713Z",
          "content": "<p>I used welch using a point shift (np.arange(100,1000,100))and got 1.62 using Genetic Programming and 1.67 with lgb - might be useful for an ensemble</p>",
          "rawMarkdown": "I used welch using a point shift (np.arange(100,1000,100))and got 1.62 using Genetic Programming and 1.67 with lgb - might be useful for an ensemble"
        }
      ]
    },
    {
      "id": 460784,
      "postDate": "2019-01-24T12:17:15.017Z",
      "content": "<p>Andrew's Script + GP == 133 features for a score of 1.461</p>",
      "rawMarkdown": "Andrew's Script + GP == 133 features for a score of 1.461",
      "votes": 4,
      "replies": [
        {
          "id": 460796,
          "postDate": "2019-01-24T12:46:41.447Z",
          "content": "<p>gaussian process?</p>",
          "rawMarkdown": "gaussian process?",
          "votes": 1
        },
        {
          "id": 460797,
          "postDate": "2019-01-24T12:48:52.273Z",
          "content": "<p>Apologies, GP = Genetic Programming</p>",
          "rawMarkdown": "Apologies, GP = Genetic Programming",
          "votes": 1
        },
        {
          "id": 460800,
          "postDate": "2019-01-24T12:52:17.163Z",
          "content": "<p>saw the previous post later. good stuff!</p>",
          "rawMarkdown": "saw the previous post later. good stuff!",
          "votes": 1
        },
        {
          "id": 460802,
          "postDate": "2019-01-24T13:03:40.190Z",
          "content": "<p>@Abhishek, GP is so 2nd nature to @Scirpus he thinks we all hear Genetic Programming when we see GP instead of gaussian process or general practitioner :-)</p>",
          "rawMarkdown": "@Abhishek, GP is so 2nd nature to @Scirpus he thinks we all hear Genetic Programming when we see GP instead of gaussian process or general practitioner :-)",
          "votes": 2
        },
        {
          "id": 460805,
          "postDate": "2019-01-24T13:07:05.463Z",
          "content": "<p>LOL so true ;)</p>",
          "rawMarkdown": "LOL so true ;)",
          "votes": 2
        },
        {
          "id": 509019,
          "postDate": "2019-04-07T07:56:59.527Z",
          "content": "<p>what is Andy's script - sorry for the newb questoin</p>",
          "rawMarkdown": "what is Andy's script - sorry for the newb questoin"
        },
        {
          "id": 509292,
          "postDate": "2019-04-07T16:26:32.577Z",
          "content": "<p>I expect that they mean the rather popular kernel by Andrew Lukyanenko , see <a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a></p>",
          "rawMarkdown": "I expect that they mean the rather popular kernel by Andrew Lukyanenko , see [https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples](https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples)"
        }
      ]
    },
    {
      "id": 455698,
      "postDate": "2019-01-14T11:42:24.620Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 464988,
      "postDate": "2019-02-02T00:52:50.840Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    }
  ],
  "comments": [
    {
      "id": 456160,
      "author_name": "Gabriel Preda",
      "author_url": "",
      "post_date": "2019-01-15T08:52:10.460000",
      "content": "<p>currently, I am using just few aggregated features, as following:\n- mean, std, var, min, max, abs_max / segment;\n- A0-A10: being the FFT module values for CC and first 10 harmonics; R1-R5: being the real components of 1st 5 harmonics;\nTotally: 21 features / LB: 1.884 (very poor).</p>\n\n<p>I am also calculated 10 times more segments than  the 4194 resulting by simply dividing the train data in 150k segments (I simply shift segments with a shift = 1/10 * segment, a technique for data multiplication we used to use in the past: we call it 'shifting aperture').</p>\n\n<p>It looks that my validation error is quite close to the LB score, but the train error is much smaller so I do have a major problem with my model, a lot to improve to it.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 456329,
          "author_name": "HorizonPicking2k18",
          "author_url": "",
          "post_date": "2019-01-15T15:26:47.050000",
          "content": "<p>Seems like you may not be treating you train / validation sets in the same way, or your validation set does not resemble the test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 458202,
          "author_name": "C4rl05/V",
          "author_url": "",
          "post_date": "2019-01-19T03:59:42.723000",
          "content": "<p>Hi! Gabriel, I been using shifting aperture, but the more shifts I add my LB Score gets worst, I'm trying to find the sweet spot to calibrate my CV with LB based on the number of shifts</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 460285,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-23T11:11:54.713000",
          "content": "<p>I used welch using a point shift (np.arange(100,1000,100))and got 1.62 using Genetic Programming and 1.67 with lgb - might be useful for an ensemble</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 460784,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-01-24T12:17:15.017000",
      "content": "<p>Andrew's Script + GP == 133 features for a score of 1.461</p>",
      "votes": 4,
      "replies": [
        {
          "id": 460796,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-01-24T12:46:41.447000",
          "content": "<p>gaussian process?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 460797,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-24T12:48:52.273000",
          "content": "<p>Apologies, GP = Genetic Programming</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 460800,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-01-24T12:52:17.163000",
          "content": "<p>saw the previous post later. good stuff!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 460802,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-01-24T13:03:40.190000",
          "content": "<p>@Abhishek, GP is so 2nd nature to @Scirpus he thinks we all hear Genetic Programming when we see GP instead of gaussian process or general practitioner :-)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 460805,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-24T13:07:05.463000",
          "content": "<p>LOL so true ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 509019,
          "author_name": "Eyas Taifour",
          "author_url": "",
          "post_date": "2019-04-07T07:56:59.527000",
          "content": "<p>what is Andy's script - sorry for the newb questoin</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 509292,
          "author_name": "Ben Nye",
          "author_url": "",
          "post_date": "2019-04-07T16:26:32.577000",
          "content": "<p>I expect that they mean the rather popular kernel by Andrew Lukyanenko , see <a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 455698,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-14T11:42:24.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 464988,
      "author_name": "QiPANDA",
      "author_url": "",
      "post_date": "2019-02-02T00:52:50.840000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "455565": "Since this is a feature building competition, it would be a good idea to know how many features is everyone using :)\n\n\nApprox. 60 features / LB 1.505\n\n",
    "456160": "currently, I am using just few aggregated features, as following:\n- mean, std, var, min, max, abs_max / segment;\n- A0-A10: being the FFT module values for CC and first 10 harmonics; R1-R5: being the real components of 1st 5 harmonics;\nTotally: 21 features / LB: 1.884 (very poor).\n\nI am also calculated 10 times more segments than  the 4194 resulting by simply dividing the train data in 150k segments (I simply shift segments with a shift = 1/10 * segment, a technique for data multiplication we used to use in the past: we call it 'shifting aperture').\n\nIt looks that my validation error is quite close to the LB score, but the train error is much smaller so I do have a major problem with my model, a lot to improve to it.",
    "460784": "Andrew's Script + GP == 133 features for a score of 1.461",
    "455698": "",
    "464988": "Thanks"
  }
}