{
  "id": 27647,
  "title": "feature selection",
  "url": "/competitions/outbrain-click-prediction/discussion/27647",
  "author_name": "",
  "post_date": "2017-01-12T15:50:41.103Z",
  "votes": null,
  "comment_count": 4,
  "views": 390,
  "content": "<p>Hello all, \nI wonder how do you do feature selection in this competition,   with this huge amount of data.  And what's the process standard to determine which feature is useful for the prediction? <br>\ncheck the correlation of x and y before using an algo, or use importance of feature like xgboost to do post remove.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "155713",
      "postDate": "01/12/2017 15:50:41",
      "content": "<p>Hello all, \nI wonder how do you do feature selection in this competition,   with this huge amount of data.  And what's the process standard to determine which feature is useful for the prediction? <br>\ncheck the correlation of x and y before using an algo, or use importance of feature like xgboost to do post remove.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hello all, \r\nI wonder how do you do feature selection in this competition,   with this huge amount of data.  And what's the process standard to determine which feature is useful for the prediction?   \r\ncheck the correlation of x and y before using an algo, or use importance of feature like xgboost to do post remove.\r\n\r\nThanks!",
      "votes": null
    },
    {
      "id": "155758",
      "postDate": "01/12/2017 18:51:41",
      "content": "<p>Your options for feature engineering on this one might be based on your memory. From what i've seen...</p>\n\n<p>Low Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo</p>\n\n<p>High Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that</p>",
      "rawMarkdown": "Your options for feature engineering on this one might be based on your memory. From what i've seen...\r\n\r\nLow Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo\r\n\r\nHigh Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that",
      "votes": null
    },
    {
      "id": "155979",
      "postDate": "01/13/2017 15:55:18",
      "content": "<p>[quote=Vape Naysh;155758]</p>\n\n<p>Your options for feature engineering on this one might be based on your memory. From what i've seen...</p>\n\n<p>Low Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo</p>\n\n<p>High Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that</p>\n\n<p>[/quote]</p>\n\n<p>so how can you know that the 5-30 features you chose is useful to take in account to do prediction? by intuition or there's some mechanism behind to help you find it. Thanks!</p>",
      "rawMarkdown": "[quote=Vape Naysh;155758]\r\n\r\nYour options for feature engineering on this one might be based on your memory. From what i've seen...\r\n\r\nLow Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo\r\n\r\nHigh Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that\r\n\r\n\r\n\r\n[/quote]\r\n\r\nso how can you know that the 5-30 features you chose is useful to take in account to do prediction? by intuition or there's some mechanism behind to help you find it. Thanks!",
      "votes": null
    },
    {
      "id": "156137",
      "postDate": "01/14/2017 13:31:49",
      "content": "<p>For me, I use part of the dataset to select the useful features. For example, using only 2000000 rows data in event.csv, which in total contains more than ten times of that number, it can be very fast to test whether one feature is useful or not with the MAP@12 evaluation metric. In my case, MSI GE62 6QF notebook, which has 16GB RAM and i7 cpu, it takes me only about 1 minute to run 5 iterations of the FTRL pipeline to decide if the feature contributes to the result.</p>",
      "rawMarkdown": "For me, I use part of the dataset to select the useful features. For example, using only 2000000 rows data in event.csv, which in total contains more than ten times of that number, it can be very fast to test whether one feature is useful or not with the MAP@12 evaluation metric. In my case, MSI GE62 6QF notebook, which has 16GB RAM and i7 cpu, it takes me only about 1 minute to run 5 iterations of the FTRL pipeline to decide if the feature contributes to the result.",
      "votes": null
    },
    {
      "id": "156181",
      "postDate": "01/14/2017 18:32:41",
      "content": "<p>@noooooo, yea pretty much what SBoost said, like any competition on here, you wont know until you try.  Just try to do things in a computational /memory efficient way and you'll be able to move through ideas more quickly. You can get a pretty good score by reading the forums and kernals on here and taking ideas.</p>",
      "rawMarkdown": "noooooo, yea pretty much what SBoost said, like any competition on here, you wont know until you try.  Just try to do things in a computational /memory efficient way and you'll be able to move through ideas more quickly. You can get a pretty good score by reading the forums and kernals on here and taking ideas.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 155758,
      "author_name": "vapenaysh",
      "author_url": "",
      "post_date": "01/12/2017 18:51:41",
      "content": "<p>Your options for feature engineering on this one might be based on your memory. From what i've seen...</p>\n\n<p>Low Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo</p>\n\n<p>High Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that</p>",
      "votes": null,
      "replies": [
        {
          "id": 155979,
          "author_name": "betterplace",
          "author_url": "",
          "post_date": "01/13/2017 15:55:18",
          "content": "<p>[quote=Vape Naysh;155758]</p>\n\n<p>Your options for feature engineering on this one might be based on your memory. From what i've seen...</p>\n\n<p>Low Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo</p>\n\n<p>High Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that</p>\n\n<p>[/quote]</p>\n\n<p>so how can you know that the 5-30 features you chose is useful to take in account to do prediction? by intuition or there's some mechanism behind to help you find it. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 156137,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "01/14/2017 13:31:49",
      "content": "<p>For me, I use part of the dataset to select the useful features. For example, using only 2000000 rows data in event.csv, which in total contains more than ten times of that number, it can be very fast to test whether one feature is useful or not with the MAP@12 evaluation metric. In my case, MSI GE62 6QF notebook, which has 16GB RAM and i7 cpu, it takes me only about 1 minute to run 5 iterations of the FTRL pipeline to decide if the feature contributes to the result.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156181,
      "author_name": "vapenaysh",
      "author_url": "",
      "post_date": "01/14/2017 18:32:41",
      "content": "<p>@noooooo, yea pretty much what SBoost said, like any competition on here, you wont know until you try.  Just try to do things in a computational /memory efficient way and you'll be able to move through ideas more quickly. You can get a pretty good score by reading the forums and kernals on here and taking ideas.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "155713": "Hello all, \r\nI wonder how do you do feature selection in this competition,   with this huge amount of data.  And what's the process standard to determine which feature is useful for the prediction?   \r\ncheck the correlation of x and y before using an algo, or use importance of feature like xgboost to do post remove.\r\n\r\nThanks!",
    "155758": "Your options for feature engineering on this one might be based on your memory. From what i've seen...\r\n\r\nLow Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo\r\n\r\nHigh Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that",
    "155979": "[quote=Vape Naysh;155758]\r\n\r\nYour options for feature engineering on this one might be based on your memory. From what i've seen...\r\n\r\nLow Memory: one-hot-encoding the categorical features through feature hashing w/ regularized algo\r\n\r\nHigh Memory: 5-30 engineered features, presumably through adjusted CTR's or something like that\r\n\r\n\r\n\r\n[/quote]\r\n\r\nso how can you know that the 5-30 features you chose is useful to take in account to do prediction? by intuition or there's some mechanism behind to help you find it. Thanks!",
    "156137": "For me, I use part of the dataset to select the useful features. For example, using only 2000000 rows data in event.csv, which in total contains more than ten times of that number, it can be very fast to test whether one feature is useful or not with the MAP@12 evaluation metric. In my case, MSI GE62 6QF notebook, which has 16GB RAM and i7 cpu, it takes me only about 1 minute to run 5 iterations of the FTRL pipeline to decide if the feature contributes to the result.",
    "156181": "noooooo, yea pretty much what SBoost said, like any competition on here, you wont know until you try.  Just try to do things in a computational /memory efficient way and you'll be able to move through ideas more quickly. You can get a pretty good score by reading the forums and kernals on here and taking ideas."
  },
  "source": "meta"
}