{
  "id": 54184,
  "title": "questions about feature selection",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54184",
  "author_name": "",
  "post_date": "2018-04-10T19:32:20.285520500Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>How you guys do the feature selection? I have been confused by this question for a long long long time. I know CV score may be a good choice, but do you do the CV for every single new feautre added in the model? The time of execution will be decades no? Especailly in this kind of competition with millions of rows. By the way, how do you do the CV in general? I am a little afraid of CV in huge dataset, cause it seems it will nevery end.</p>\n\n<p>I have seen some kernels use the distribution of data within different class in function of the feature to see if a feature is significant or not, I have tried, but I saw nearly no difference.....</p>\n\n<p>Or the feature importance in the tree model? Can we really exclude the feautre with little importance? Or a Z-score?</p>",
  "messages": [
    {
      "id": "311855",
      "postDate": "04/10/2018 19:32:20",
      "content": "<p>How you guys do the feature selection? I have been confused by this question for a long long long time. I know CV score may be a good choice, but do you do the CV for every single new feautre added in the model? The time of execution will be decades no? Especailly in this kind of competition with millions of rows. By the way, how do you do the CV in general? I am a little afraid of CV in huge dataset, cause it seems it will nevery end.</p>\n\n<p>I have seen some kernels use the distribution of data within different class in function of the feature to see if a feature is significant or not, I have tried, but I saw nearly no difference.....</p>\n\n<p>Or the feature importance in the tree model? Can we really exclude the feautre with little importance? Or a Z-score?</p>",
      "rawMarkdown": "How you guys do the feature selection? I have been confused by this question for a long long long time. I know CV score may be a good choice, but do you do the CV for every single new feautre added in the model? The time of execution will be decades no? Especailly in this kind of competition with millions of rows. By the way, how do you do the CV in general? I am a little afraid of CV in huge dataset, cause it seems it will nevery end.\n\nI have seen some kernels use the distribution of data within different class in function of the feature to see if a feature is significant or not, I have tried, but I saw nearly no difference.....\n\nOr the feature importance in the tree model? Can we really exclude the feautre with little importance? Or a Z-score?",
      "votes": null
    },
    {
      "id": "311857",
      "postDate": "04/10/2018 19:43:44",
      "content": "<p>Load all discussions from the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion\"><strong>forum</strong></a> - just keep scrolling down until there aren't any more left - and search for word \"feature\" in that page. As a general rule, if you come up with a question late in the competition, there is a fair chance that it has already been asked in answered.</p>\n\n<p>There are at least a dozen topics that deal with features. I will list only few of them so you have something to do as well:</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\"><strong>link 1</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53246\"><strong>link 2</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53820\"><strong>link 3</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53752\"><strong>link 4</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53397\"><strong>link 5</strong></a></p>",
      "rawMarkdown": "Load all discussions from the [__forum__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion) - just keep scrolling down until there aren't any more left - and search for word \"feature\" in that page. As a general rule, if you come up with a question late in the competition, there is a fair chance that it has already been asked in answered.\n\nThere are at least a dozen topics that deal with features. I will list only few of them so you have something to do as well:\n\n[__link 1__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634)\n\n[__link 2__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53246)\n\n[__link 3__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53820)\n\n[__link 4__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53752)\n\n[__link 5__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53397)",
      "votes": null
    },
    {
      "id": "311951",
      "postDate": "04/11/2018 00:59:19",
      "content": "<blockquote>\n  <p>but do you do the CV for every single new feautre added in the model?</p>\n</blockquote>\n\n<p>You better =). </p>\n\n<blockquote>\n  <p>The time of execution will be decades no?</p>\n</blockquote>\n\n<p>Yup. Luckily for gbm, you can purchase a 24-core blade server w/ 64gb ram off of craigslist / ebay for $200-$400. I Imagine what our cousins over in DSB2018 have to go through........</p>\n\n<p>Also you can use a subset of the data. But picking the right subset = critical . . .</p>",
      "rawMarkdown": "&gt; but do you do the CV for every single new feautre added in the model?\n\nYou better =). \n\n&gt; The time of execution will be decades no?\n\nYup. Luckily for gbm, you can purchase a 24-core blade server w/ 64gb ram off of craigslist / ebay for $200-$400. I Imagine what our cousins over in DSB2018 have to go through........\n\nAlso you can use a subset of the data. But picking the right subset = critical . . .",
      "votes": null
    },
    {
      "id": "312134",
      "postDate": "04/11/2018 08:52:02",
      "content": "<p>Thx, I have read some of them, I guess I just don't want to admit that a good feature selection is from experience....</p>",
      "rawMarkdown": "Thx, I have read some of them, I guess I just don't want to admit that a good feature selection is from experience....",
      "votes": null
    },
    {
      "id": "312442",
      "postDate": "04/11/2018 18:59:57",
      "content": "<p>Is it worth buying something like this more than upgrading a recent setup ? how do old xeon with ddr3 ram compare to, let's say for example AMD ryzen 1700 + ddr4 ram ? I know the price is not the same but I have no idea about the perf</p>\n\n<p>ps:  any link to that kind of machine would be appreciated :)</p>",
      "rawMarkdown": "Is it worth buying something like this more than upgrading a recent setup ? how do old xeon with ddr3 ram compare to, let's say for example AMD ryzen 1700 + ddr4 ram ? I know the price is not the same but I have no idea about the perf\n\nps:  any link to that kind of machine would be appreciated :)",
      "votes": null
    },
    {
      "id": "312500",
      "postDate": "04/11/2018 21:28:36",
      "content": "<p>how about a hadoop cluster, that will be much better</p>",
      "rawMarkdown": "how about a hadoop cluster, that will be much better",
      "votes": null
    },
    {
      "id": "312599",
      "postDate": "04/12/2018 02:56:16",
      "content": "<p>I wouldn't recommend putting  cash down on anything unless you're \"In the Gold\" when it comes to rankings. For sample servers, check out the Poweredge series on ebay / craigslist / local pawnshops. I take my own advice so I can't speak to i7 vs Xeon w/ DDR3; but if I move up a few more spots I'll give you a full report :P</p>",
      "rawMarkdown": "I wouldn't recommend putting  cash down on anything unless you're \"In the Gold\" when it comes to rankings. For sample servers, check out the Poweredge series on ebay / craigslist / local pawnshops. I take my own advice so I can't speak to i7 vs Xeon w/ DDR3; but if I move up a few more spots I'll give you a full report :P",
      "votes": null
    },
    {
      "id": "314274",
      "postDate": "04/15/2018 04:47:34",
      "content": "<p>linustechtips had a funny video on this: <a href=\"https://www.youtube.com/watch?v=XAuOmJVZeQI\">https://www.youtube.com/watch?v=XAuOmJVZeQI</a></p>\n\n<p>Basically it's just better to upgrade to one of the new multi-core CPUs, threadripper/i9 if you have the budget. These will be cheaper, faster, and use half the power than the xeon processors from a few years ago.</p>\n\n<p>AWS spot instances are also a good way to get extra computing if you need some. For example, the new c5.9xlarge instances with 36cores&amp;72GB RAM cost $7.2 a day to run, so you can run them for some days for each competition. Those also have the new AVX 512 instructions, which make vectorized parts of code run 8-times faster, including any model code running on top of CUDA&amp;deep learning libraries.</p>",
      "rawMarkdown": "linustechtips had a funny video on this: https://www.youtube.com/watch?v=XAuOmJVZeQI\n\nBasically it's just better to upgrade to one of the new multi-core CPUs, threadripper/i9 if you have the budget. These will be cheaper, faster, and use half the power than the xeon processors from a few years ago.\n\nAWS spot instances are also a good way to get extra computing if you need some. For example, the new c5.9xlarge instances with 36cores&amp;72GB RAM cost $7.2 a day to run, so you can run them for some days for each competition. Those also have the new AVX 512 instructions, which make vectorized parts of code run 8-times faster, including any model code running on top of CUDA&amp;deep learning libraries.",
      "votes": null
    },
    {
      "id": "314276",
      "postDate": "04/15/2018 05:03:50",
      "content": "<p>You didn't seem to get a solid answer on this one.</p>\n\n<p>This competition is prediction on new time-organized data, so it's essentially an online learning problem. This means that cross-validation is the wrong approach to do validation, including feature selection. What's used most often in the online learning setting is held-out validation on the last time period before the test data, in this case this could be the last day before the test day. A common variant of this is to train and test in a stream fashion, where you first test on each instance, and then update your classifier. This gives estimates across the entire dataset, without wasting any of the data for a separate validation set, or violating the held-out testing principle. You can also do this in a mini-batch fashion, and use an online validation measure that weights the latest data more, as is done in my script (<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752\">https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752</a>). Cross-validation can still get you started, but you should move to held-out validation if you want a properly optimized model.</p>\n\n<p>Once you have a validation setup, you can use feature importances from classifiers to do selection. For example the ELI5 library (<a href=\"https://github.com/TeamHG-Memex/eli5\">https://github.com/TeamHG-Memex/eli5</a>) gives feature importances and visualizes them. You might want to look at features that are stable across different time samples in the training data. You can use multiple measures (logloss and AUC), and then finally verify on the public leaderboard. Private vs. public leaderboard differences are larger in time-organized competitions, so you have to trust equally your local measures and intuition.</p>",
      "rawMarkdown": "You didn't seem to get a solid answer on this one.\n\nThis competition is prediction on new time-organized data, so it's essentially an online learning problem. This means that cross-validation is the wrong approach to do validation, including feature selection. What's used most often in the online learning setting is held-out validation on the last time period before the test data, in this case this could be the last day before the test day. A common variant of this is to train and test in a stream fashion, where you first test on each instance, and then update your classifier. This gives estimates across the entire dataset, without wasting any of the data for a separate validation set, or violating the held-out testing principle. You can also do this in a mini-batch fashion, and use an online validation measure that weights the latest data more, as is done in my script (https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752). Cross-validation can still get you started, but you should move to held-out validation if you want a properly optimized model.\n\nOnce you have a validation setup, you can use feature importances from classifiers to do selection. For example the ELI5 library (https://github.com/TeamHG-Memex/eli5) gives feature importances and visualizes them. You might want to look at features that are stable across different time samples in the training data. You can use multiple measures (logloss and AUC), and then finally verify on the public leaderboard. Private vs. public leaderboard differences are larger in time-organized competitions, so you have to trust equally your local measures and intuition.",
      "votes": null
    },
    {
      "id": "314411",
      "postDate": "04/15/2018 14:20:38",
      "content": "<p>Quite funny yes :) thank you for the good answer!</p>",
      "rawMarkdown": "Quite funny yes :) thank you for the good answer!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 311857,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/10/2018 19:43:44",
      "content": "<p>Load all discussions from the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion\"><strong>forum</strong></a> - just keep scrolling down until there aren't any more left - and search for word \"feature\" in that page. As a general rule, if you come up with a question late in the competition, there is a fair chance that it has already been asked in answered.</p>\n\n<p>There are at least a dozen topics that deal with features. I will list only few of them so you have something to do as well:</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\"><strong>link 1</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53246\"><strong>link 2</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53820\"><strong>link 3</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53752\"><strong>link 4</strong></a></p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53397\"><strong>link 5</strong></a></p>",
      "votes": null,
      "replies": [
        {
          "id": 312134,
          "author_name": "marrvolo",
          "author_url": "",
          "post_date": "04/11/2018 08:52:02",
          "content": "<p>Thx, I have read some of them, I guess I just don't want to admit that a good feature selection is from experience....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 311951,
      "author_name": "authman",
      "author_url": "",
      "post_date": "04/11/2018 00:59:19",
      "content": "<blockquote>\n  <p>but do you do the CV for every single new feautre added in the model?</p>\n</blockquote>\n\n<p>You better =). </p>\n\n<blockquote>\n  <p>The time of execution will be decades no?</p>\n</blockquote>\n\n<p>Yup. Luckily for gbm, you can purchase a 24-core blade server w/ 64gb ram off of craigslist / ebay for $200-$400. I Imagine what our cousins over in DSB2018 have to go through........</p>\n\n<p>Also you can use a subset of the data. But picking the right subset = critical . . .</p>",
      "votes": null,
      "replies": [
        {
          "id": 312442,
          "author_name": "fl2ooo",
          "author_url": "",
          "post_date": "04/11/2018 18:59:57",
          "content": "<p>Is it worth buying something like this more than upgrading a recent setup ? how do old xeon with ddr3 ram compare to, let's say for example AMD ryzen 1700 + ddr4 ram ? I know the price is not the same but I have no idea about the perf</p>\n\n<p>ps:  any link to that kind of machine would be appreciated :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312500,
          "author_name": "marrvolo",
          "author_url": "",
          "post_date": "04/11/2018 21:28:36",
          "content": "<p>how about a hadoop cluster, that will be much better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312599,
          "author_name": "authman",
          "author_url": "",
          "post_date": "04/12/2018 02:56:16",
          "content": "<p>I wouldn't recommend putting  cash down on anything unless you're \"In the Gold\" when it comes to rankings. For sample servers, check out the Poweredge series on ebay / craigslist / local pawnshops. I take my own advice so I can't speak to i7 vs Xeon w/ DDR3; but if I move up a few more spots I'll give you a full report :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314274,
          "author_name": "anttip",
          "author_url": "",
          "post_date": "04/15/2018 04:47:34",
          "content": "<p>linustechtips had a funny video on this: <a href=\"https://www.youtube.com/watch?v=XAuOmJVZeQI\">https://www.youtube.com/watch?v=XAuOmJVZeQI</a></p>\n\n<p>Basically it's just better to upgrade to one of the new multi-core CPUs, threadripper/i9 if you have the budget. These will be cheaper, faster, and use half the power than the xeon processors from a few years ago.</p>\n\n<p>AWS spot instances are also a good way to get extra computing if you need some. For example, the new c5.9xlarge instances with 36cores&amp;72GB RAM cost $7.2 a day to run, so you can run them for some days for each competition. Those also have the new AVX 512 instructions, which make vectorized parts of code run 8-times faster, including any model code running on top of CUDA&amp;deep learning libraries.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314411,
          "author_name": "fl2ooo",
          "author_url": "",
          "post_date": "04/15/2018 14:20:38",
          "content": "<p>Quite funny yes :) thank you for the good answer!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 314276,
      "author_name": "anttip",
      "author_url": "",
      "post_date": "04/15/2018 05:03:50",
      "content": "<p>You didn't seem to get a solid answer on this one.</p>\n\n<p>This competition is prediction on new time-organized data, so it's essentially an online learning problem. This means that cross-validation is the wrong approach to do validation, including feature selection. What's used most often in the online learning setting is held-out validation on the last time period before the test data, in this case this could be the last day before the test day. A common variant of this is to train and test in a stream fashion, where you first test on each instance, and then update your classifier. This gives estimates across the entire dataset, without wasting any of the data for a separate validation set, or violating the held-out testing principle. You can also do this in a mini-batch fashion, and use an online validation measure that weights the latest data more, as is done in my script (<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752\">https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752</a>). Cross-validation can still get you started, but you should move to held-out validation if you want a properly optimized model.</p>\n\n<p>Once you have a validation setup, you can use feature importances from classifiers to do selection. For example the ELI5 library (<a href=\"https://github.com/TeamHG-Memex/eli5\">https://github.com/TeamHG-Memex/eli5</a>) gives feature importances and visualizes them. You might want to look at features that are stable across different time samples in the training data. You can use multiple measures (logloss and AUC), and then finally verify on the public leaderboard. Private vs. public leaderboard differences are larger in time-organized competitions, so you have to trust equally your local measures and intuition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "311855": "How you guys do the feature selection? I have been confused by this question for a long long long time. I know CV score may be a good choice, but do you do the CV for every single new feautre added in the model? The time of execution will be decades no? Especailly in this kind of competition with millions of rows. By the way, how do you do the CV in general? I am a little afraid of CV in huge dataset, cause it seems it will nevery end.\n\nI have seen some kernels use the distribution of data within different class in function of the feature to see if a feature is significant or not, I have tried, but I saw nearly no difference.....\n\nOr the feature importance in the tree model? Can we really exclude the feautre with little importance? Or a Z-score?",
    "311857": "Load all discussions from the [__forum__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion) - just keep scrolling down until there aren't any more left - and search for word \"feature\" in that page. As a general rule, if you come up with a question late in the competition, there is a fair chance that it has already been asked in answered.\n\nThere are at least a dozen topics that deal with features. I will list only few of them so you have something to do as well:\n\n[__link 1__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634)\n\n[__link 2__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53246)\n\n[__link 3__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53820)\n\n[__link 4__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53752)\n\n[__link 5__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53397)",
    "311951": "&gt; but do you do the CV for every single new feautre added in the model?\n\nYou better =). \n\n&gt; The time of execution will be decades no?\n\nYup. Luckily for gbm, you can purchase a 24-core blade server w/ 64gb ram off of craigslist / ebay for $200-$400. I Imagine what our cousins over in DSB2018 have to go through........\n\nAlso you can use a subset of the data. But picking the right subset = critical . . .",
    "312134": "Thx, I have read some of them, I guess I just don't want to admit that a good feature selection is from experience....",
    "312442": "Is it worth buying something like this more than upgrading a recent setup ? how do old xeon with ddr3 ram compare to, let's say for example AMD ryzen 1700 + ddr4 ram ? I know the price is not the same but I have no idea about the perf\n\nps:  any link to that kind of machine would be appreciated :)",
    "312500": "how about a hadoop cluster, that will be much better",
    "312599": "I wouldn't recommend putting  cash down on anything unless you're \"In the Gold\" when it comes to rankings. For sample servers, check out the Poweredge series on ebay / craigslist / local pawnshops. I take my own advice so I can't speak to i7 vs Xeon w/ DDR3; but if I move up a few more spots I'll give you a full report :P",
    "314274": "linustechtips had a funny video on this: https://www.youtube.com/watch?v=XAuOmJVZeQI\n\nBasically it's just better to upgrade to one of the new multi-core CPUs, threadripper/i9 if you have the budget. These will be cheaper, faster, and use half the power than the xeon processors from a few years ago.\n\nAWS spot instances are also a good way to get extra computing if you need some. For example, the new c5.9xlarge instances with 36cores&amp;72GB RAM cost $7.2 a day to run, so you can run them for some days for each competition. Those also have the new AVX 512 instructions, which make vectorized parts of code run 8-times faster, including any model code running on top of CUDA&amp;deep learning libraries.",
    "314276": "You didn't seem to get a solid answer on this one.\n\nThis competition is prediction on new time-organized data, so it's essentially an online learning problem. This means that cross-validation is the wrong approach to do validation, including feature selection. What's used most often in the online learning setting is held-out validation on the last time period before the test data, in this case this could be the last day before the test day. A common variant of this is to train and test in a stream fashion, where you first test on each instance, and then update your classifier. This gives estimates across the entire dataset, without wasting any of the data for a separate validation set, or violating the held-out testing principle. You can also do this in a mini-batch fashion, and use an online validation measure that weights the latest data more, as is done in my script (https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9752). Cross-validation can still get you started, but you should move to held-out validation if you want a properly optimized model.\n\nOnce you have a validation setup, you can use feature importances from classifiers to do selection. For example the ELI5 library (https://github.com/TeamHG-Memex/eli5) gives feature importances and visualizes them. You might want to look at features that are stable across different time samples in the training data. You can use multiple measures (logloss and AUC), and then finally verify on the public leaderboard. Private vs. public leaderboard differences are larger in time-organized competitions, so you have to trust equally your local measures and intuition.",
    "314411": "Quite funny yes :) thank you for the good answer!"
  },
  "source": "meta"
}