{
  "id": 16987,
  "title": "Improvements after feature selection",
  "url": "/competitions/dato-native/discussion/16987",
  "author_name": "",
  "post_date": "2015-10-13T21:54:20.773Z",
  "votes": null,
  "comment_count": 2,
  "views": 918,
  "content": "<p>This is my first participation in a Kaggle competition so I was curious about a couple things (since the competition is now almost over).  For my analysis I used a Random Forest with a large number of features.  My process for improvement was to come up with a new idea for additional features and then add them into the the calculation.  Almost always this improved the score, however, as I added more features, the resulting improvement dropped off substantially.  It got to the point where I started to look for other options.  So then I used stacking to combine multiple Random Forest runs together to form a better prediction.  Again, this helped a little bit but not much.  I guess my question is, do you spend more time thinking about and adding additional features or more time tweaking or changing your prediction method?  How does the timeline change what you work on?  Towards the end, are you doing more things like stacking or changing parameters to see what kind of accuracy you can get?</p>\n\n<p>Thanks in advance.</p>",
  "messages": [
    {
      "id": "96041",
      "postDate": "10/13/2015 21:54:20",
      "content": "<p>This is my first participation in a Kaggle competition so I was curious about a couple things (since the competition is now almost over).  For my analysis I used a Random Forest with a large number of features.  My process for improvement was to come up with a new idea for additional features and then add them into the the calculation.  Almost always this improved the score, however, as I added more features, the resulting improvement dropped off substantially.  It got to the point where I started to look for other options.  So then I used stacking to combine multiple Random Forest runs together to form a better prediction.  Again, this helped a little bit but not much.  I guess my question is, do you spend more time thinking about and adding additional features or more time tweaking or changing your prediction method?  How does the timeline change what you work on?  Towards the end, are you doing more things like stacking or changing parameters to see what kind of accuracy you can get?</p>\n\n<p>Thanks in advance.</p>",
      "rawMarkdown": "This is my first participation in a Kaggle competition so I was curious about a couple things (since the competition is now almost over).  For my analysis I used a Random Forest with a large number of features.  My process for improvement was to come up with a new idea for additional features and then add them into the the calculation.  Almost always this improved the score, however, as I added more features, the resulting improvement dropped off substantially.  It got to the point where I started to look for other options.  So then I used stacking to combine multiple Random Forest runs together to form a better prediction.  Again, this helped a little bit but not much.  I guess my question is, do you spend more time thinking about and adding additional features or more time tweaking or changing your prediction method?  How does the timeline change what you work on?  Towards the end, are you doing more things like stacking or changing parameters to see what kind of accuracy you can get?\r\n\r\nThanks in advance.",
      "votes": null
    },
    {
      "id": "96050",
      "postDate": "10/13/2015 23:23:44",
      "content": "<p>Welcome to kaggle. You chose a tough competition to start with. </p>\n\n<p>When you combine two similar things together, the outcome is usually too similar to get substantial gains. So either you can use similar RF instances on a different subset of the data (or a dataset that's generated differently from the raw data, like in this competition), or you use completely different models on the same originating dataset. Those usually gives the most lift from your blends.</p>\n\n<p>What you could look into are automated methods for model tweaking (grid search for example). Those can take a long time, so you may want to choose the minimal representative subset of your data and let it run on that. Then you take the best of that and run that on a slightly larger set to tweak them some more until you home in on something that you're comfy with. </p>\n\n<p>What's common is that people work on features early on, but as time progresses, people get less happy about changing things around and spend the last of the competition tweaking in the hope that they find the golden combination. What I recommend is to get a feel early on how much extra you can squeeze by tweaking, this tells you that there's something you may be missing that others are doing.</p>\n\n<p>If you are looking for significant extra performance, it's unlikely that tweaking is going to do this for you, so it's time to look back at what you've done, figure out new features and figure out perhaps how to make things a lot more efficient, so you can run the models a lot faster and using a lot less memory. Then you learn a lot more about the data. In this competition for example, it can be extremely expensive to do a single run (20+ hours) if you've structured things a little bit simple without intermediary features.</p>",
      "rawMarkdown": "Welcome to kaggle. You chose a tough competition to start with. \r\n\r\nWhen you combine two similar things together, the outcome is usually too similar to get substantial gains. So either you can use similar RF instances on a different subset of the data (or a dataset that's generated differently from the raw data, like in this competition), or you use completely different models on the same originating dataset. Those usually gives the most lift from your blends.\r\n\r\nWhat you could look into are automated methods for model tweaking (grid search for example). Those can take a long time, so you may want to choose the minimal representative subset of your data and let it run on that. Then you take the best of that and run that on a slightly larger set to tweak them some more until you home in on something that you're comfy with. \r\n\r\nWhat's common is that people work on features early on, but as time progresses, people get less happy about changing things around and spend the last of the competition tweaking in the hope that they find the golden combination. What I recommend is to get a feel early on how much extra you can squeeze by tweaking, this tells you that there's something you may be missing that others are doing.\r\n\r\nIf you are looking for significant extra performance, it's unlikely that tweaking is going to do this for you, so it's time to look back at what you've done, figure out new features and figure out perhaps how to make things a lot more efficient, so you can run the models a lot faster and using a lot less memory. Then you learn a lot more about the data. In this competition for example, it can be extremely expensive to do a single run (20+ hours) if you've structured things a little bit simple without intermediary features.",
      "votes": null
    },
    {
      "id": "96055",
      "postDate": "10/14/2015 02:02:31",
      "content": "<p>Hi firefly2442,</p>\n\n<p>Welcome!, and regarding your questions:</p>\n\n<blockquote>\n  <p>do you spend more time thinking about and adding additional features\n  or more time tweaking or changing your prediction method?</p>\n</blockquote>\n\n<p>In this particular challenge, I think it was all about feature engineering, our best single model is 0.98422 and while not super awesome, it is enough for Top 15 (up to this moment).</p>\n\n<p>It was ridiculous hard to me how much more complex the solution need to be (in our case), in order to gain meaningless improvement, we only improved by 0.00172 vs our single model.</p>\n\n<blockquote>\n  <p>How does the timeline change what you work on? Towards the end, are\n  you doing more things like stacking or changing parameters to see what\n  kind of accuracy you can get?</p>\n</blockquote>\n\n<p>I think it's just matter of apply common sense here, depends on how things are going.</p>",
      "rawMarkdown": "Hi firefly2442,\r\n\r\nWelcome!, and regarding your questions:\r\n\r\n> do you spend more time thinking about and adding additional features\r\n> or more time tweaking or changing your prediction method?\r\n\r\nIn this particular challenge, I think it was all about feature engineering, our best single model is 0.98422 and while not super awesome, it is enough for Top 15 (up to this moment).\r\n\r\nIt was ridiculous hard to me how much more complex the solution need to be (in our case), in order to gain meaningless improvement, we only improved by 0.00172 vs our single model.\r\n\r\n\r\n> How does the timeline change what you work on? Towards the end, are\r\n> you doing more things like stacking or changing parameters to see what\r\n> kind of accuracy you can get?\r\n\r\nI think it's just matter of apply common sense here, depends on how things are going.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 96050,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "10/13/2015 23:23:44",
      "content": "<p>Welcome to kaggle. You chose a tough competition to start with. </p>\n\n<p>When you combine two similar things together, the outcome is usually too similar to get substantial gains. So either you can use similar RF instances on a different subset of the data (or a dataset that's generated differently from the raw data, like in this competition), or you use completely different models on the same originating dataset. Those usually gives the most lift from your blends.</p>\n\n<p>What you could look into are automated methods for model tweaking (grid search for example). Those can take a long time, so you may want to choose the minimal representative subset of your data and let it run on that. Then you take the best of that and run that on a slightly larger set to tweak them some more until you home in on something that you're comfy with. </p>\n\n<p>What's common is that people work on features early on, but as time progresses, people get less happy about changing things around and spend the last of the competition tweaking in the hope that they find the golden combination. What I recommend is to get a feel early on how much extra you can squeeze by tweaking, this tells you that there's something you may be missing that others are doing.</p>\n\n<p>If you are looking for significant extra performance, it's unlikely that tweaking is going to do this for you, so it's time to look back at what you've done, figure out new features and figure out perhaps how to make things a lot more efficient, so you can run the models a lot faster and using a lot less memory. Then you learn a lot more about the data. In this competition for example, it can be extremely expensive to do a single run (20+ hours) if you've structured things a little bit simple without intermediary features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 96055,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/14/2015 02:02:31",
      "content": "<p>Hi firefly2442,</p>\n\n<p>Welcome!, and regarding your questions:</p>\n\n<blockquote>\n  <p>do you spend more time thinking about and adding additional features\n  or more time tweaking or changing your prediction method?</p>\n</blockquote>\n\n<p>In this particular challenge, I think it was all about feature engineering, our best single model is 0.98422 and while not super awesome, it is enough for Top 15 (up to this moment).</p>\n\n<p>It was ridiculous hard to me how much more complex the solution need to be (in our case), in order to gain meaningless improvement, we only improved by 0.00172 vs our single model.</p>\n\n<blockquote>\n  <p>How does the timeline change what you work on? Towards the end, are\n  you doing more things like stacking or changing parameters to see what\n  kind of accuracy you can get?</p>\n</blockquote>\n\n<p>I think it's just matter of apply common sense here, depends on how things are going.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "96041": "This is my first participation in a Kaggle competition so I was curious about a couple things (since the competition is now almost over).  For my analysis I used a Random Forest with a large number of features.  My process for improvement was to come up with a new idea for additional features and then add them into the the calculation.  Almost always this improved the score, however, as I added more features, the resulting improvement dropped off substantially.  It got to the point where I started to look for other options.  So then I used stacking to combine multiple Random Forest runs together to form a better prediction.  Again, this helped a little bit but not much.  I guess my question is, do you spend more time thinking about and adding additional features or more time tweaking or changing your prediction method?  How does the timeline change what you work on?  Towards the end, are you doing more things like stacking or changing parameters to see what kind of accuracy you can get?\r\n\r\nThanks in advance.",
    "96050": "Welcome to kaggle. You chose a tough competition to start with. \r\n\r\nWhen you combine two similar things together, the outcome is usually too similar to get substantial gains. So either you can use similar RF instances on a different subset of the data (or a dataset that's generated differently from the raw data, like in this competition), or you use completely different models on the same originating dataset. Those usually gives the most lift from your blends.\r\n\r\nWhat you could look into are automated methods for model tweaking (grid search for example). Those can take a long time, so you may want to choose the minimal representative subset of your data and let it run on that. Then you take the best of that and run that on a slightly larger set to tweak them some more until you home in on something that you're comfy with. \r\n\r\nWhat's common is that people work on features early on, but as time progresses, people get less happy about changing things around and spend the last of the competition tweaking in the hope that they find the golden combination. What I recommend is to get a feel early on how much extra you can squeeze by tweaking, this tells you that there's something you may be missing that others are doing.\r\n\r\nIf you are looking for significant extra performance, it's unlikely that tweaking is going to do this for you, so it's time to look back at what you've done, figure out new features and figure out perhaps how to make things a lot more efficient, so you can run the models a lot faster and using a lot less memory. Then you learn a lot more about the data. In this competition for example, it can be extremely expensive to do a single run (20+ hours) if you've structured things a little bit simple without intermediary features.",
    "96055": "Hi firefly2442,\r\n\r\nWelcome!, and regarding your questions:\r\n\r\n> do you spend more time thinking about and adding additional features\r\n> or more time tweaking or changing your prediction method?\r\n\r\nIn this particular challenge, I think it was all about feature engineering, our best single model is 0.98422 and while not super awesome, it is enough for Top 15 (up to this moment).\r\n\r\nIt was ridiculous hard to me how much more complex the solution need to be (in our case), in order to gain meaningless improvement, we only improved by 0.00172 vs our single model.\r\n\r\n\r\n> How does the timeline change what you work on? Towards the end, are\r\n> you doing more things like stacking or changing parameters to see what\r\n> kind of accuracy you can get?\r\n\r\nI think it's just matter of apply common sense here, depends on how things are going."
  },
  "source": "meta"
}