{
  "id": 56986,
  "title": "Iteration cycles, Hyper-parameter tuning",
  "url": "/competitions/avito-demand-prediction/discussion/56986",
  "author_name": "",
  "post_date": "2018-05-17T17:56:33.476368600Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Do you have a specific strategy for testing new ideas? Since we have 4 main channels (text, categories, numerical, image) do you create individual \"small\" models for checking if a new idea helps?  I would be thankful for any insights. Do you use only a sample of the data to speed up training?</p>",
  "messages": [
    {
      "id": "329970",
      "postDate": "05/17/2018 17:56:33",
      "content": "<p>Do you have a specific strategy for testing new ideas? Since we have 4 main channels (text, categories, numerical, image) do you create individual \"small\" models for checking if a new idea helps?  I would be thankful for any insights. Do you use only a sample of the data to speed up training?</p>",
      "rawMarkdown": "Do you have a specific strategy for testing new ideas? Since we have 4 main channels (text, categories, numerical, image) do you create individual \"small\" models for checking if a new idea helps?  I would be thankful for any insights. Do you use only a sample of the data to speed up training?",
      "votes": null
    },
    {
      "id": "330023",
      "postDate": "05/17/2018 21:11:41",
      "content": "<p>I've only just started working with this data, so feel free to take my ideas with a grain of salt :)</p>\n\n<p>For feature engineering so far, I'm using greedy forward feature selection with consistent base model parameters and a single 10% validation holdout. I believe that relatively early on / through most of your competition work cycle it's a better use of time to iterate on feature engineering than to focus on hyperparameter tuning, so long as you still have good feature ideas to explore. I think there are a few key criteria that you can ideally try to meet in order to perform reliable and efficient feature testing:</p>\n\n<ol>\n<li>Base model parameters and validation scheme held constant, to isolate the predictive impact of adding/removing features</li>\n<li>Validation setup robust enough to detect real improvements in features, but fast enough to allow for rapid iteration (this is why I use only 1 holdout for now). You can decide on a cutoff of \"significance\" in RMSE difference to justify feature(s) as being worth including, and possibly run a more robust validation loop if you think it's more of a grey area case </li>\n<li>Model executes quickly, allowing for many experiments </li>\n<li>Test a small group of features at a time by adding them to your current best model. This is a nice compromise between granularity in experimentation and speed. I usually like to test features together that are of a similar style (e.g. say you add a few image quality features and try them out together)</li>\n<li>Accumulate features instead of treating them individually, to make sure that new features are a real value add instead of a proxy for information that you've already captured with other features. The problem with creating a \"small\" model like you mention is that you might find some standalone features that look very nice this way, but adding them to your overall model only adds redundant information with no gain. </li>\n</ol>\n\n<p>Luckily, lightgbm is an outstanding tool for facilitating this style of workflow. With a relatively high learning rate (.1+) you can blaze through 1.5 million records on a 16 GB / 4 core laptop. And being able to easily access feature importance can give you nice corroborating evidence that feature(s) are worth including (though be careful here, since a high importance scoring feature can easily be one that the model is overfitting to and ought be excluded). Even if your best results are from a different model (say a neural net), it's nice to have a tool like this on hand for more rapid exploration.</p>",
      "rawMarkdown": "I've only just started working with this data, so feel free to take my ideas with a grain of salt :)\n\nFor feature engineering so far, I'm using greedy forward feature selection with consistent base model parameters and a single 10% validation holdout. I believe that relatively early on / through most of your competition work cycle it's a better use of time to iterate on feature engineering than to focus on hyperparameter tuning, so long as you still have good feature ideas to explore. I think there are a few key criteria that you can ideally try to meet in order to perform reliable and efficient feature testing:\n\n 1. Base model parameters and validation scheme held constant, to isolate the predictive impact of adding/removing features\n 2. Validation setup robust enough to detect real improvements in features, but fast enough to allow for rapid iteration (this is why I use only 1 holdout for now). You can decide on a cutoff of \"significance\" in RMSE difference to justify feature(s) as being worth including, and possibly run a more robust validation loop if you think it's more of a grey area case \n 3. Model executes quickly, allowing for many experiments \n 4. Test a small group of features at a time by adding them to your current best model. This is a nice compromise between granularity in experimentation and speed. I usually like to test features together that are of a similar style (e.g. say you add a few image quality features and try them out together)\n 5. Accumulate features instead of treating them individually, to make sure that new features are a real value add instead of a proxy for information that you've already captured with other features. The problem with creating a \"small\" model like you mention is that you might find some standalone features that look very nice this way, but adding them to your overall model only adds redundant information with no gain. \n\nLuckily, lightgbm is an outstanding tool for facilitating this style of workflow. With a relatively high learning rate (.1+) you can blaze through 1.5 million records on a 16 GB / 4 core laptop. And being able to easily access feature importance can give you nice corroborating evidence that feature(s) are worth including (though be careful here, since a high importance scoring feature can easily be one that the model is overfitting to and ought be excluded). Even if your best results are from a different model (say a neural net), it's nice to have a tool like this on hand for more rapid exploration.",
      "votes": null
    },
    {
      "id": "330430",
      "postDate": "05/18/2018 19:37:21",
      "content": "<p>Thank you very much, that really helps.</p>",
      "rawMarkdown": "Thank you very much, that really helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 330023,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/17/2018 21:11:41",
      "content": "<p>I've only just started working with this data, so feel free to take my ideas with a grain of salt :)</p>\n\n<p>For feature engineering so far, I'm using greedy forward feature selection with consistent base model parameters and a single 10% validation holdout. I believe that relatively early on / through most of your competition work cycle it's a better use of time to iterate on feature engineering than to focus on hyperparameter tuning, so long as you still have good feature ideas to explore. I think there are a few key criteria that you can ideally try to meet in order to perform reliable and efficient feature testing:</p>\n\n<ol>\n<li>Base model parameters and validation scheme held constant, to isolate the predictive impact of adding/removing features</li>\n<li>Validation setup robust enough to detect real improvements in features, but fast enough to allow for rapid iteration (this is why I use only 1 holdout for now). You can decide on a cutoff of \"significance\" in RMSE difference to justify feature(s) as being worth including, and possibly run a more robust validation loop if you think it's more of a grey area case </li>\n<li>Model executes quickly, allowing for many experiments </li>\n<li>Test a small group of features at a time by adding them to your current best model. This is a nice compromise between granularity in experimentation and speed. I usually like to test features together that are of a similar style (e.g. say you add a few image quality features and try them out together)</li>\n<li>Accumulate features instead of treating them individually, to make sure that new features are a real value add instead of a proxy for information that you've already captured with other features. The problem with creating a \"small\" model like you mention is that you might find some standalone features that look very nice this way, but adding them to your overall model only adds redundant information with no gain. </li>\n</ol>\n\n<p>Luckily, lightgbm is an outstanding tool for facilitating this style of workflow. With a relatively high learning rate (.1+) you can blaze through 1.5 million records on a 16 GB / 4 core laptop. And being able to easily access feature importance can give you nice corroborating evidence that feature(s) are worth including (though be careful here, since a high importance scoring feature can easily be one that the model is overfitting to and ought be excluded). Even if your best results are from a different model (say a neural net), it's nice to have a tool like this on hand for more rapid exploration.</p>",
      "votes": null,
      "replies": [
        {
          "id": 330430,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/18/2018 19:37:21",
          "content": "<p>Thank you very much, that really helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "329970": "Do you have a specific strategy for testing new ideas? Since we have 4 main channels (text, categories, numerical, image) do you create individual \"small\" models for checking if a new idea helps?  I would be thankful for any insights. Do you use only a sample of the data to speed up training?",
    "330023": "I've only just started working with this data, so feel free to take my ideas with a grain of salt :)\n\nFor feature engineering so far, I'm using greedy forward feature selection with consistent base model parameters and a single 10% validation holdout. I believe that relatively early on / through most of your competition work cycle it's a better use of time to iterate on feature engineering than to focus on hyperparameter tuning, so long as you still have good feature ideas to explore. I think there are a few key criteria that you can ideally try to meet in order to perform reliable and efficient feature testing:\n\n 1. Base model parameters and validation scheme held constant, to isolate the predictive impact of adding/removing features\n 2. Validation setup robust enough to detect real improvements in features, but fast enough to allow for rapid iteration (this is why I use only 1 holdout for now). You can decide on a cutoff of \"significance\" in RMSE difference to justify feature(s) as being worth including, and possibly run a more robust validation loop if you think it's more of a grey area case \n 3. Model executes quickly, allowing for many experiments \n 4. Test a small group of features at a time by adding them to your current best model. This is a nice compromise between granularity in experimentation and speed. I usually like to test features together that are of a similar style (e.g. say you add a few image quality features and try them out together)\n 5. Accumulate features instead of treating them individually, to make sure that new features are a real value add instead of a proxy for information that you've already captured with other features. The problem with creating a \"small\" model like you mention is that you might find some standalone features that look very nice this way, but adding them to your overall model only adds redundant information with no gain. \n\nLuckily, lightgbm is an outstanding tool for facilitating this style of workflow. With a relatively high learning rate (.1+) you can blaze through 1.5 million records on a 16 GB / 4 core laptop. And being able to easily access feature importance can give you nice corroborating evidence that feature(s) are worth including (though be careful here, since a high importance scoring feature can easily be one that the model is overfitting to and ought be excluded). Even if your best results are from a different model (say a neural net), it's nice to have a tool like this on hand for more rapid exploration.",
    "330430": "Thank you very much, that really helps."
  },
  "source": "meta"
}