{
  "id": 339112,
  "title": "Can anything beat lightgbm and xgboost?",
  "url": "/competitions/amex-default-prediction/discussion/339112",
  "author_name": "",
  "post_date": "2022-07-23T12:08:36.498717500Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3874081%2F740bbd31f1cf9dae2461ccf613f2624e%2Ftrue%20dat.jpg?generation=1658576488993746&amp;alt=media\" alt=\"\"></p>\n<p><strong>Answer</strong> : Yes and no. </p>\n<p>for any competition strong features can enhance any model so in theory you can get a better score from any random model if you have strong features… That's how the small can steal the ball.</p>\n<p>Specifically for this competition I have noticed that not that much attention and finesse has been thought of in terms of feature engineering it's just the same old <strong>\"brute force\"</strong> feature engineering where you aggregate the values of a specific customer_ID and get the mean ,std,min and max(so on and so forth). Don't get me wrong it has proven effective especially with the large dataset at our disposal but in doing that aren't we losing subtle information? </p>\n<p>The mammoth task of even trying to think of features to engineer is hard on any given day now trying to do that on a dataset this huge!. That might be a factor for the lack of creativity in terms of features one can engineer.There have been a few interesting features to engineer put up in the discussions but for the most part everyone seems sure of the models they have and it just seems hyper-parameter tuning , ensembling and blending predictions is what they believe can help them win it all… Because we are in the final stretch who really is still feature engineering ? </p>\n<p>for one Me</p>\n<p>I believe strong features &gt; strong model</p>\n<p>maybe someone has a different opinion. let me know.</p>",
  "messages": [
    {
      "id": "1867637",
      "postDate": "07/23/2022 12:08:36",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3874081%2F740bbd31f1cf9dae2461ccf613f2624e%2Ftrue%20dat.jpg?generation=1658576488993746&amp;alt=media\" alt=\"\"></p>\n<p><strong>Answer</strong> : Yes and no. </p>\n<p>for any competition strong features can enhance any model so in theory you can get a better score from any random model if you have strong features… That's how the small can steal the ball.</p>\n<p>Specifically for this competition I have noticed that not that much attention and finesse has been thought of in terms of feature engineering it's just the same old <strong>\"brute force\"</strong> feature engineering where you aggregate the values of a specific customer_ID and get the mean ,std,min and max(so on and so forth). Don't get me wrong it has proven effective especially with the large dataset at our disposal but in doing that aren't we losing subtle information? </p>\n<p>The mammoth task of even trying to think of features to engineer is hard on any given day now trying to do that on a dataset this huge!. That might be a factor for the lack of creativity in terms of features one can engineer.There have been a few interesting features to engineer put up in the discussions but for the most part everyone seems sure of the models they have and it just seems hyper-parameter tuning , ensembling and blending predictions is what they believe can help them win it all… Because we are in the final stretch who really is still feature engineering ? </p>\n<p>for one Me</p>\n<p>I believe strong features &gt; strong model</p>\n<p>maybe someone has a different opinion. let me know.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3874081%2F740bbd31f1cf9dae2461ccf613f2624e%2Ftrue%20dat.jpg?generation=1658576488993746&alt=media)\n\n\n**Answer** : Yes and no. \n\nfor any competition strong features can enhance any model so in theory you can get a better score from any random model if you have strong features... That's how the small can steal the ball.\n\nSpecifically for this competition I have noticed that not that much attention and finesse has been thought of in terms of feature engineering it's just the same old **\"brute force\"** feature engineering where you aggregate the values of a specific customer_ID and get the mean ,std,min and max(so on and so forth). Don't get me wrong it has proven effective especially with the large dataset at our disposal but in doing that aren't we losing subtle information? \n\nThe mammoth task of even trying to think of features to engineer is hard on any given day now trying to do that on a dataset this huge!. That might be a factor for the lack of creativity in terms of features one can engineer.There have been a few interesting features to engineer put up in the discussions but for the most part everyone seems sure of the models they have and it just seems hyper-parameter tuning , ensembling and blending predictions is what they believe can help them win it all... Because we are in the final stretch who really is still feature engineering ? \n\nfor one Me\n\nI believe strong features > strong model\n\nmaybe someone has a different opinion. let me know.",
      "votes": null
    },
    {
      "id": "1867823",
      "postDate": "07/23/2022 14:25:06",
      "content": "<p>I think you are right, feature engineering is not a one time activity but an iterative process based on the model results and scope for improvement. You may continue to do this task till the end of the competition to improve on your contemporary scores.</p>",
      "rawMarkdown": "I think you are right, feature engineering is not a one time activity but an iterative process based on the model results and scope for improvement. You may continue to do this task till the end of the competition to improve on your contemporary scores.",
      "votes": null
    },
    {
      "id": "1868113",
      "postDate": "07/23/2022 18:03:02",
      "content": "<p>It's not surprising that feature engineering is difficult here. There are a lot of very strong features in the raw data (we even get an internal model already in <code>P_2</code>!), the column descriptions are mostly anonymized making it harder to apply domain reasoning, and there aren't many training labels to support truly extensive FE -- a few hundred thousand is really pretty small. I think a lot of people are trying very hard to engineer new features but not having much luck, especially when the publicly known features set a pretty high bar. I'd disagree that these are just \"brute force\" and instead suggest that they already do a great job capturing the bulk of the signal in the data with a simple methodology. After all, kernels with just these features are within .002 of the top LB scores.</p>\n<p>That said, I'm almost certain that top scorers have discovered some differentiating features -- that would usually be my guess, but you can also find people hinting at it when describing their workflow and single model results. You'll see that level of creativity shared after the competition ends :)</p>",
      "rawMarkdown": "It's not surprising that feature engineering is difficult here. There are a lot of very strong features in the raw data (we even get an internal model already in `P_2 `!), the column descriptions are mostly anonymized making it harder to apply domain reasoning, and there aren't many training labels to support truly extensive FE -- a few hundred thousand is really pretty small. I think a lot of people are trying very hard to engineer new features but not having much luck, especially when the publicly known features set a pretty high bar. I'd disagree that these are just \"brute force\" and instead suggest that they already do a great job capturing the bulk of the signal in the data with a simple methodology. After all, kernels with just these features are within .002 of the top LB scores.\n\nThat said, I'm almost certain that top scorers have discovered some differentiating features -- that would usually be my guess, but you can also find people hinting at it when describing their workflow and single model results. You'll see that level of creativity shared after the competition ends :)",
      "votes": null
    },
    {
      "id": "1868203",
      "postDate": "07/23/2022 19:29:27",
      "content": "<p>Feature engineering is a difficult process here as you have to do that iteratively to improve the scores </p>",
      "rawMarkdown": "Feature engineering is a difficult process here as you have to do that iteratively to improve the scores",
      "votes": null
    },
    {
      "id": "1868339",
      "postDate": "07/23/2022 23:21:02",
      "content": "<p>It is difficult to beat boosted trees when dealing with tabular data. I think neural networks would be the only possible contender, but it is paramount to find data representations that works well with them.</p>\n<p>Pretty sure that most of the top scores are ensembles. I get 0.798 by ensembling 10-11 models, none of which is over 0.795 on public LB, and some are as low as 0.781. Those that have a single robust model at 0.799 should easily be able to get an ensemble score that is 0.001-0.002 higher than the best individual model.</p>",
      "rawMarkdown": "It is difficult to beat boosted trees when dealing with tabular data. I think neural networks would be the only possible contender, but it is paramount to find data representations that works well with them.\n\nPretty sure that most of the top scores are ensembles. I get 0.798 by ensembling 10-11 models, none of which is over 0.795 on public LB, and some are as low as 0.781. Those that have a single robust model at 0.799 should easily be able to get an ensemble score that is 0.001-0.002 higher than the best individual model.",
      "votes": null
    },
    {
      "id": "1869535",
      "postDate": "07/24/2022 20:25:00",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/wuuthraad\" target=\"_blank\">@wuuthraad</a>, thanks for the post; I believe in feature engineering for this competition. However, it is difficult to aggregate or build complex features quickly enough on a Kernel due to the massive amount of data; I am planning to develop a representative subset that I can experiment with and then extrapolate to the entire dataset in a distributed way.</p>",
      "rawMarkdown": "Hello @wuuthraad, thanks for the post; I believe in feature engineering for this competition. However, it is difficult to aggregate or build complex features quickly enough on a Kernel due to the massive amount of data; I am planning to develop a representative subset that I can experiment with and then extrapolate to the entire dataset in a distributed way.",
      "votes": null
    },
    {
      "id": "1871362",
      "postDate": "07/26/2022 08:29:45",
      "content": "<p>I completely agree with you. I think that feature engineering is really important and can make a big difference in the final results.<br>\nI also think that, in this particular competition, there is still a lot of room for improvement in terms of features. For example, I think that there could be a lot of information in features we simply don't the meaning of.. </p>",
      "rawMarkdown": "I completely agree with you. I think that feature engineering is really important and can make a big difference in the final results.\nI also think that, in this particular competition, there is still a lot of room for improvement in terms of features. For example, I think that there could be a lot of information in features we simply don't the meaning of..",
      "votes": null
    },
    {
      "id": "1874956",
      "postDate": "07/28/2022 16:13:05",
      "content": "<p>I agree with you</p>",
      "rawMarkdown": "I agree with you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1867823,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/23/2022 14:25:06",
      "content": "<p>I think you are right, feature engineering is not a one time activity but an iterative process based on the model results and scope for improvement. You may continue to do this task till the end of the competition to improve on your contemporary scores.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1868113,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "07/23/2022 18:03:02",
      "content": "<p>It's not surprising that feature engineering is difficult here. There are a lot of very strong features in the raw data (we even get an internal model already in <code>P_2</code>!), the column descriptions are mostly anonymized making it harder to apply domain reasoning, and there aren't many training labels to support truly extensive FE -- a few hundred thousand is really pretty small. I think a lot of people are trying very hard to engineer new features but not having much luck, especially when the publicly known features set a pretty high bar. I'd disagree that these are just \"brute force\" and instead suggest that they already do a great job capturing the bulk of the signal in the data with a simple methodology. After all, kernels with just these features are within .002 of the top LB scores.</p>\n<p>That said, I'm almost certain that top scorers have discovered some differentiating features -- that would usually be my guess, but you can also find people hinting at it when describing their workflow and single model results. You'll see that level of creativity shared after the competition ends :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1868203,
      "author_name": "omkarbhosale0606",
      "author_url": "",
      "post_date": "07/23/2022 19:29:27",
      "content": "<p>Feature engineering is a difficult process here as you have to do that iteratively to improve the scores </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1868339,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/23/2022 23:21:02",
      "content": "<p>It is difficult to beat boosted trees when dealing with tabular data. I think neural networks would be the only possible contender, but it is paramount to find data representations that works well with them.</p>\n<p>Pretty sure that most of the top scores are ensembles. I get 0.798 by ensembling 10-11 models, none of which is over 0.795 on public LB, and some are as low as 0.781. Those that have a single robust model at 0.799 should easily be able to get an ensemble score that is 0.001-0.002 higher than the best individual model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869535,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "07/24/2022 20:25:00",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/wuuthraad\" target=\"_blank\">@wuuthraad</a>, thanks for the post; I believe in feature engineering for this competition. However, it is difficult to aggregate or build complex features quickly enough on a Kernel due to the massive amount of data; I am planning to develop a representative subset that I can experiment with and then extrapolate to the entire dataset in a distributed way.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1871362,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/26/2022 08:29:45",
      "content": "<p>I completely agree with you. I think that feature engineering is really important and can make a big difference in the final results.<br>\nI also think that, in this particular competition, there is still a lot of room for improvement in terms of features. For example, I think that there could be a lot of information in features we simply don't the meaning of.. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1874956,
      "author_name": "zhehaoliang",
      "author_url": "",
      "post_date": "07/28/2022 16:13:05",
      "content": "<p>I agree with you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1867637": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3874081%2F740bbd31f1cf9dae2461ccf613f2624e%2Ftrue%20dat.jpg?generation=1658576488993746&alt=media)\n\n\n**Answer** : Yes and no. \n\nfor any competition strong features can enhance any model so in theory you can get a better score from any random model if you have strong features... That's how the small can steal the ball.\n\nSpecifically for this competition I have noticed that not that much attention and finesse has been thought of in terms of feature engineering it's just the same old **\"brute force\"** feature engineering where you aggregate the values of a specific customer_ID and get the mean ,std,min and max(so on and so forth). Don't get me wrong it has proven effective especially with the large dataset at our disposal but in doing that aren't we losing subtle information? \n\nThe mammoth task of even trying to think of features to engineer is hard on any given day now trying to do that on a dataset this huge!. That might be a factor for the lack of creativity in terms of features one can engineer.There have been a few interesting features to engineer put up in the discussions but for the most part everyone seems sure of the models they have and it just seems hyper-parameter tuning , ensembling and blending predictions is what they believe can help them win it all... Because we are in the final stretch who really is still feature engineering ? \n\nfor one Me\n\nI believe strong features > strong model\n\nmaybe someone has a different opinion. let me know.",
    "1867823": "I think you are right, feature engineering is not a one time activity but an iterative process based on the model results and scope for improvement. You may continue to do this task till the end of the competition to improve on your contemporary scores.",
    "1868113": "It's not surprising that feature engineering is difficult here. There are a lot of very strong features in the raw data (we even get an internal model already in `P_2 `!), the column descriptions are mostly anonymized making it harder to apply domain reasoning, and there aren't many training labels to support truly extensive FE -- a few hundred thousand is really pretty small. I think a lot of people are trying very hard to engineer new features but not having much luck, especially when the publicly known features set a pretty high bar. I'd disagree that these are just \"brute force\" and instead suggest that they already do a great job capturing the bulk of the signal in the data with a simple methodology. After all, kernels with just these features are within .002 of the top LB scores.\n\nThat said, I'm almost certain that top scorers have discovered some differentiating features -- that would usually be my guess, but you can also find people hinting at it when describing their workflow and single model results. You'll see that level of creativity shared after the competition ends :)",
    "1868203": "Feature engineering is a difficult process here as you have to do that iteratively to improve the scores",
    "1868339": "It is difficult to beat boosted trees when dealing with tabular data. I think neural networks would be the only possible contender, but it is paramount to find data representations that works well with them.\n\nPretty sure that most of the top scores are ensembles. I get 0.798 by ensembling 10-11 models, none of which is over 0.795 on public LB, and some are as low as 0.781. Those that have a single robust model at 0.799 should easily be able to get an ensemble score that is 0.001-0.002 higher than the best individual model.",
    "1869535": "Hello @wuuthraad, thanks for the post; I believe in feature engineering for this competition. However, it is difficult to aggregate or build complex features quickly enough on a Kernel due to the massive amount of data; I am planning to develop a representative subset that I can experiment with and then extrapolate to the entire dataset in a distributed way.",
    "1871362": "I completely agree with you. I think that feature engineering is really important and can make a big difference in the final results.\nI also think that, in this particular competition, there is still a lot of room for improvement in terms of features. For example, I think that there could be a lot of information in features we simply don't the meaning of..",
    "1874956": "I agree with you"
  },
  "source": "meta"
}