{
  "id": 348304,
  "title": "How to choose between hundreds of machine learning algorithms?",
  "url": "/competitions/amex-default-prediction/discussion/348304",
  "author_name": "Aakash Yadav",
  "post_date": "2022-08-27T20:07:33.514000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>From classical machine learning to deep learning there are hundreds of algorithms and more and more algorithms continue to publish. So, In competitions like this, time is of utmost importance, therefore, how do you choose which algorithm to start from i.e. whether to start from gradient boosting, neural networks, or just logistic regression is enough? You cannot just try each one and check its performance on the validation set, Because, you have to also do feature engineering and hyperparameter tuning therefore, one cannot just test all the algorithms. So, are there any rules that the X algorithm works best with the Y type of dataset under Z assumptions or conditions? </p>\n<p>I request top rankers of this competition and other kagglers to answer this question as I wanted to know how you choose which algorithm to spend time on. </p>",
  "messages": [
    {
      "id": 1916376,
      "postDate": "2022-08-27T20:07:33.513Z",
      "content": "<p>From classical machine learning to deep learning there are hundreds of algorithms and more and more algorithms continue to publish. So, In competitions like this, time is of utmost importance, therefore, how do you choose which algorithm to start from i.e. whether to start from gradient boosting, neural networks, or just logistic regression is enough? You cannot just try each one and check its performance on the validation set, Because, you have to also do feature engineering and hyperparameter tuning therefore, one cannot just test all the algorithms. So, are there any rules that the X algorithm works best with the Y type of dataset under Z assumptions or conditions? </p>\n<p>I request top rankers of this competition and other kagglers to answer this question as I wanted to know how you choose which algorithm to spend time on. </p>",
      "rawMarkdown": "From classical machine learning to deep learning there are hundreds of algorithms and more and more algorithms continue to publish. So, In competitions like this, time is of utmost importance, therefore, how do you choose which algorithm to start from i.e. whether to start from gradient boosting, neural networks, or just logistic regression is enough? You cannot just try each one and check its performance on the validation set, Because, you have to also do feature engineering and hyperparameter tuning therefore, one cannot just test all the algorithms. So, are there any rules that the X algorithm works best with the Y type of dataset under Z assumptions or conditions? \n\nI request top rankers of this competition and other kagglers to answer this question as I wanted to know how you choose which algorithm to spend time on. \n",
      "votes": 3
    },
    {
      "id": 1916457,
      "postDate": "2022-08-27T23:09:33.503Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/ydvaakash\" target=\"_blank\">@ydvaakash</a>. Start simple and add complexity as you go; also, build pipelines to iterate faster between model types.</p>",
      "rawMarkdown": "Hello, @ydvaakash. Start simple and add complexity as you go; also, build pipelines to iterate faster between model types.",
      "votes": 2
    },
    {
      "id": 1916645,
      "postDate": "2022-08-28T04:00:55.707Z",
      "content": "<p>Don't choose, just start with other people's work. Sometimes other people will even compile the best options for you:<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846</a> </p>\n<p>Now you are choosing between a dozen options at most, AND can get a sense of what they are like and what score they've achieved. It's much more approachable and you can start with one, learn about it, and as you learn more, in future you will be able to make more informed decision on your own. :)</p>",
      "rawMarkdown": "Don't choose, just start with other people's work. Sometimes other people will even compile the best options for you:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328846 \n\nNow you are choosing between a dozen options at most, AND can get a sense of what they are like and what score they've achieved. It's much more approachable and you can start with one, learn about it, and as you learn more, in future you will be able to make more informed decision on your own. :)",
      "replies": [
        {
          "id": 1916914,
          "postDate": "2022-08-28T09:15:44.380Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> for your answer. Actually, I adopted this methodology only, and then it got me thinking about why did they think of making ensembles of catboost, xgboost etc., did they go through all other algorithms first? </p>\n<p>Correct me if I am wrong, I believe correct approach would be here to try the popular algorithms first and not go through all the published algorithms, even if someone is getting a low score with that famous algo, because let's say I choose LGBM as my baseline, then I could get a better score as my feature engineering approach could be different from someone who is getting a lower score. And, then as correctly mentioned above by <a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a> we should build upon these chosen baseline algorithms. Then in the end, if we have some time,  we could maybe try out newer published algorithms.</p>",
          "rawMarkdown": "Thanks, @roberthatch for your answer. Actually, I adopted this methodology only, and then it got me thinking about why did they think of making ensembles of catboost, xgboost etc., did they go through all other algorithms first? \n\nCorrect me if I am wrong, I believe correct approach would be here to try the popular algorithms first and not go through all the published algorithms, even if someone is getting a low score with that famous algo, because let's say I choose LGBM as my baseline, then I could get a better score as my feature engineering approach could be different from someone who is getting a lower score. And, then as correctly mentioned above by @cv13j0 we should build upon these chosen baseline algorithms. Then in the end, if we have some time,  we could maybe try out newer published algorithms.\n"
        }
      ]
    },
    {
      "id": 1918288,
      "postDate": "2022-08-29T13:17:09.160Z",
      "content": "<p>For me, for any tabular data the goto algo #1 are boosted trees - LightGBM/XGBoost. They are relatively easy to use and tune and I believe they have the best performance per time spent rate, if you're not quite proficient with NNs.</p>\n<p>That said, it seems to me that if you want to do well in kaggle competitions, you should try as many different approaches as you can, keep the same validation scheme, keep out-of-fold validation predictions along with test predictions, and then ensemble all these models (leverage all the different approaches).</p>\n<p>I personally feel that for majority of real life usecases with tabular data, you will benefit from learning boosted trees the most. (You should definitely understand the logistic regression too though!)</p>",
      "rawMarkdown": "For me, for any tabular data the goto algo #1 are boosted trees - LightGBM/XGBoost. They are relatively easy to use and tune and I believe they have the best performance per time spent rate, if you're not quite proficient with NNs.\n\nThat said, it seems to me that if you want to do well in kaggle competitions, you should try as many different approaches as you can, keep the same validation scheme, keep out-of-fold validation predictions along with test predictions, and then ensemble all these models (leverage all the different approaches).\n\nI personally feel that for majority of real life usecases with tabular data, you will benefit from learning boosted trees the most. (You should definitely understand the logistic regression too though!)"
    },
    {
      "id": 1916993,
      "postDate": "2022-08-28T10:22:05.680Z",
      "content": "<p>I found following notes in my machine learning course in regards to this question:</p>\n<ol>\n<li>If the dataset is linear then the following popular classification algorithms will work:\n   a) Logistic Regression\n   b) LDA\n   c) Naive Byes\n   d) Support Vector machines<ol>\n<li>If the dataset is non-linear then the following popular classification algorithms will work:<br>\na) Neural Networks<br>\nb) Decision Trees<br>\nc) Random Forest<br>\nd) K-Nearer Neighbors<br>\ne) SVM( Non-Linear Kernel)</li></ol></li>\n</ol>",
      "rawMarkdown": "I found following notes in my machine learning course in regards to this question:\n  1.  If the dataset is linear then the following popular classification algorithms will work:\n       a) Logistic Regression\n       b) LDA\n       c) Naive Byes\n       d) Support Vector machines\n2.  If the dataset is non-linear then the following popular classification algorithms will work:\n     a) Neural Networks\n     b) Decision Trees\n     c) Random Forest\n     d) K-Nearer Neighbors\n     e) SVM( Non-Linear Kernel)",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1917656,
          "postDate": "2022-08-29T00:43:13.040Z",
          "content": "<p>This is good advice. With experience, we learn the \"personalities\" of different models. Then when we enter a new competition and perform EDA, we sense the \"personality\" of the data and choose a model with the same \"personality\"!</p>\n<p>Generally, NN and GBT should be your two default models to try. These work for all data and you should persist to make these work. When data is small, you can try the basic ML like logistic regression, SVM, etc. It isn't too hard to build a dozen models in the first week (which include exotic models like TabNet, DAE, DeepInsight, etc) of a competition to try the dozen most popular models. (And it is helpful to fork public notebooks for examples of new models too). Then see which combinations form the best ensemble and/or which single model achieves the best CV/LB. Next focus on improving the best models.</p>\n<p>The above applies to tabular data. Of course for images and NLP, we use the common pretrained models from timm respository and hugging face.</p>\n<p>In addition to best models, it is important to understand how to leverage all the data that Kaggle provides. For example if they provide test data, find a way to use it! Also consider what external data to acquire, what data augmentation to use, what special learning schedules, losses, sample weighting, preprocessing, post processing, etc should be used. Many competitions are won by this stuff.</p>",
          "rawMarkdown": "This is good advice. With experience, we learn the \"personalities\" of different models. Then when we enter a new competition and perform EDA, we sense the \"personality\" of the data and choose a model with the same \"personality\"!\n\nGenerally, NN and GBT should be your two default models to try. These work for all data and you should persist to make these work. When data is small, you can try the basic ML like logistic regression, SVM, etc. It isn't too hard to build a dozen models in the first week (which include exotic models like TabNet, DAE, DeepInsight, etc) of a competition to try the dozen most popular models. (And it is helpful to fork public notebooks for examples of new models too). Then see which combinations form the best ensemble and/or which single model achieves the best CV/LB. Next focus on improving the best models.\n\nThe above applies to tabular data. Of course for images and NLP, we use the common pretrained models from timm respository and hugging face.\n\nIn addition to best models, it is important to understand how to leverage all the data that Kaggle provides. For example if they provide test data, find a way to use it! Also consider what external data to acquire, what data augmentation to use, what special learning schedules, losses, sample weighting, preprocessing, post processing, etc should be used. Many competitions are won by this stuff.",
          "votes": 6
        },
        {
          "id": 1917681,
          "postDate": "2022-08-29T01:23:14.277Z",
          "content": "<p>Yes. This makes sense.</p>",
          "rawMarkdown": "Yes. This makes sense.",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1917920,
          "postDate": "2022-08-29T06:51:44.647Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. If possible, could you elaborate on this, \"if they provide test data, find a way to use it!\" ?</p>",
          "rawMarkdown": "Thanks @cdeotte. If possible, could you elaborate on this, \"if they provide test data, find a way to use it!\" ?",
          "votes": 1
        },
        {
          "id": 1918674,
          "postDate": "2022-08-29T18:28:30.073Z",
          "content": "<p>Kaggle has two types of competitions. There are (1) code competitions and (2) CSV competitions. In CSV competitions, Kaggle provides us with train data and <strong>all the test data</strong> (whether it be tabular, image, or NLP data). The train data has targets and the test data does not have targets. (In code competitions, we do not get to see the test data).</p>\n<p>In CSV Kaggle competitions, we can <strong>always use</strong> the test data to train our models and <strong>it will boost</strong> our CV LB scores. (In real life this practice isn't always possible and/or acceptable). </p>\n<p>You might think \"how can test data help because we don't have the labels?\". None-the-less, we have columns and columns of features like <code>P_2</code> and <code>D_39</code> etc. There is always benefit using the columns of test data even without the labels. Because there is information therein. We need to think of creative ways to extract this information.</p>\n<p>(The basic concept of why unlabeled data helps is explained in my pseudo label notebook <a href=\"https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969\" target=\"_blank\">here</a> from a past tabular data competition. And in computer vision and NLP competitions, there is always benefit in using the test images and/or test text even without labels. This is the idea behind using pretrained CNN on ImageNet images and/or using pretrained NLP transformers on text from the internet. More data helps models.)</p>",
          "rawMarkdown": "Kaggle has two types of competitions. There are (1) code competitions and (2) CSV competitions. In CSV competitions, Kaggle provides us with train data and **all the test data** (whether it be tabular, image, or NLP data). The train data has targets and the test data does not have targets. (In code competitions, we do not get to see the test data).\n\nIn CSV Kaggle competitions, we can **always use** the test data to train our models and **it will boost** our CV LB scores. (In real life this practice isn't always possible and/or acceptable). \n\nYou might think \"how can test data help because we don't have the labels?\". None-the-less, we have columns and columns of features like `P_2` and `D_39` etc. There is always benefit using the columns of test data even without the labels. Because there is information therein. We need to think of creative ways to extract this information.\n\n(The basic concept of why unlabeled data helps is explained in my pseudo label notebook [here][1] from a past tabular data competition. And in computer vision and NLP competitions, there is always benefit in using the test images and/or test text even without labels. This is the idea behind using pretrained CNN on ImageNet images and/or using pretrained NLP transformers on text from the internet. More data helps models.)\n\n[1]: https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969",
          "votes": 2
        },
        {
          "id": 1930674,
          "postDate": "2022-09-08T05:52:47.003Z",
          "content": "<p>What a useful things it is!</p>",
          "rawMarkdown": "What a useful things it is!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1916457,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-08-27T23:09:33.503000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/ydvaakash\" target=\"_blank\">@ydvaakash</a>. Start simple and add complexity as you go; also, build pipelines to iterate faster between model types.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1916645,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-08-28T04:00:55.707000",
      "content": "<p>Don't choose, just start with other people's work. Sometimes other people will even compile the best options for you:<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846</a> </p>\n<p>Now you are choosing between a dozen options at most, AND can get a sense of what they are like and what score they've achieved. It's much more approachable and you can start with one, learn about it, and as you learn more, in future you will be able to make more informed decision on your own. :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1916914,
          "author_name": "Aakash Yadav",
          "author_url": "",
          "post_date": "2022-08-28T09:15:44.380000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> for your answer. Actually, I adopted this methodology only, and then it got me thinking about why did they think of making ensembles of catboost, xgboost etc., did they go through all other algorithms first? </p>\n<p>Correct me if I am wrong, I believe correct approach would be here to try the popular algorithms first and not go through all the published algorithms, even if someone is getting a low score with that famous algo, because let's say I choose LGBM as my baseline, then I could get a better score as my feature engineering approach could be different from someone who is getting a lower score. And, then as correctly mentioned above by <a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a> we should build upon these chosen baseline algorithms. Then in the end, if we have some time,  we could maybe try out newer published algorithms.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1918288,
      "author_name": "MartinBarus",
      "author_url": "",
      "post_date": "2022-08-29T13:17:09.160000",
      "content": "<p>For me, for any tabular data the goto algo #1 are boosted trees - LightGBM/XGBoost. They are relatively easy to use and tune and I believe they have the best performance per time spent rate, if you're not quite proficient with NNs.</p>\n<p>That said, it seems to me that if you want to do well in kaggle competitions, you should try as many different approaches as you can, keep the same validation scheme, keep out-of-fold validation predictions along with test predictions, and then ensemble all these models (leverage all the different approaches).</p>\n<p>I personally feel that for majority of real life usecases with tabular data, you will benefit from learning boosted trees the most. (You should definitely understand the logistic regression too though!)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1916993,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-28T10:22:05.680000",
      "content": "<p>I found following notes in my machine learning course in regards to this question:</p>\n<ol>\n<li>If the dataset is linear then the following popular classification algorithms will work:\n   a) Logistic Regression\n   b) LDA\n   c) Naive Byes\n   d) Support Vector machines<ol>\n<li>If the dataset is non-linear then the following popular classification algorithms will work:<br>\na) Neural Networks<br>\nb) Decision Trees<br>\nc) Random Forest<br>\nd) K-Nearer Neighbors<br>\ne) SVM( Non-Linear Kernel)</li></ol></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1917656,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-29T00:43:13.040000",
          "content": "<p>This is good advice. With experience, we learn the \"personalities\" of different models. Then when we enter a new competition and perform EDA, we sense the \"personality\" of the data and choose a model with the same \"personality\"!</p>\n<p>Generally, NN and GBT should be your two default models to try. These work for all data and you should persist to make these work. When data is small, you can try the basic ML like logistic regression, SVM, etc. It isn't too hard to build a dozen models in the first week (which include exotic models like TabNet, DAE, DeepInsight, etc) of a competition to try the dozen most popular models. (And it is helpful to fork public notebooks for examples of new models too). Then see which combinations form the best ensemble and/or which single model achieves the best CV/LB. Next focus on improving the best models.</p>\n<p>The above applies to tabular data. Of course for images and NLP, we use the common pretrained models from timm respository and hugging face.</p>\n<p>In addition to best models, it is important to understand how to leverage all the data that Kaggle provides. For example if they provide test data, find a way to use it! Also consider what external data to acquire, what data augmentation to use, what special learning schedules, losses, sample weighting, preprocessing, post processing, etc should be used. Many competitions are won by this stuff.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1917681,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-29T01:23:14.277000",
          "content": "<p>Yes. This makes sense.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1917920,
          "author_name": "Aakash Yadav",
          "author_url": "",
          "post_date": "2022-08-29T06:51:44.647000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. If possible, could you elaborate on this, \"if they provide test data, find a way to use it!\" ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1918674,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-29T18:28:30.073000",
          "content": "<p>Kaggle has two types of competitions. There are (1) code competitions and (2) CSV competitions. In CSV competitions, Kaggle provides us with train data and <strong>all the test data</strong> (whether it be tabular, image, or NLP data). The train data has targets and the test data does not have targets. (In code competitions, we do not get to see the test data).</p>\n<p>In CSV Kaggle competitions, we can <strong>always use</strong> the test data to train our models and <strong>it will boost</strong> our CV LB scores. (In real life this practice isn't always possible and/or acceptable). </p>\n<p>You might think \"how can test data help because we don't have the labels?\". None-the-less, we have columns and columns of features like <code>P_2</code> and <code>D_39</code> etc. There is always benefit using the columns of test data even without the labels. Because there is information therein. We need to think of creative ways to extract this information.</p>\n<p>(The basic concept of why unlabeled data helps is explained in my pseudo label notebook <a href=\"https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969\" target=\"_blank\">here</a> from a past tabular data competition. And in computer vision and NLP competitions, there is always benefit in using the test images and/or test text even without labels. This is the idea behind using pretrained CNN on ImageNet images and/or using pretrained NLP transformers on text from the internet. More data helps models.)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1930674,
          "author_name": "Jackson You",
          "author_url": "",
          "post_date": "2022-09-08T05:52:47.003000",
          "content": "<p>What a useful things it is!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1916376": "From classical machine learning to deep learning there are hundreds of algorithms and more and more algorithms continue to publish. So, In competitions like this, time is of utmost importance, therefore, how do you choose which algorithm to start from i.e. whether to start from gradient boosting, neural networks, or just logistic regression is enough? You cannot just try each one and check its performance on the validation set, Because, you have to also do feature engineering and hyperparameter tuning therefore, one cannot just test all the algorithms. So, are there any rules that the X algorithm works best with the Y type of dataset under Z assumptions or conditions? \n\nI request top rankers of this competition and other kagglers to answer this question as I wanted to know how you choose which algorithm to spend time on. \n",
    "1916457": "Hello, @ydvaakash. Start simple and add complexity as you go; also, build pipelines to iterate faster between model types.",
    "1916645": "Don't choose, just start with other people's work. Sometimes other people will even compile the best options for you:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328846 \n\nNow you are choosing between a dozen options at most, AND can get a sense of what they are like and what score they've achieved. It's much more approachable and you can start with one, learn about it, and as you learn more, in future you will be able to make more informed decision on your own. :)",
    "1918288": "For me, for any tabular data the goto algo #1 are boosted trees - LightGBM/XGBoost. They are relatively easy to use and tune and I believe they have the best performance per time spent rate, if you're not quite proficient with NNs.\n\nThat said, it seems to me that if you want to do well in kaggle competitions, you should try as many different approaches as you can, keep the same validation scheme, keep out-of-fold validation predictions along with test predictions, and then ensemble all these models (leverage all the different approaches).\n\nI personally feel that for majority of real life usecases with tabular data, you will benefit from learning boosted trees the most. (You should definitely understand the logistic regression too though!)",
    "1916993": "I found following notes in my machine learning course in regards to this question:\n  1.  If the dataset is linear then the following popular classification algorithms will work:\n       a) Logistic Regression\n       b) LDA\n       c) Naive Byes\n       d) Support Vector machines\n2.  If the dataset is non-linear then the following popular classification algorithms will work:\n     a) Neural Networks\n     b) Decision Trees\n     c) Random Forest\n     d) K-Nearer Neighbors\n     e) SVM( Non-Linear Kernel)"
  }
}