{
  "id": 215834,
  "title": "Ensemble learning question",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/215834",
  "author_name": "Alien",
  "post_date": "2021-01-31T12:54:12.483000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am new to kaggle, and I have only recently started to learn deep learning, and I want to learn through competitions. I have some questions about Ensemble learning. My current score is 0.966, which is an average of confidence through my model and 0.965 model(public notebook). If I find the best result by adjusting the weights, does it make sense to do so? Will it overfitting?</p>\n<p>I have two models, combined with the 0.965 model, can improve the performance. Next, I want to try the three-model Ensemble learning and use the weighted average and voting ensemble. Regarding the auc score, is it the one with the most confidence after voting? </p>\n<p>Thank you very much in advance!</p>",
  "messages": [
    {
      "id": 1179530,
      "postDate": "2021-01-31T15:58:41.737Z",
      "content": "<p>Most of the times, simble averaging just works best. However if you get saturated on the models which you make ensemble, only then it will be vital to try more sophisticated averaging like weigth averaging, confidence voting averaging, power averaging.</p>",
      "rawMarkdown": "Most of the times, simble averaging just works best. However if you get saturated on the models which you make ensemble, only then it will be vital to try more sophisticated averaging like weigth averaging, confidence voting averaging, power averaging.",
      "votes": 7,
      "replies": [
        {
          "id": 1179926,
          "postDate": "2021-02-01T01:33:27.213Z",
          "content": "<p>When I google ensemble learning, I get many complicated methods. Now I got it! Most of the times, simble averaging just works best. Thank you for your answer. </p>",
          "rawMarkdown": "When I google ensemble learning, I get many complicated methods. Now I got it! Most of the times, simble averaging just works best. Thank you for your answer. "
        },
        {
          "id": 1220657,
          "postDate": "2021-02-28T08:17:42.060Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1179725,
      "postDate": "2021-01-31T18:58:45.163Z",
      "content": "<p>When you want to go beyond simple averages (and there's even a <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">good question</a> as to what averages one should use), the standard way to avoid overfitting would be, as usual, based on cross validation. You fit all models to identical cross validation schemes (i.e. the same observations are in the training and validation parts of each fold for all the models). Then you use the out of fold predictions to try out ways of combining the predictions. You, again, use the same splits for cross validation to pick hyper parameters (e.g. how much weights for each model should be shrunk towards a simple average). In the <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">thread I linked above</a>, there's also some ideas as to what models one could use to combine predictions.</p>",
      "rawMarkdown": "When you want to go beyond simple averages (and there's even a [good question](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221) as to what averages one should use), the standard way to avoid overfitting would be, as usual, based on cross validation. You fit all models to identical cross validation schemes (i.e. the same observations are in the training and validation parts of each fold for all the models). Then you use the out of fold predictions to try out ways of combining the predictions. You, again, use the same splits for cross validation to pick hyper parameters (e.g. how much weights for each model should be shrunk towards a simple average). In the [thread I linked above](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221), there's also some ideas as to what models one could use to combine predictions.",
      "votes": 3,
      "replies": [
        {
          "id": 1179922,
          "postDate": "2021-02-01T01:27:38.223Z",
          "content": "<p>Thanks for your answer. I learned a lot from your answer.</p>",
          "rawMarkdown": "Thanks for your answer. I learned a lot from your answer."
        }
      ]
    },
    {
      "id": 1179247,
      "postDate": "2021-01-31T12:54:12.483Z",
      "content": "<p>I am new to kaggle, and I have only recently started to learn deep learning, and I want to learn through competitions. I have some questions about Ensemble learning. My current score is 0.966, which is an average of confidence through my model and 0.965 model(public notebook). If I find the best result by adjusting the weights, does it make sense to do so? Will it overfitting?</p>\n<p>I have two models, combined with the 0.965 model, can improve the performance. Next, I want to try the three-model Ensemble learning and use the weighted average and voting ensemble. Regarding the auc score, is it the one with the most confidence after voting? </p>\n<p>Thank you very much in advance!</p>",
      "rawMarkdown": "I am new to kaggle, and I have only recently started to learn deep learning, and I want to learn through competitions. I have some questions about Ensemble learning. My current score is 0.966, which is an average of confidence through my model and 0.965 model(public notebook). If I find the best result by adjusting the weights, does it make sense to do so? Will it overfitting?\n\nI have two models, combined with the 0.965 model, can improve the performance. Next, I want to try the three-model Ensemble learning and use the weighted average and voting ensemble. Regarding the auc score, is it the one with the most confidence after voting? \n\nThank you very much in advance!",
      "votes": 1
    },
    {
      "id": 1220658,
      "postDate": "2021-02-28T08:17:51.067Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1179530,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-31T15:58:41.737000",
      "content": "<p>Most of the times, simble averaging just works best. However if you get saturated on the models which you make ensemble, only then it will be vital to try more sophisticated averaging like weigth averaging, confidence voting averaging, power averaging.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1179926,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-02-01T01:33:27.213000",
          "content": "<p>When I google ensemble learning, I get many complicated methods. Now I got it! Most of the times, simble averaging just works best. Thank you for your answer. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220657,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-28T08:17:42.060000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1179725,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-01-31T18:58:45.163000",
      "content": "<p>When you want to go beyond simple averages (and there's even a <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">good question</a> as to what averages one should use), the standard way to avoid overfitting would be, as usual, based on cross validation. You fit all models to identical cross validation schemes (i.e. the same observations are in the training and validation parts of each fold for all the models). Then you use the out of fold predictions to try out ways of combining the predictions. You, again, use the same splits for cross validation to pick hyper parameters (e.g. how much weights for each model should be shrunk towards a simple average). In the <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">thread I linked above</a>, there's also some ideas as to what models one could use to combine predictions.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1179922,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-02-01T01:27:38.223000",
          "content": "<p>Thanks for your answer. I learned a lot from your answer.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1220658,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-28T08:17:51.067000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1179530": "Most of the times, simble averaging just works best. However if you get saturated on the models which you make ensemble, only then it will be vital to try more sophisticated averaging like weigth averaging, confidence voting averaging, power averaging.",
    "1179725": "When you want to go beyond simple averages (and there's even a [good question](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221) as to what averages one should use), the standard way to avoid overfitting would be, as usual, based on cross validation. You fit all models to identical cross validation schemes (i.e. the same observations are in the training and validation parts of each fold for all the models). Then you use the out of fold predictions to try out ways of combining the predictions. You, again, use the same splits for cross validation to pick hyper parameters (e.g. how much weights for each model should be shrunk towards a simple average). In the [thread I linked above](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221), there's also some ideas as to what models one could use to combine predictions.",
    "1179247": "I am new to kaggle, and I have only recently started to learn deep learning, and I want to learn through competitions. I have some questions about Ensemble learning. My current score is 0.966, which is an average of confidence through my model and 0.965 model(public notebook). If I find the best result by adjusting the weights, does it make sense to do so? Will it overfitting?\n\nI have two models, combined with the 0.965 model, can improve the performance. Next, I want to try the three-model Ensemble learning and use the weighted average and voting ensemble. Regarding the auc score, is it the one with the most confidence after voting? \n\nThank you very much in advance!",
    "1220658": ""
  }
}