{
  "id": 174602,
  "title": "Ensemble Techniques?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174602",
  "author_name": "",
  "post_date": "2020-08-14T09:11:22.007735100Z",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>With a few days to go and after many many experiments and building dozens of Cross Validation (CV) models over the last three months, I now have nine, what I hope are, solid sets of predictions from nine different models.</p>\n<p>For all nine sets of predictions I have AUC and Loss scores from Out Of Fold (OOF) predictions, Mean CV predictions and fully trained model predictions including training scores and Kaggle Leaderboard (LB) scores.</p>\n<p>I have used some trial and error on my ensembles by observing all my AUC scores and also observing which models create good diversity in predictions against other models with similar AUC scores and weighting these accordingly.</p>\n<p>Can anyone recommend a more procedural way of building an ensemble in this situation? For example has anyone used Standard Deviation as a way to measure diversity for their ensembles?</p>\n<p>I have seen and tried ensemble methods used where we have target values, such as the best OOF ensembles with and without meta data, but not where we don't have target values.</p>\n<p>Many thanks and good luck everyone.</p>\n<p>ps this is my first Kaggle competition and it has been an awesome experience, particularly observing how people share and contribute to ideas, hat tipped to an incredible community here.</p>",
  "messages": [
    {
      "id": "970179",
      "postDate": "08/14/2020 09:11:22",
      "content": "<p>With a few days to go and after many many experiments and building dozens of Cross Validation (CV) models over the last three months, I now have nine, what I hope are, solid sets of predictions from nine different models.</p>\n<p>For all nine sets of predictions I have AUC and Loss scores from Out Of Fold (OOF) predictions, Mean CV predictions and fully trained model predictions including training scores and Kaggle Leaderboard (LB) scores.</p>\n<p>I have used some trial and error on my ensembles by observing all my AUC scores and also observing which models create good diversity in predictions against other models with similar AUC scores and weighting these accordingly.</p>\n<p>Can anyone recommend a more procedural way of building an ensemble in this situation? For example has anyone used Standard Deviation as a way to measure diversity for their ensembles?</p>\n<p>I have seen and tried ensemble methods used where we have target values, such as the best OOF ensembles with and without meta data, but not where we don't have target values.</p>\n<p>Many thanks and good luck everyone.</p>\n<p>ps this is my first Kaggle competition and it has been an awesome experience, particularly observing how people share and contribute to ideas, hat tipped to an incredible community here.</p>",
      "rawMarkdown": "With a few days to go and after many many experiments and building dozens of Cross Validation (CV) models over the last three months, I now have nine, what I hope are, solid sets of predictions from nine different models.\n\nFor all nine sets of predictions I have AUC and Loss scores from Out Of Fold (OOF) predictions, Mean CV predictions and fully trained model predictions including training scores and Kaggle Leaderboard (LB) scores.\n\nI have used some trial and error on my ensembles by observing all my AUC scores and also observing which models create good diversity in predictions against other models with similar AUC scores and weighting these accordingly.\n\nCan anyone recommend a more procedural way of building an ensemble in this situation? For example has anyone used Standard Deviation as a way to measure diversity for their ensembles?\n\nI have seen and tried ensemble methods used where we have target values, such as the best OOF ensembles with and without meta data, but not where we don't have target values.\n\nMany thanks and good luck everyone.\n\nps this is my first Kaggle competition and it has been an awesome experience, particularly observing how people share and contribute to ideas, hat tipped to an incredible community here.",
      "votes": null
    },
    {
      "id": "970190",
      "postDate": "08/14/2020 09:17:59",
      "content": "<p>384th rank is your first kaggle competition is awesome. Hope you get that top 10% for a bronze.</p>",
      "rawMarkdown": "384th rank is your first kaggle competition is awesome. Hope you get that top 10% for a bronze.",
      "votes": null
    },
    {
      "id": "970204",
      "postDate": "08/14/2020 09:26:32",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/fahad7\" target=\"_blank\">@fahad7</a>. Three months of around five hours a day, I'll be pleased to be in the top half, which was my original goal.</p>",
      "rawMarkdown": "Thanks @fahad7. Three months of around five hours a day, I'll be pleased to be in the top half, which was my original goal.",
      "votes": null
    },
    {
      "id": "970209",
      "postDate": "08/14/2020 09:27:58",
      "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> opened a discussion before presenting different methods, you can also see what people recommend or what worked and what did not:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a></p>",
      "rawMarkdown": "sirishks opened a discussion before presenting different methods, you can also see what people recommend or what worked and what did not:\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653",
      "votes": null
    },
    {
      "id": "970216",
      "postDate": "08/14/2020 09:30:41",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I'll check that out.</p>",
      "rawMarkdown": "Thanks @aliabdin1 I'll check that out.",
      "votes": null
    },
    {
      "id": "970403",
      "postDate": "08/14/2020 12:31:53",
      "content": "<p>I myself am a newbie here, I am ensembling all my good models just by Overfitting the CV. 😄 Please don't judge me for this, this is the first time I am doing ensembling and just following the quote, \"Always trust your CV\". </p>\n<p>I have made a flowchart for the same and I am using weighted power ensemble with ranking.</p>",
      "rawMarkdown": "I myself am a newbie here, I am ensembling all my good models just by Overfitting the CV. 😄 Please don't judge me for this, this is the first time I am doing ensembling and just following the quote, \"Always trust your CV\". \n\nI have made a flowchart for the same and I am using weighted power ensemble with ranking.",
      "votes": null
    },
    {
      "id": "970415",
      "postDate": "08/14/2020 12:48:37",
      "content": "<p>You and me both <a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a>! I'll check out the weighted power ensemble with ranking technique.</p>",
      "rawMarkdown": "You and me both @sarques! I'll check out the weighted power ensemble with ranking technique.",
      "votes": null
    },
    {
      "id": "970420",
      "postDate": "08/14/2020 12:51:42",
      "content": "<p>Good luck with that! :)</p>",
      "rawMarkdown": "Good luck with that! :)",
      "votes": null
    },
    {
      "id": "970497",
      "postDate": "08/14/2020 14:05:37",
      "content": "<p>I'm currently using Weighted Averaging using relatively arbitrary weights based on my CV scores. Interesting, even applying a small weight to smaller models with lower CV scores gives an overall boost of c. 0.0010. </p>\n<p>I will certainly experiment with Rank Averaging with a great explanation here <a href=\"url\" target=\"_blank\">https://www.kaggle.com/niteshx2/improve-blending-using-rankdata</a> from <a href=\"https://www.kaggle.com/niteshchoudhary\" target=\"_blank\">@niteshchoudhary</a>.</p>\n<p>I'm hesitant with Power Averaging due to correlation issues, I have deliberately created diversity in my models to try and ensure that the weakness in some models are covered by strengths in others and vice-versa. There could be value in performing Power Averaging in small groups of models which have stronger correlation and then using Rank Averaging on these group subsets.</p>",
      "rawMarkdown": "I'm currently using Weighted Averaging using relatively arbitrary weights based on my CV scores. Interesting, even applying a small weight to smaller models with lower CV scores gives an overall boost of c. 0.0010. \n\nI will certainly experiment with Rank Averaging with a great explanation here [https://www.kaggle.com/niteshx2/improve-blending-using-rankdata](url) from @niteshchoudhary.\n\nI'm hesitant with Power Averaging due to correlation issues, I have deliberately created diversity in my models to try and ensure that the weakness in some models are covered by strengths in others and vice-versa. There could be value in performing Power Averaging in small groups of models which have stronger correlation and then using Rank Averaging on these group subsets.",
      "votes": null
    },
    {
      "id": "970523",
      "postDate": "08/14/2020 14:30:23",
      "content": "<p>Yup, it definitely is good, At first I used only 0.90+ CV models and when combined using this, I got a CV score of about 0.92+. It is rocking in that way!</p>",
      "rawMarkdown": "Yup, it definitely is good, At first I used only 0.90+ CV models and when combined using this, I got a CV score of about 0.92+. It is rocking in that way!",
      "votes": null
    },
    {
      "id": "970625",
      "postDate": "08/14/2020 16:06:15",
      "content": "<p>That sounds good, but focus more on the model itself. It could be, that you are overfitting the LB with your tiny adjustments. That means you will have a higher score in public set but later on the private set it could be way lower. Thats why it is suggested to tune your results just slightly and focus on the model itself and your data more. Especially in this competition we have a very imbalanced dataset and the public LB consists of only 30% of the complete test data. </p>",
      "rawMarkdown": "That sounds good, but focus more on the model itself. It could be, that you are overfitting the LB with your tiny adjustments. That means you will have a higher score in public set but later on the private set it could be way lower. Thats why it is suggested to tune your results just slightly and focus on the model itself and your data more. Especially in this competition we have a very imbalanced dataset and the public LB consists of only 30% of the complete test data.",
      "votes": null
    },
    {
      "id": "970849",
      "postDate": "08/14/2020 21:49:46",
      "content": "<p>Resisting the temptation to chase a Public LB score and focus instead on solid models derived from experimentation and stable CV scores with a standard ensemble is very wise advice <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> . That temptation is hard to resist though!</p>",
      "rawMarkdown": "Resisting the temptation to chase a Public LB score and focus instead on solid models derived from experimentation and stable CV scores with a standard ensemble is very wise advice @aliabdin1 . That temptation is hard to resist though!",
      "votes": null
    },
    {
      "id": "970893",
      "postDate": "08/15/2020 00:44:08",
      "content": "<p>Have you considered creating a model using the out of fold predictions as input. For example, logistic regression. Or MLP. This is what I do and has worked well for me.</p>",
      "rawMarkdown": "Have you considered creating a model using the out of fold predictions as input. For example, logistic regression. Or MLP. This is what I do and has worked well for me.",
      "votes": null
    },
    {
      "id": "970987",
      "postDate": "08/15/2020 04:30:17",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> What I am doing is <code>overfitting</code> the <code>CV</code> scores, isn't it right, :) </p>\n<p>I'm am not even considering the LB score for the ensemble, I'm just taking the CV scores into consideration.</p>",
      "rawMarkdown": "Hey @aliabdin1 What I am doing is `overfitting` the `CV` scores, isn't it right, :) \n\nI'm am not even considering the LB score for the ensemble, I'm just taking the CV scores into consideration.",
      "votes": null
    },
    {
      "id": "972082",
      "postDate": "08/16/2020 08:06:53",
      "content": "<p>Having spent some time and thought on this over the last two days, I have decided to use a weighted average ensemble with the following strategy. I have deliberately worked on trying to build uncorrelated yet promising models over the last few months. </p>\n<ol>\n<li>Observe and then group together/bucket any highly correlated predictions from different models/inputs.</li>\n<li>Use the cross validation score for each model in each group/bucket and determine an average CV score for each group/bucket</li>\n<li>Build an ensemble assigning a higher weighting to groups/buckets with higher average CV scores.</li>\n<li>Don’t be put off by models or groups of models with lower CV scores (Eg meta only models or smaller CNN networks with smaller dimension inputs) as these will add important diversity to the ensemble. </li>\n<li>Don’t focus on LB scores but do use these as a guide to determine stability of CV scores. </li>\n</ol>\n<p>I haven’t used power weighting as I already have diversity and power weighting appears to be much better suited to highly correlated groups of ensembles. </p>\n<p>I am considering using a rank ensemble as AUC is determined by the correct rank and separation of target values but I need to research this more. I may attempt a weighted rank strategy similar to the above. </p>\n<p>Good luck everyone. </p>",
      "rawMarkdown": "Having spent some time and thought on this over the last two days, I have decided to use a weighted average ensemble with the following strategy. I have deliberately worked on trying to build uncorrelated yet promising models over the last few months. \n\n1. Observe and then group together/bucket any highly correlated predictions from different models/inputs.\n2. Use the cross validation score for each model in each group/bucket and determine an average CV score for each group/bucket\n3. Build an ensemble assigning a higher weighting to groups/buckets with higher average CV scores.\n4. Don’t be put off by models or groups of models with lower CV scores (Eg meta only models or smaller CNN networks with smaller dimension inputs) as these will add important diversity to the ensemble. \n5. Don’t focus on LB scores but do use these as a guide to determine stability of CV scores. \n\nI haven’t used power weighting as I already have diversity and power weighting appears to be much better suited to highly correlated groups of ensembles. \n\nI am considering using a rank ensemble as AUC is determined by the correct rank and separation of target values but I need to research this more. I may attempt a weighted rank strategy similar to the above. \n\nGood luck everyone.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 970190,
      "author_name": "fahad7",
      "author_url": "",
      "post_date": "08/14/2020 09:17:59",
      "content": "<p>384th rank is your first kaggle competition is awesome. Hope you get that top 10% for a bronze.</p>",
      "votes": null,
      "replies": [
        {
          "id": 970204,
          "author_name": "jsyphil",
          "author_url": "",
          "post_date": "08/14/2020 09:26:32",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/fahad7\" target=\"_blank\">@fahad7</a>. Three months of around five hours a day, I'll be pleased to be in the top half, which was my original goal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970209,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/14/2020 09:27:58",
      "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> opened a discussion before presenting different methods, you can also see what people recommend or what worked and what did not:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 970216,
          "author_name": "jsyphil",
          "author_url": "",
          "post_date": "08/14/2020 09:30:41",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I'll check that out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970403,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/14/2020 12:31:53",
      "content": "<p>I myself am a newbie here, I am ensembling all my good models just by Overfitting the CV. 😄 Please don't judge me for this, this is the first time I am doing ensembling and just following the quote, \"Always trust your CV\". </p>\n<p>I have made a flowchart for the same and I am using weighted power ensemble with ranking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 970415,
          "author_name": "jsyphil",
          "author_url": "",
          "post_date": "08/14/2020 12:48:37",
          "content": "<p>You and me both <a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a>! I'll check out the weighted power ensemble with ranking technique.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970420,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/14/2020 12:51:42",
          "content": "<p>Good luck with that! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970497,
      "author_name": "jsyphil",
      "author_url": "",
      "post_date": "08/14/2020 14:05:37",
      "content": "<p>I'm currently using Weighted Averaging using relatively arbitrary weights based on my CV scores. Interesting, even applying a small weight to smaller models with lower CV scores gives an overall boost of c. 0.0010. </p>\n<p>I will certainly experiment with Rank Averaging with a great explanation here <a href=\"url\" target=\"_blank\">https://www.kaggle.com/niteshx2/improve-blending-using-rankdata</a> from <a href=\"https://www.kaggle.com/niteshchoudhary\" target=\"_blank\">@niteshchoudhary</a>.</p>\n<p>I'm hesitant with Power Averaging due to correlation issues, I have deliberately created diversity in my models to try and ensure that the weakness in some models are covered by strengths in others and vice-versa. There could be value in performing Power Averaging in small groups of models which have stronger correlation and then using Rank Averaging on these group subsets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 970523,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/14/2020 14:30:23",
          "content": "<p>Yup, it definitely is good, At first I used only 0.90+ CV models and when combined using this, I got a CV score of about 0.92+. It is rocking in that way!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970625,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "08/14/2020 16:06:15",
          "content": "<p>That sounds good, but focus more on the model itself. It could be, that you are overfitting the LB with your tiny adjustments. That means you will have a higher score in public set but later on the private set it could be way lower. Thats why it is suggested to tune your results just slightly and focus on the model itself and your data more. Especially in this competition we have a very imbalanced dataset and the public LB consists of only 30% of the complete test data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970849,
          "author_name": "jsyphil",
          "author_url": "",
          "post_date": "08/14/2020 21:49:46",
          "content": "<p>Resisting the temptation to chase a Public LB score and focus instead on solid models derived from experimentation and stable CV scores with a standard ensemble is very wise advice <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> . That temptation is hard to resist though!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970893,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/15/2020 00:44:08",
          "content": "<p>Have you considered creating a model using the out of fold predictions as input. For example, logistic regression. Or MLP. This is what I do and has worked well for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970987,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/15/2020 04:30:17",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> What I am doing is <code>overfitting</code> the <code>CV</code> scores, isn't it right, :) </p>\n<p>I'm am not even considering the LB score for the ensemble, I'm just taking the CV scores into consideration.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 972082,
      "author_name": "jsyphil",
      "author_url": "",
      "post_date": "08/16/2020 08:06:53",
      "content": "<p>Having spent some time and thought on this over the last two days, I have decided to use a weighted average ensemble with the following strategy. I have deliberately worked on trying to build uncorrelated yet promising models over the last few months. </p>\n<ol>\n<li>Observe and then group together/bucket any highly correlated predictions from different models/inputs.</li>\n<li>Use the cross validation score for each model in each group/bucket and determine an average CV score for each group/bucket</li>\n<li>Build an ensemble assigning a higher weighting to groups/buckets with higher average CV scores.</li>\n<li>Don’t be put off by models or groups of models with lower CV scores (Eg meta only models or smaller CNN networks with smaller dimension inputs) as these will add important diversity to the ensemble. </li>\n<li>Don’t focus on LB scores but do use these as a guide to determine stability of CV scores. </li>\n</ol>\n<p>I haven’t used power weighting as I already have diversity and power weighting appears to be much better suited to highly correlated groups of ensembles. </p>\n<p>I am considering using a rank ensemble as AUC is determined by the correct rank and separation of target values but I need to research this more. I may attempt a weighted rank strategy similar to the above. </p>\n<p>Good luck everyone. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "970179": "With a few days to go and after many many experiments and building dozens of Cross Validation (CV) models over the last three months, I now have nine, what I hope are, solid sets of predictions from nine different models.\n\nFor all nine sets of predictions I have AUC and Loss scores from Out Of Fold (OOF) predictions, Mean CV predictions and fully trained model predictions including training scores and Kaggle Leaderboard (LB) scores.\n\nI have used some trial and error on my ensembles by observing all my AUC scores and also observing which models create good diversity in predictions against other models with similar AUC scores and weighting these accordingly.\n\nCan anyone recommend a more procedural way of building an ensemble in this situation? For example has anyone used Standard Deviation as a way to measure diversity for their ensembles?\n\nI have seen and tried ensemble methods used where we have target values, such as the best OOF ensembles with and without meta data, but not where we don't have target values.\n\nMany thanks and good luck everyone.\n\nps this is my first Kaggle competition and it has been an awesome experience, particularly observing how people share and contribute to ideas, hat tipped to an incredible community here.",
    "970190": "384th rank is your first kaggle competition is awesome. Hope you get that top 10% for a bronze.",
    "970204": "Thanks @fahad7. Three months of around five hours a day, I'll be pleased to be in the top half, which was my original goal.",
    "970209": "sirishks opened a discussion before presenting different methods, you can also see what people recommend or what worked and what did not:\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653",
    "970216": "Thanks @aliabdin1 I'll check that out.",
    "970403": "I myself am a newbie here, I am ensembling all my good models just by Overfitting the CV. 😄 Please don't judge me for this, this is the first time I am doing ensembling and just following the quote, \"Always trust your CV\". \n\nI have made a flowchart for the same and I am using weighted power ensemble with ranking.",
    "970415": "You and me both @sarques! I'll check out the weighted power ensemble with ranking technique.",
    "970420": "Good luck with that! :)",
    "970497": "I'm currently using Weighted Averaging using relatively arbitrary weights based on my CV scores. Interesting, even applying a small weight to smaller models with lower CV scores gives an overall boost of c. 0.0010. \n\nI will certainly experiment with Rank Averaging with a great explanation here [https://www.kaggle.com/niteshx2/improve-blending-using-rankdata](url) from @niteshchoudhary.\n\nI'm hesitant with Power Averaging due to correlation issues, I have deliberately created diversity in my models to try and ensure that the weakness in some models are covered by strengths in others and vice-versa. There could be value in performing Power Averaging in small groups of models which have stronger correlation and then using Rank Averaging on these group subsets.",
    "970523": "Yup, it definitely is good, At first I used only 0.90+ CV models and when combined using this, I got a CV score of about 0.92+. It is rocking in that way!",
    "970625": "That sounds good, but focus more on the model itself. It could be, that you are overfitting the LB with your tiny adjustments. That means you will have a higher score in public set but later on the private set it could be way lower. Thats why it is suggested to tune your results just slightly and focus on the model itself and your data more. Especially in this competition we have a very imbalanced dataset and the public LB consists of only 30% of the complete test data.",
    "970849": "Resisting the temptation to chase a Public LB score and focus instead on solid models derived from experimentation and stable CV scores with a standard ensemble is very wise advice @aliabdin1 . That temptation is hard to resist though!",
    "970893": "Have you considered creating a model using the out of fold predictions as input. For example, logistic regression. Or MLP. This is what I do and has worked well for me.",
    "970987": "Hey @aliabdin1 What I am doing is `overfitting` the `CV` scores, isn't it right, :) \n\nI'm am not even considering the LB score for the ensemble, I'm just taking the CV scores into consideration.",
    "972082": "Having spent some time and thought on this over the last two days, I have decided to use a weighted average ensemble with the following strategy. I have deliberately worked on trying to build uncorrelated yet promising models over the last few months. \n\n1. Observe and then group together/bucket any highly correlated predictions from different models/inputs.\n2. Use the cross validation score for each model in each group/bucket and determine an average CV score for each group/bucket\n3. Build an ensemble assigning a higher weighting to groups/buckets with higher average CV scores.\n4. Don’t be put off by models or groups of models with lower CV scores (Eg meta only models or smaller CNN networks with smaller dimension inputs) as these will add important diversity to the ensemble. \n5. Don’t focus on LB scores but do use these as a guide to determine stability of CV scores. \n\nI haven’t used power weighting as I already have diversity and power weighting appears to be much better suited to highly correlated groups of ensembles. \n\nI am considering using a rank ensemble as AUC is determined by the correct rank and separation of target values but I need to research this more. I may attempt a weighted rank strategy similar to the above. \n\nGood luck everyone."
  },
  "source": "meta"
}