{
  "id": 47507,
  "title": "How to choose weights for ensembling models",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47507",
  "author_name": "",
  "post_date": "2018-01-15T17:17:58.232338700Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I have nine individual models with LB scored 85, 85, 86,  87, 87, 88, 88, 88, 88, I am skipping models which gave LB scores less than 85. I have tried ensembling these models but I am not seeing any good results. I have tried with sqrt(prob) and majority voting but it gave me only few position jumps.\nI am new to kaggle competition can you give me some suggestions how to choose weights or other ensembling methods that I can use. Note that the models are not very uncorrelated, these models are trained with different data augmentation techniques and and pseudo labeling files in test set</p>",
  "messages": [
    {
      "id": "268797",
      "postDate": "01/15/2018 17:17:58",
      "content": "<p>I have nine individual models with LB scored 85, 85, 86,  87, 87, 88, 88, 88, 88, I am skipping models which gave LB scores less than 85. I have tried ensembling these models but I am not seeing any good results. I have tried with sqrt(prob) and majority voting but it gave me only few position jumps.\nI am new to kaggle competition can you give me some suggestions how to choose weights or other ensembling methods that I can use. Note that the models are not very uncorrelated, these models are trained with different data augmentation techniques and and pseudo labeling files in test set</p>",
      "rawMarkdown": "I have nine individual models with LB scored 85, 85, 86,  87, 87, 88, 88, 88, 88, I am skipping models which gave LB scores less than 85. I have tried ensembling these models but I am not seeing any good results. I have tried with sqrt(prob) and majority voting but it gave me only few position jumps.\nI am new to kaggle competition can you give me some suggestions how to choose weights or other ensembling methods that I can use. Note that the models are not very uncorrelated, these models are trained with different data augmentation techniques and and pseudo labeling files in test set",
      "votes": null
    },
    {
      "id": "268804",
      "postDate": "01/15/2018 17:42:58",
      "content": "<p>same here, this is also our first time and we have no idea how to ensemble. We had around 15 models &gt;= 87% LB which was enough to get a place around 30 with majority voting.</p>",
      "rawMarkdown": "same here, this is also our first time and we have no idea how to ensemble. We had around 15 models &gt;= 87% LB which was enough to get a place around 30 with majority voting.",
      "votes": null
    },
    {
      "id": "268809",
      "postDate": "01/15/2018 17:51:37",
      "content": "<p>some great idea in here, if you have time to implement them ;) \n<a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>",
      "rawMarkdown": "some great idea in here, if you have time to implement them ;) \nhttp://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/",
      "votes": null
    },
    {
      "id": "268811",
      "postDate": "01/15/2018 17:54:00",
      "content": "<p>will see, very less time though</p>",
      "rawMarkdown": "will see, very less time though",
      "votes": null
    },
    {
      "id": "268812",
      "postDate": "01/15/2018 17:55:06",
      "content": "<p>How correlated your models are?</p>",
      "rawMarkdown": "How correlated your models are?",
      "votes": null
    },
    {
      "id": "268822",
      "postDate": "01/15/2018 18:13:34",
      "content": "<p>Pearson's correlation scores are between 0.88 and 0.93.</p>",
      "rawMarkdown": "Pearson's correlation scores are between 0.88 and 0.93.",
      "votes": null
    },
    {
      "id": "268886",
      "postDate": "01/15/2018 22:03:15",
      "content": "<p>actually, this is almost a trial and error things. It is in fact the key of winning kaggle competitions. And I strongly suggest you try different weight combinations after the competition the see the effects on public and private LB by making post-competition submissions. </p>\n\n<p>Here are my suggestions:</p>\n\n<p>[\"correct\" way to do]</p>\n\n<ul>\n<li><p>have a validation set. This means that at the start of competition, you should have already planned for ensemble and decide how to split the dataset for train/valid the model and train/valid the ensemble</p></li>\n<li><p>for me, in this competition, i kept a \"clean validation set\" that  is never used in training model or ensemble. I synthesize more validation samples by more extreme augmentation to test the limits.  Especially, i synthesize more silence and unknown in validation (e.g. random mix unknown words by superimpose or cut and paste)</p></li>\n<li><p>make few baseline ensemble, just average and majority vote. It is important to do this because you will  be using this to compare different weighing method.  </p></li>\n<li><p>most of the LB samples are \"stable easily correct\" and do not change that much with different weighing. You may want to identify a subset of LB samples that are \"unstable\" and observe how their label changes with weighing changes. </p></li>\n</ul>\n\n<p>[post-competition]</p>\n\n<p>you can think of it like this (greedy approach):</p>\n\n<ul>\n<li><p>start with ensemble = 0  (empty ). </p></li>\n<li><p>then, ensemble  = model with the highest LB score . </p></li>\n<li><p>if i make ensemble  +=  a_k*model_k, results is going to change. How to describe the change?  You have to think of some \"observed features\". E.g.  let change =  {f1, f2, f3 ... } , where f1 = % change in labels whose original confidence = 0.5 to 0.6, where f2 = % of label is is silence, f3 ...</p></li>\n<li><p>now, based on these \"observed features\", can we predict if the LB scores can improve or not? If you have enough submission slots (e.g. post-competition), you can collect data to \"train\" a LB score predictor.</p></li>\n</ul>\n\n<p>in this previous competition, i use % change of labels to judge if ensemble would work better or not. (but it is not so accurate in this speech competition)</p>",
      "rawMarkdown": "actually, this is almost a trial and error things. It is in fact the key of winning kaggle competitions. And I strongly suggest you try different weight combinations after the competition the see the effects on public and private LB by making post-competition submissions. \n\nHere are my suggestions:\n\n[\"correct\" way to do]\n\n- have a validation set. This means that at the start of competition, you should have already planned for ensemble and decide how to split the dataset for train/valid the model and train/valid the ensemble\n\n- for me, in this competition, i kept a \"clean validation set\" that  is never used in training model or ensemble. I synthesize more validation samples by more extreme augmentation to test the limits.  Especially, i synthesize more silence and unknown in validation (e.g. random mix unknown words by superimpose or cut and paste)\n\n- make few baseline ensemble, just average and majority vote. It is important to do this because you will  be using this to compare different weighing method.  \n\n- most of the LB samples are \"stable easily correct\" and do not change that much with different weighing. You may want to identify a subset of LB samples that are \"unstable\" and observe how their label changes with weighing changes. \n\n[post-competition]\n\nyou can think of it like this (greedy approach):\n\n\n- start with ensemble = 0  (empty ). \n\n- then, ensemble  = model with the highest LB score . \n\n- if i make ensemble  +=  a_k*model_k, results is going to change. How to describe the change?  You have to think of some \"observed features\". E.g.  let change =  {f1, f2, f3 ... } , where f1 = % change in labels whose original confidence = 0.5 to 0.6, where f2 = % of label is is silence, f3 ...\n\n- now, based on these \"observed features\", can we predict if the LB scores can improve or not? If you have enough submission slots (e.g. post-competition), you can collect data to \"train\" a LB score predictor.\n\nin this previous competition, i use % change of labels to judge if ensemble would work better or not. (but it is not so accurate in this speech competition)",
      "votes": null
    },
    {
      "id": "269013",
      "postDate": "01/16/2018 03:55:16",
      "content": "<p>Thanks Heng, for the insight, left with few submissions to try out all these, certainly it will help me in future competition.</p>",
      "rawMarkdown": "Thanks Heng, for the insight, left with few submissions to try out all these, certainly it will help me in future competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 268804,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "01/15/2018 17:42:58",
      "content": "<p>same here, this is also our first time and we have no idea how to ensemble. We had around 15 models &gt;= 87% LB which was enough to get a place around 30 with majority voting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 268812,
          "author_name": "princerk",
          "author_url": "",
          "post_date": "01/15/2018 17:55:06",
          "content": "<p>How correlated your models are?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268822,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "01/15/2018 18:13:34",
          "content": "<p>Pearson's correlation scores are between 0.88 and 0.93.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268886,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/15/2018 22:03:15",
          "content": "<p>actually, this is almost a trial and error things. It is in fact the key of winning kaggle competitions. And I strongly suggest you try different weight combinations after the competition the see the effects on public and private LB by making post-competition submissions. </p>\n\n<p>Here are my suggestions:</p>\n\n<p>[\"correct\" way to do]</p>\n\n<ul>\n<li><p>have a validation set. This means that at the start of competition, you should have already planned for ensemble and decide how to split the dataset for train/valid the model and train/valid the ensemble</p></li>\n<li><p>for me, in this competition, i kept a \"clean validation set\" that  is never used in training model or ensemble. I synthesize more validation samples by more extreme augmentation to test the limits.  Especially, i synthesize more silence and unknown in validation (e.g. random mix unknown words by superimpose or cut and paste)</p></li>\n<li><p>make few baseline ensemble, just average and majority vote. It is important to do this because you will  be using this to compare different weighing method.  </p></li>\n<li><p>most of the LB samples are \"stable easily correct\" and do not change that much with different weighing. You may want to identify a subset of LB samples that are \"unstable\" and observe how their label changes with weighing changes. </p></li>\n</ul>\n\n<p>[post-competition]</p>\n\n<p>you can think of it like this (greedy approach):</p>\n\n<ul>\n<li><p>start with ensemble = 0  (empty ). </p></li>\n<li><p>then, ensemble  = model with the highest LB score . </p></li>\n<li><p>if i make ensemble  +=  a_k*model_k, results is going to change. How to describe the change?  You have to think of some \"observed features\". E.g.  let change =  {f1, f2, f3 ... } , where f1 = % change in labels whose original confidence = 0.5 to 0.6, where f2 = % of label is is silence, f3 ...</p></li>\n<li><p>now, based on these \"observed features\", can we predict if the LB scores can improve or not? If you have enough submission slots (e.g. post-competition), you can collect data to \"train\" a LB score predictor.</p></li>\n</ul>\n\n<p>in this previous competition, i use % change of labels to judge if ensemble would work better or not. (but it is not so accurate in this speech competition)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269013,
          "author_name": "princerk",
          "author_url": "",
          "post_date": "01/16/2018 03:55:16",
          "content": "<p>Thanks Heng, for the insight, left with few submissions to try out all these, certainly it will help me in future competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 268809,
      "author_name": "liamsch",
      "author_url": "",
      "post_date": "01/15/2018 17:51:37",
      "content": "<p>some great idea in here, if you have time to implement them ;) \n<a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 268811,
          "author_name": "princerk",
          "author_url": "",
          "post_date": "01/15/2018 17:54:00",
          "content": "<p>will see, very less time though</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "268797": "I have nine individual models with LB scored 85, 85, 86,  87, 87, 88, 88, 88, 88, I am skipping models which gave LB scores less than 85. I have tried ensembling these models but I am not seeing any good results. I have tried with sqrt(prob) and majority voting but it gave me only few position jumps.\nI am new to kaggle competition can you give me some suggestions how to choose weights or other ensembling methods that I can use. Note that the models are not very uncorrelated, these models are trained with different data augmentation techniques and and pseudo labeling files in test set",
    "268804": "same here, this is also our first time and we have no idea how to ensemble. We had around 15 models &gt;= 87% LB which was enough to get a place around 30 with majority voting.",
    "268809": "some great idea in here, if you have time to implement them ;) \nhttp://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/",
    "268811": "will see, very less time though",
    "268812": "How correlated your models are?",
    "268822": "Pearson's correlation scores are between 0.88 and 0.93.",
    "268886": "actually, this is almost a trial and error things. It is in fact the key of winning kaggle competitions. And I strongly suggest you try different weight combinations after the competition the see the effects on public and private LB by making post-competition submissions. \n\nHere are my suggestions:\n\n[\"correct\" way to do]\n\n- have a validation set. This means that at the start of competition, you should have already planned for ensemble and decide how to split the dataset for train/valid the model and train/valid the ensemble\n\n- for me, in this competition, i kept a \"clean validation set\" that  is never used in training model or ensemble. I synthesize more validation samples by more extreme augmentation to test the limits.  Especially, i synthesize more silence and unknown in validation (e.g. random mix unknown words by superimpose or cut and paste)\n\n- make few baseline ensemble, just average and majority vote. It is important to do this because you will  be using this to compare different weighing method.  \n\n- most of the LB samples are \"stable easily correct\" and do not change that much with different weighing. You may want to identify a subset of LB samples that are \"unstable\" and observe how their label changes with weighing changes. \n\n[post-competition]\n\nyou can think of it like this (greedy approach):\n\n\n- start with ensemble = 0  (empty ). \n\n- then, ensemble  = model with the highest LB score . \n\n- if i make ensemble  +=  a_k*model_k, results is going to change. How to describe the change?  You have to think of some \"observed features\". E.g.  let change =  {f1, f2, f3 ... } , where f1 = % change in labels whose original confidence = 0.5 to 0.6, where f2 = % of label is is silence, f3 ...\n\n- now, based on these \"observed features\", can we predict if the LB scores can improve or not? If you have enough submission slots (e.g. post-competition), you can collect data to \"train\" a LB score predictor.\n\nin this previous competition, i use % change of labels to judge if ensemble would work better or not. (but it is not so accurate in this speech competition)",
    "269013": "Thanks Heng, for the insight, left with few submissions to try out all these, certainly it will help me in future competition."
  },
  "source": "meta"
}