{
  "id": 115641,
  "title": "multi model ensemble questions",
  "url": "/competitions/understanding_cloud_organization/discussion/115641",
  "author_name": "liuze",
  "post_date": "2019-11-04T10:20:42.245000",
  "votes": 6,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I am confused with model ensemble, I trained different model use great kernel [https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools]\nand most model got 0.645+, if I want to ensemble the result, should I average each model's logit output? or the union of mask?</p>",
  "messages": [
    {
      "id": 665147,
      "postDate": "2019-11-04T17:15:21.590Z",
      "content": "<p>I think that there are three main approaches:\n- as Chris said - averaging of raw predictions;\n- if you train models with the same architecture, you could average their weights in all layers;\n- predict masks by all models and make some kind of combination - maybe take all pixels in all masks or pixels which are masks in more than a half of the models;</p>",
      "rawMarkdown": "I think that there are three main approaches:\n- as Chris said - averaging of raw predictions;\n- if you train models with the same architecture, you could average their weights in all layers;\n- predict masks by all models and make some kind of combination - maybe take all pixels in all masks or pixels which are masks in more than a half of the models;",
      "votes": 7,
      "replies": [
        {
          "id": 665182,
          "postDate": "2019-11-04T17:47:25.363Z",
          "content": "<p>Approach 2 is interesting. I've never tried that. I guess if you do 5-Folds, you can convert the five separate models into one model for quicker inference. How does the  performance of that single model compare with averaging raw predictions of the five models?</p>",
          "rawMarkdown": "Approach 2 is interesting. I've never tried that. I guess if you do 5-Folds, you can convert the five separate models into one model for quicker inference. How does the  performance of that single model compare with averaging raw predictions of the five models?",
          "votes": 1
        },
        {
          "id": 665189,
          "postDate": "2019-11-04T18:05:08.140Z",
          "content": "<p>I think it depends on the situation.\nIn the writeup of the 12th place in severstal competition it worked:\n<a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114309#latest-661404\">https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114309#latest-661404</a>\n<code>\nTrick that helps to improve models quality for about 0.3-0.5 % points for all models - best checkpoints weights average (weights!, not predictions). During traning 5 best checkpoints have been saved and then differnet combinations of them evaluated to find the best candidates to average\n</code></p>",
          "rawMarkdown": "I think it depends on the situation.\nIn the writeup of the 12th place in severstal competition it worked:\nhttps://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114309#latest-661404\n```\nTrick that helps to improve models quality for about 0.3-0.5 % points for all models - best checkpoints weights average (weights!, not predictions). During traning 5 best checkpoints have been saved and then differnet combinations of them evaluated to find the best candidates to average\n```",
          "votes": 2
        },
        {
          "id": 665191,
          "postDate": "2019-11-04T18:06:19.533Z",
          "content": "<p>@Andrew, do you know if catalyst or pytorch-toolbelt or SMP allows easy weight averaging or does one have to parse weights?</p>",
          "rawMarkdown": "@Andrew, do you know if catalyst or pytorch-toolbelt or SMP allows easy weight averaging or does one have to parse weights?"
        },
        {
          "id": 665331,
          "postDate": "2019-11-04T22:11:33.680Z",
          "content": "<p>SWA(stochastic weight averaging)? I have never tried it. </p>",
          "rawMarkdown": "SWA(stochastic weight averaging)? I have never tried it. "
        },
        {
          "id": 665414,
          "postDate": "2019-11-05T01:34:45.067Z",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a> hi cao, you are the first now, keep it. don't overfit this time :)</p>",
          "rawMarkdown": "@naivelamb hi cao, you are the first now, keep it. don't overfit this time :)"
        },
        {
          "id": 665431,
          "postDate": "2019-11-05T02:26:00.597Z",
          "content": "<p><a href=\"/pheadrus\">@pheadrus</a> sadly I don't know about easy way to do it.</p>",
          "rawMarkdown": "@pheadrus sadly I don't know about easy way to do it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 664868,
      "postDate": "2019-11-04T10:20:42.247Z",
      "content": "<p>I am confused with model ensemble, I trained different model use great kernel [https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools]\nand most model got 0.645+, if I want to ensemble the result, should I average each model's logit output? or the union of mask?</p>",
      "rawMarkdown": "I am confused with model ensemble, I trained different model use great kernel [https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools]\nand most model got 0.645+, if I want to ensemble the result, should I average each model's logit output? or the union of mask?",
      "votes": 6
    },
    {
      "id": 665139,
      "postDate": "2019-11-04T17:09:08.773Z",
      "content": "<p>I think it's best to ensemble the masks before converting into <code>np.uint8</code>, i.e. use the probabilities where <code>mask_after = (mask_probabilities&gt;0.4).astype(np.uint8)</code>.  </p>\n\n<p>You can use simple average like <code>mask_new = (mask_1 + mask_2)/2.0</code> of you can do temperature shaping <code>mask_new = (mask_1**0.5 + mask_2**0.5)/2.0</code> explained by Heng <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/107716\">here</a></p>",
      "rawMarkdown": "I think it's best to ensemble the masks before converting into `np.uint8`, i.e. use the probabilities where `mask_after = (mask_probabilities&gt;0.4).astype(np.uint8)`.  \n  \nYou can use simple average like `mask_new = (mask_1 + mask_2)/2.0` of you can do temperature shaping `mask_new = (mask_1**0.5 + mask_2**0.5)/2.0` explained by Heng [here][1]\n\n[1]: https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/107716",
      "votes": 3,
      "replies": [
        {
          "id": 665337,
          "postDate": "2019-11-04T22:22:31.980Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , have you tried temperature ensembling here? I tried once and got weird results</p>",
          "rawMarkdown": "@cdeotte , have you tried temperature ensembling here? I tried once and got weird results"
        },
        {
          "id": 665376,
          "postDate": "2019-11-05T00:27:39.663Z",
          "content": "<p>I have not tried it yet</p>",
          "rawMarkdown": "I have not tried it yet"
        },
        {
          "id": 665413,
          "postDate": "2019-11-05T01:33:40.620Z",
          "content": "<p>good idea, I would try, as <a href=\"/raspstephan\">@raspstephan</a> said, [https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587#601668]\nthe labels are merge from crowds, so maybe merge different model's output maybe match the label</p>",
          "rawMarkdown": "good idea, I would try, as @raspstephan said, [https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587#601668]\nthe labels are merge from crowds, so maybe merge different model's output maybe match the label"
        },
        {
          "id": 666541,
          "postDate": "2019-11-06T07:47:25.440Z",
          "content": "<p>Hi Chris!I'm confuesd if we should apply sigmoid function on our mask predictions when doing ensembles on our model,coz I found this in <a href=\"https://www.kaggle.com/khornlund/sever-ensemble-classification\">carl's solution</a> in steel detection,thx a lot</p>",
          "rawMarkdown": "Hi Chris!I'm confuesd if we should apply sigmoid function on our mask predictions when doing ensembles on our model,coz I found this in [carl's solution](https://www.kaggle.com/khornlund/sever-ensemble-classification) in steel detection,thx a lot"
        },
        {
          "id": 666934,
          "postDate": "2019-11-06T16:29:01.097Z",
          "content": "<p>You can apply any function you want in the formula: <code>mask_new = ( func(mask_1) + func(mask_2) )/2.0</code>. Basic average uses <code>func(x) = x</code> and temperature shaping uses <code>func(x) = x**0.5</code>. You could use sigmoid as in <code>func(x) = sigmoid(x)</code>. Afterward, pick a threshold and <code>final_mask = (mask_new&gt;threshold).astype(uint8)</code></p>",
          "rawMarkdown": "You can apply any function you want in the formula: `mask_new = ( func(mask_1) + func(mask_2) )/2.0`. Basic average uses `func(x) = x` and temperature shaping uses `func(x) = x**0.5`. You could use sigmoid as in `func(x) = sigmoid(x)`. Afterward, pick a threshold and `final_mask = (mask_new&gt;threshold).astype(uint8)`"
        },
        {
          "id": 668193,
          "postDate": "2019-11-08T05:24:14.417Z",
          "content": "<p>Oh,Chris,that's so nice of you,BTW,thx a lot</p>",
          "rawMarkdown": "Oh,Chris,that's so nice of you,BTW,thx a lot"
        }
      ]
    },
    {
      "id": 665434,
      "postDate": "2019-11-05T02:28:17.697Z",
      "content": "<p>I use voting on predicted masks, i.e. 1 1 0 of three models on pixel (0,0), threshold&gt;=2, voting result 1. This gives boost on LB up to .005, not yet compared with other methods...</p>",
      "rawMarkdown": "I use voting on predicted masks, i.e. 1 1 0 of three models on pixel (0,0), threshold&gt;=2, voting result 1. This gives boost on LB up to .005, not yet compared with other methods...",
      "votes": 4,
      "replies": [
        {
          "id": 668145,
          "postDate": "2019-11-08T02:58:25.943Z",
          "content": "<p>oh, 0.005 boost !! how many modes are you voting</p>",
          "rawMarkdown": "oh, 0.005 boost !! how many modes are you voting"
        },
        {
          "id": 668198,
          "postDate": "2019-11-08T05:32:13.290Z",
          "content": "<p>3\nsubmit 1: model1=.658, model2=.656, model3=.656, ensemble=.663\nsubmit 2: model1=.661, model2=.658, model3=.658, ensemble=.659</p>",
          "rawMarkdown": "3\nsubmit 1: model1=.658, model2=.656, model3=.656, ensemble=.663\nsubmit 2: model1=.661, model2=.658, model3=.658, ensemble=.659"
        },
        {
          "id": 668218,
          "postDate": "2019-11-08T06:10:35.017Z",
          "content": "<p>Wow,very impressive on getting your single model getting.661,I think you can do better ensemble based on this best single model like a .666 in total maybe?</p>",
          "rawMarkdown": "Wow,very impressive on getting your single model getting.661,I think you can do better ensemble based on this best single model like a .666 in total maybe?"
        },
        {
          "id": 668236,
          "postDate": "2019-11-08T06:28:33.003Z",
          "content": "<p>Is this single fold models or multi folds combined? For the individual model numbers</p>",
          "rawMarkdown": "Is this single fold models or multi folds combined? For the individual model numbers"
        },
        {
          "id": 668266,
          "postDate": "2019-11-08T07:43:19.833Z",
          "content": "<p><code>submit 2: model1=.661, model2=.658, model3=.658, ensemble=.659</code>\nin this submit, ensemble score is decreased?</p>",
          "rawMarkdown": "```submit 2: model1=.661, model2=.658, model3=.658, ensemble=.659```\nin this submit, ensemble score is decreased?"
        },
        {
          "id": 668286,
          "postDate": "2019-11-08T08:14:39.370Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 668379,
          "postDate": "2019-11-08T11:13:10.267Z",
          "content": "<p>yes, threshold&gt;=2, maybe two weaker models accidentally beat the stronger one</p>",
          "rawMarkdown": "yes, threshold&gt;=2, maybe two weaker models accidentally beat the stronger one"
        },
        {
          "id": 668411,
          "postDate": "2019-11-08T12:02:13.823Z",
          "content": "<p>Right now,My best single model is .654,but ater ensembling four models and applying classifier to my result,I got 661 for now ,I wonder could it be overfitting the public data?</p>",
          "rawMarkdown": "Right now,My best single model is .654,but ater ensembling four models and applying classifier to my result,I got 661 for now ,I wonder could it be overfitting the public data?"
        },
        {
          "id": 668457,
          "postDate": "2019-11-08T12:59:46.157Z",
          "content": "<p>I have a lot of folds scoring .650~.660, ensemble best score .663\nand I see in forum best single fold (segment only) got .664</p>",
          "rawMarkdown": "I have a lot of folds scoring .650~.660, ensemble best score .663\nand I see in forum best single fold (segment only) got .664"
        }
      ]
    },
    {
      "id": 664922,
      "postDate": "2019-11-04T12:10:33Z",
      "content": "<p>I think you need to average the probabilities and not the logits. </p>",
      "rawMarkdown": "I think you need to average the probabilities and not the logits. "
    }
  ],
  "comments": [
    {
      "id": 665147,
      "author_name": "Andrey Lukyanenko",
      "author_url": "",
      "post_date": "2019-11-04T17:15:21.590000",
      "content": "<p>I think that there are three main approaches:\n- as Chris said - averaging of raw predictions;\n- if you train models with the same architecture, you could average their weights in all layers;\n- predict masks by all models and make some kind of combination - maybe take all pixels in all masks or pixels which are masks in more than a half of the models;</p>",
      "votes": 7,
      "replies": [
        {
          "id": 665182,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-04T17:47:25.363000",
          "content": "<p>Approach 2 is interesting. I've never tried that. I guess if you do 5-Folds, you can convert the five separate models into one model for quicker inference. How does the  performance of that single model compare with averaging raw predictions of the five models?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665189,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-11-04T18:05:08.140000",
          "content": "<p>I think it depends on the situation.\nIn the writeup of the 12th place in severstal competition it worked:\n<a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114309#latest-661404\">https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114309#latest-661404</a>\n<code>\nTrick that helps to improve models quality for about 0.3-0.5 % points for all models - best checkpoints weights average (weights!, not predictions). During traning 5 best checkpoints have been saved and then differnet combinations of them evaluated to find the best candidates to average\n</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665191,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2019-11-04T18:06:19.533000",
          "content": "<p>@Andrew, do you know if catalyst or pytorch-toolbelt or SMP allows easy weight averaging or does one have to parse weights?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665331,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-04T22:11:33.680000",
          "content": "<p>SWA(stochastic weight averaging)? I have never tried it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665414,
          "author_name": "liuze",
          "author_url": "",
          "post_date": "2019-11-05T01:34:45.067000",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a> hi cao, you are the first now, keep it. don't overfit this time :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665431,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-11-05T02:26:00.597000",
          "content": "<p><a href=\"/pheadrus\">@pheadrus</a> sadly I don't know about easy way to do it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 665139,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-04T17:09:08.773000",
      "content": "<p>I think it's best to ensemble the masks before converting into <code>np.uint8</code>, i.e. use the probabilities where <code>mask_after = (mask_probabilities&gt;0.4).astype(np.uint8)</code>.  </p>\n\n<p>You can use simple average like <code>mask_new = (mask_1 + mask_2)/2.0</code> of you can do temperature shaping <code>mask_new = (mask_1**0.5 + mask_2**0.5)/2.0</code> explained by Heng <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/107716\">here</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 665337,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-11-04T22:22:31.980000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , have you tried temperature ensembling here? I tried once and got weird results</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665376,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-05T00:27:39.663000",
          "content": "<p>I have not tried it yet</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665413,
          "author_name": "liuze",
          "author_url": "",
          "post_date": "2019-11-05T01:33:40.620000",
          "content": "<p>good idea, I would try, as <a href=\"/raspstephan\">@raspstephan</a> said, [https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587#601668]\nthe labels are merge from crowds, so maybe merge different model's output maybe match the label</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666541,
          "author_name": "zhou dream",
          "author_url": "",
          "post_date": "2019-11-06T07:47:25.440000",
          "content": "<p>Hi Chris!I'm confuesd if we should apply sigmoid function on our mask predictions when doing ensembles on our model,coz I found this in <a href=\"https://www.kaggle.com/khornlund/sever-ensemble-classification\">carl's solution</a> in steel detection,thx a lot</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666934,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-06T16:29:01.097000",
          "content": "<p>You can apply any function you want in the formula: <code>mask_new = ( func(mask_1) + func(mask_2) )/2.0</code>. Basic average uses <code>func(x) = x</code> and temperature shaping uses <code>func(x) = x**0.5</code>. You could use sigmoid as in <code>func(x) = sigmoid(x)</code>. Afterward, pick a threshold and <code>final_mask = (mask_new&gt;threshold).astype(uint8)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668193,
          "author_name": "zhou dream",
          "author_url": "",
          "post_date": "2019-11-08T05:24:14.417000",
          "content": "<p>Oh,Chris,that's so nice of you,BTW,thx a lot</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 665434,
      "author_name": "Endi Niu",
      "author_url": "",
      "post_date": "2019-11-05T02:28:17.697000",
      "content": "<p>I use voting on predicted masks, i.e. 1 1 0 of three models on pixel (0,0), threshold&gt;=2, voting result 1. This gives boost on LB up to .005, not yet compared with other methods...</p>",
      "votes": 4,
      "replies": [
        {
          "id": 668145,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-11-08T02:58:25.943000",
          "content": "<p>oh, 0.005 boost !! how many modes are you voting</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668198,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-08T05:32:13.290000",
          "content": "<p>3\nsubmit 1: model1=.658, model2=.656, model3=.656, ensemble=.663\nsubmit 2: model1=.661, model2=.658, model3=.658, ensemble=.659</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668218,
          "author_name": "Overfit Queen",
          "author_url": "",
          "post_date": "2019-11-08T06:10:35.017000",
          "content": "<p>Wow,very impressive on getting your single model getting.661,I think you can do better ensemble based on this best single model like a .666 in total maybe?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668236,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-11-08T06:28:33.003000",
          "content": "<p>Is this single fold models or multi folds combined? For the individual model numbers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668266,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-11-08T07:43:19.833000",
          "content": "<p><code>submit 2: model1=.661, model2=.658, model3=.658, ensemble=.659</code>\nin this submit, ensemble score is decreased?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668286,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-08T08:14:39.370000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668379,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-08T11:13:10.267000",
          "content": "<p>yes, threshold&gt;=2, maybe two weaker models accidentally beat the stronger one</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668411,
          "author_name": "Overfit Queen",
          "author_url": "",
          "post_date": "2019-11-08T12:02:13.823000",
          "content": "<p>Right now,My best single model is .654,but ater ensembling four models and applying classifier to my result,I got 661 for now ,I wonder could it be overfitting the public data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668457,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-08T12:59:46.157000",
          "content": "<p>I have a lot of folds scoring .650~.660, ensemble best score .663\nand I see in forum best single fold (segment only) got .664</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 664922,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2019-11-04T12:10:33",
      "content": "<p>I think you need to average the probabilities and not the logits. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "665147": "I think that there are three main approaches:\n- as Chris said - averaging of raw predictions;\n- if you train models with the same architecture, you could average their weights in all layers;\n- predict masks by all models and make some kind of combination - maybe take all pixels in all masks or pixels which are masks in more than a half of the models;",
    "664868": "I am confused with model ensemble, I trained different model use great kernel [https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools]\nand most model got 0.645+, if I want to ensemble the result, should I average each model's logit output? or the union of mask?",
    "665139": "I think it's best to ensemble the masks before converting into `np.uint8`, i.e. use the probabilities where `mask_after = (mask_probabilities&gt;0.4).astype(np.uint8)`.  \n  \nYou can use simple average like `mask_new = (mask_1 + mask_2)/2.0` of you can do temperature shaping `mask_new = (mask_1**0.5 + mask_2**0.5)/2.0` explained by Heng [here][1]\n\n[1]: https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/107716",
    "665434": "I use voting on predicted masks, i.e. 1 1 0 of three models on pixel (0,0), threshold&gt;=2, voting result 1. This gives boost on LB up to .005, not yet compared with other methods...",
    "664922": "I think you need to average the probabilities and not the logits. "
  }
}