{
  "id": 39383,
  "title": "How to ensemble my several models?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39383",
  "author_name": "",
  "post_date": "2017-09-13T03:58:36.864305200Z",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have two models which respectively trained by 1152x1152,1024x1024 . I use <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523\">Peter's codes</a> to train these models . And I have only a gtx1060 gpu, it is very slow to generate the submit file(about 9 hours for 1152x1152 model, and about 6 hours for 1024x1024 model)  . I don't know how to ensemble them and generate the submit file.  Any suggestion?  </p>",
  "messages": [
    {
      "id": "220788",
      "postDate": "09/13/2017 03:58:36",
      "content": "<p>I have two models which respectively trained by 1152x1152,1024x1024 . I use <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523\">Peter's codes</a> to train these models . And I have only a gtx1060 gpu, it is very slow to generate the submit file(about 9 hours for 1152x1152 model, and about 6 hours for 1024x1024 model)  . I don't know how to ensemble them and generate the submit file.  Any suggestion?  </p>",
      "rawMarkdown": "I have two models which respectively trained by 1152x1152,1024x1024 . I use [Peter's codes][1] to train these models . And I have only a gtx1060 gpu, it is very slow to generate the submit file(about 9 hours for 1152x1152 model, and about 6 hours for 1024x1024 model)  . I don't know how to ensemble them and generate the submit file.  Any suggestion?  \n\n\n  [1]: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523",
      "votes": null
    },
    {
      "id": "220811",
      "postDate": "09/13/2017 06:24:39",
      "content": "<p>The easiest way to do an ensemble is to create predictions for both models. Then average the predictions and create a new submission using the average predictions.</p>",
      "rawMarkdown": "The easiest way to do an ensemble is to create predictions for both models. Then average the predictions and create a new submission using the average predictions.",
      "votes": null
    },
    {
      "id": "220853",
      "postDate": "09/13/2017 09:41:06",
      "content": "<p>Thanks.  Since I already had two submit files , maybe I can try to read these files and get the  average predictions.</p>",
      "rawMarkdown": "Thanks.  Since I already had two submit files , maybe I can try to read these files and get the  average predictions.",
      "votes": null
    },
    {
      "id": "220871",
      "postDate": "09/13/2017 10:51:13",
      "content": "<p>I haven't really done ensembles before, but I assumed that one should average out the outputs from the two (or more) networks before the soft_max operation?</p>",
      "rawMarkdown": "I haven't really done ensembles before, but I assumed that one should average out the outputs from the two (or more) networks before the soft_max operation?",
      "votes": null
    },
    {
      "id": "221012",
      "postDate": "09/13/2017 19:10:25",
      "content": "<p>Im not sure about that, but in this competition you dont need softmax.</p>",
      "rawMarkdown": "Im not sure about that, but in this competition you dont need softmax.",
      "votes": null
    },
    {
      "id": "221113",
      "postDate": "09/14/2017 06:20:51",
      "content": "<p>average gets 0.0002 improvement of my LB score</p>",
      "rawMarkdown": "average gets 0.0002 improvement of my LB score",
      "votes": null
    },
    {
      "id": "222280",
      "postDate": "09/18/2017 09:44:56",
      "content": "<p>any improvement with the ensemble ?</p>",
      "rawMarkdown": "any improvement with the ensemble ?",
      "votes": null
    },
    {
      "id": "222336",
      "postDate": "09/18/2017 13:51:53",
      "content": "<p>If the models have the same size you can try concatenating the outputs before the final convolution (usually a sigmoid/softmax in unet) and see if that works. Haven't had much success for my unets but I ran limited trials. Also if you're limited in memory this may hurt because you're running parallel models and joining the outputs. There is also the cost of testing on a larger network... There are other options, searching for somthing like 'keras join models', 'keras stacking models' or whatnot, depending on your framework may be a good read.</p>\n\n<p>Averaging out the outputs is the easiest to implement though, you just pick your best models and, well, average the output, i.e., the predictions before thresholding. More guaranteed to give you some gains too. You can also try voting. Haven't done either for lack of time but if you have a couple of models with good/similar scores it's worth a try. If one model is better then the other try weighing it in favor of the others. If you're really datasciency you can vary the weights (maybe do a grid search) and compute dice for the test set to try and find the ideal combination. Gains are incremental though but if you've given up on finding a better model it's worth a shot. </p>",
      "rawMarkdown": "If the models have the same size you can try concatenating the outputs before the final convolution (usually a sigmoid/softmax in unet) and see if that works. Haven't had much success for my unets but I ran limited trials. Also if you're limited in memory this may hurt because you're running parallel models and joining the outputs. There is also the cost of testing on a larger network... There are other options, searching for somthing like 'keras join models', 'keras stacking models' or whatnot, depending on your framework may be a good read.\n\nAveraging out the outputs is the easiest to implement though, you just pick your best models and, well, average the output, i.e., the predictions before thresholding. More guaranteed to give you some gains too. You can also try voting. Haven't done either for lack of time but if you have a couple of models with good/similar scores it's worth a try. If one model is better then the other try weighing it in favor of the others. If you're really datasciency you can vary the weights (maybe do a grid search) and compute dice for the test set to try and find the ideal combination. Gains are incremental though but if you've given up on finding a better model it's worth a shot.",
      "votes": null
    },
    {
      "id": "222483",
      "postDate": "09/19/2017 01:38:24",
      "content": "<p>Try many times, all failured.  : (   I tried to average  two models' prediction with 1024x1024  solution likes these:\nmodel.load_weights(filepath='weights/best_weights5_1626.hdf5')<br>\npreds_list=do_predict(1024,k,model)<br>\npreds1=preds_list<br></p>\n\n<p>model.load_weights(filepath='weights/best_weights_h_1024_pro3.hdf5')<br>\n preds_list=do_predict(1024,k,model)<br>\n    preds1=np.array(preds1)<br>\n    preds_list=np.array(preds_list)<br>\n    preds2=preds1+preds_list<br></p>\n\n<pre><code>preds2=preds2/2 &lt;br&gt;\n</code></pre>\n\n<p>My pc has 32G memory and a gtx1060(6g gpu memory) .  To run the code ,I split the test data into 201 parts(one part has only 500 pictures).  But I always got   gpu memory errors after several epochs.  It seemed as if there was no way out.</p>",
      "rawMarkdown": "Try many times, all failured.  : (   I tried to average  two models' prediction with 1024x1024  solution likes these:\nmodel.load_weights(filepath='weights/best_weights5_1626.hdf5')<br>\npreds_list=do_predict(1024,k,model)<br>\npreds1=preds_list<br>\n\nmodel.load_weights(filepath='weights/best_weights_h_1024_pro3.hdf5')<br>\n preds_list=do_predict(1024,k,model)<br>\n    preds1=np.array(preds1)<br>\n    preds_list=np.array(preds_list)<br>\n    preds2=preds1+preds_list<br>\n\n    preds2=preds2/2 <br>\n\n My pc has 32G memory and a gtx1060(6g gpu memory) .  To run the code ,I split the test data into 201 parts(one part has only 500 pictures).  But I always got   gpu memory errors after several epochs.  It seemed as if there was no way out.",
      "votes": null
    },
    {
      "id": "222506",
      "postDate": "09/19/2017 03:50:48",
      "content": "<p>You can set up a generator that loads the test images in batches then convert the preds to rles.</p>",
      "rawMarkdown": "You can set up a generator that loads the test images in batches then convert the preds to rles.",
      "votes": null
    },
    {
      "id": "222508",
      "postDate": "09/19/2017 04:15:37",
      "content": "<p>do we really require a gpu for predictions since it could be done on cpu? too since it is 100k forward passes if I am not wrong.</p>",
      "rawMarkdown": "do we really require a gpu for predictions since it could be done on cpu? too since it is 100k forward passes if I am not wrong.",
      "votes": null
    },
    {
      "id": "222540",
      "postDate": "09/19/2017 06:23:40",
      "content": "<p>Instead of trying to generate complete predictions for different models and saving in-memory,  I suppose you can change your prediction code so as to predict on a mini-batch using all of your models and take the average of the sigmoid outputs for individual items and then apply the threshold. This worked for me (3 models, 8GB GPU) I am not sure if an ensemble of different resolutions for the same model would work well. For me, even an ensemble of three different models with voting did not do much better than the best model.</p>",
      "rawMarkdown": "Instead of trying to generate complete predictions for different models and saving in-memory,  I suppose you can change your prediction code so as to predict on a mini-batch using all of your models and take the average of the sigmoid outputs for individual items and then apply the threshold. This worked for me (3 models, 8GB GPU) I am not sure if an ensemble of different resolutions for the same model would work well. For me, even an ensemble of three different models with voting did not do much better than the best model.",
      "votes": null
    },
    {
      "id": "223965",
      "postDate": "09/24/2017 13:28:55",
      "content": "<p>You could 1) save your predicted masks of different models to folders then read these masks and do majority voting:</p>\n\n<pre><code>model_prediction_list = [glob.glob(model+'/*.gif') for model in model_list]\nresult = np.zeros((chunk_size, H, W))\n# iterate through each model\nfor model in model_predicton_list:\n  # iterate through predictions in each model\n  for i, m in enumerate(model):\n    m = imread(m)\n    result[i] = result[i] + m \nresult = (result &gt; (len(model_list // 2))).astype(np.uint8)\n</code></pre>\n\n<p>or 2) you can decode the rle files and use the same technique. I do not know whether there are faster and more memory-saving ways to do it.</p>\n\n<p>I tried averaging the probabilities(before thresholding) of each model, it does not give me a better result. </p>",
      "rawMarkdown": "You could 1) save your predicted masks of different models to folders then read these masks and do majority voting:\n\n    model_prediction_list = [glob.glob(model+'/*.gif') for model in model_list]\n    result = np.zeros((chunk_size, H, W))\n    # iterate through each model\n    for model in model_predicton_list:\n      # iterate through predictions in each model\n      for i, m in enumerate(model):\n        m = imread(m)\n        result[i] = result[i] + m \n    result = (result &gt; (len(model_list // 2))).astype(np.uint8)\nor 2) you can decode the rle files and use the same technique. I do not know whether there are faster and more memory-saving ways to do it.\n\nI tried averaging the probabilities(before thresholding) of each model, it does not give me a better result.",
      "votes": null
    },
    {
      "id": "224055",
      "postDate": "09/25/2017 00:00:04",
      "content": "<p>yash_r's suggestion worked for me. I disabled gpu with \n<code>export CUDA_VISIBLE_DEVICES=\"\"</code> and reloaded the Python kernel. It's taking ~5 hours to predict and ensemble two models.</p>",
      "rawMarkdown": "yash_r's suggestion worked for me. I disabled gpu with \n`export CUDA_VISIBLE_DEVICES=\"\"` and reloaded the Python kernel. It's taking ~5 hours to predict and ensemble two models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 220811,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "09/13/2017 06:24:39",
      "content": "<p>The easiest way to do an ensemble is to create predictions for both models. Then average the predictions and create a new submission using the average predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 220871,
          "author_name": "adamhart",
          "author_url": "",
          "post_date": "09/13/2017 10:51:13",
          "content": "<p>I haven't really done ensembles before, but I assumed that one should average out the outputs from the two (or more) networks before the soft_max operation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 221012,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "09/13/2017 19:10:25",
          "content": "<p>Im not sure about that, but in this competition you dont need softmax.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 221113,
          "author_name": "zhihang",
          "author_url": "",
          "post_date": "09/14/2017 06:20:51",
          "content": "<p>average gets 0.0002 improvement of my LB score</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 220853,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "09/13/2017 09:41:06",
      "content": "<p>Thanks.  Since I already had two submit files , maybe I can try to read these files and get the  average predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222280,
      "author_name": "anuragreddygv323",
      "author_url": "",
      "post_date": "09/18/2017 09:44:56",
      "content": "<p>any improvement with the ensemble ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222336,
      "author_name": "fpaboim",
      "author_url": "",
      "post_date": "09/18/2017 13:51:53",
      "content": "<p>If the models have the same size you can try concatenating the outputs before the final convolution (usually a sigmoid/softmax in unet) and see if that works. Haven't had much success for my unets but I ran limited trials. Also if you're limited in memory this may hurt because you're running parallel models and joining the outputs. There is also the cost of testing on a larger network... There are other options, searching for somthing like 'keras join models', 'keras stacking models' or whatnot, depending on your framework may be a good read.</p>\n\n<p>Averaging out the outputs is the easiest to implement though, you just pick your best models and, well, average the output, i.e., the predictions before thresholding. More guaranteed to give you some gains too. You can also try voting. Haven't done either for lack of time but if you have a couple of models with good/similar scores it's worth a try. If one model is better then the other try weighing it in favor of the others. If you're really datasciency you can vary the weights (maybe do a grid search) and compute dice for the test set to try and find the ideal combination. Gains are incremental though but if you've given up on finding a better model it's worth a shot. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222483,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "09/19/2017 01:38:24",
      "content": "<p>Try many times, all failured.  : (   I tried to average  two models' prediction with 1024x1024  solution likes these:\nmodel.load_weights(filepath='weights/best_weights5_1626.hdf5')<br>\npreds_list=do_predict(1024,k,model)<br>\npreds1=preds_list<br></p>\n\n<p>model.load_weights(filepath='weights/best_weights_h_1024_pro3.hdf5')<br>\n preds_list=do_predict(1024,k,model)<br>\n    preds1=np.array(preds1)<br>\n    preds_list=np.array(preds_list)<br>\n    preds2=preds1+preds_list<br></p>\n\n<pre><code>preds2=preds2/2 &lt;br&gt;\n</code></pre>\n\n<p>My pc has 32G memory and a gtx1060(6g gpu memory) .  To run the code ,I split the test data into 201 parts(one part has only 500 pictures).  But I always got   gpu memory errors after several epochs.  It seemed as if there was no way out.</p>",
      "votes": null,
      "replies": [
        {
          "id": 222506,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "09/19/2017 03:50:48",
          "content": "<p>You can set up a generator that loads the test images in batches then convert the preds to rles.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 222508,
          "author_name": "remidi",
          "author_url": "",
          "post_date": "09/19/2017 04:15:37",
          "content": "<p>do we really require a gpu for predictions since it could be done on cpu? too since it is 100k forward passes if I am not wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 222540,
          "author_name": "datasec",
          "author_url": "",
          "post_date": "09/19/2017 06:23:40",
          "content": "<p>Instead of trying to generate complete predictions for different models and saving in-memory,  I suppose you can change your prediction code so as to predict on a mini-batch using all of your models and take the average of the sigmoid outputs for individual items and then apply the threshold. This worked for me (3 models, 8GB GPU) I am not sure if an ensemble of different resolutions for the same model would work well. For me, even an ensemble of three different models with voting did not do much better than the best model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224055,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "09/25/2017 00:00:04",
          "content": "<p>yash_r's suggestion worked for me. I disabled gpu with \n<code>export CUDA_VISIBLE_DEVICES=\"\"</code> and reloaded the Python kernel. It's taking ~5 hours to predict and ensemble two models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 223965,
      "author_name": "junhongxu",
      "author_url": "",
      "post_date": "09/24/2017 13:28:55",
      "content": "<p>You could 1) save your predicted masks of different models to folders then read these masks and do majority voting:</p>\n\n<pre><code>model_prediction_list = [glob.glob(model+'/*.gif') for model in model_list]\nresult = np.zeros((chunk_size, H, W))\n# iterate through each model\nfor model in model_predicton_list:\n  # iterate through predictions in each model\n  for i, m in enumerate(model):\n    m = imread(m)\n    result[i] = result[i] + m \nresult = (result &gt; (len(model_list // 2))).astype(np.uint8)\n</code></pre>\n\n<p>or 2) you can decode the rle files and use the same technique. I do not know whether there are faster and more memory-saving ways to do it.</p>\n\n<p>I tried averaging the probabilities(before thresholding) of each model, it does not give me a better result. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "220788": "I have two models which respectively trained by 1152x1152,1024x1024 . I use [Peter's codes][1] to train these models . And I have only a gtx1060 gpu, it is very slow to generate the submit file(about 9 hours for 1152x1152 model, and about 6 hours for 1024x1024 model)  . I don't know how to ensemble them and generate the submit file.  Any suggestion?  \n\n\n  [1]: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37523",
    "220811": "The easiest way to do an ensemble is to create predictions for both models. Then average the predictions and create a new submission using the average predictions.",
    "220853": "Thanks.  Since I already had two submit files , maybe I can try to read these files and get the  average predictions.",
    "220871": "I haven't really done ensembles before, but I assumed that one should average out the outputs from the two (or more) networks before the soft_max operation?",
    "221012": "Im not sure about that, but in this competition you dont need softmax.",
    "221113": "average gets 0.0002 improvement of my LB score",
    "222280": "any improvement with the ensemble ?",
    "222336": "If the models have the same size you can try concatenating the outputs before the final convolution (usually a sigmoid/softmax in unet) and see if that works. Haven't had much success for my unets but I ran limited trials. Also if you're limited in memory this may hurt because you're running parallel models and joining the outputs. There is also the cost of testing on a larger network... There are other options, searching for somthing like 'keras join models', 'keras stacking models' or whatnot, depending on your framework may be a good read.\n\nAveraging out the outputs is the easiest to implement though, you just pick your best models and, well, average the output, i.e., the predictions before thresholding. More guaranteed to give you some gains too. You can also try voting. Haven't done either for lack of time but if you have a couple of models with good/similar scores it's worth a try. If one model is better then the other try weighing it in favor of the others. If you're really datasciency you can vary the weights (maybe do a grid search) and compute dice for the test set to try and find the ideal combination. Gains are incremental though but if you've given up on finding a better model it's worth a shot.",
    "222483": "Try many times, all failured.  : (   I tried to average  two models' prediction with 1024x1024  solution likes these:\nmodel.load_weights(filepath='weights/best_weights5_1626.hdf5')<br>\npreds_list=do_predict(1024,k,model)<br>\npreds1=preds_list<br>\n\nmodel.load_weights(filepath='weights/best_weights_h_1024_pro3.hdf5')<br>\n preds_list=do_predict(1024,k,model)<br>\n    preds1=np.array(preds1)<br>\n    preds_list=np.array(preds_list)<br>\n    preds2=preds1+preds_list<br>\n\n    preds2=preds2/2 <br>\n\n My pc has 32G memory and a gtx1060(6g gpu memory) .  To run the code ,I split the test data into 201 parts(one part has only 500 pictures).  But I always got   gpu memory errors after several epochs.  It seemed as if there was no way out.",
    "222506": "You can set up a generator that loads the test images in batches then convert the preds to rles.",
    "222508": "do we really require a gpu for predictions since it could be done on cpu? too since it is 100k forward passes if I am not wrong.",
    "222540": "Instead of trying to generate complete predictions for different models and saving in-memory,  I suppose you can change your prediction code so as to predict on a mini-batch using all of your models and take the average of the sigmoid outputs for individual items and then apply the threshold. This worked for me (3 models, 8GB GPU) I am not sure if an ensemble of different resolutions for the same model would work well. For me, even an ensemble of three different models with voting did not do much better than the best model.",
    "223965": "You could 1) save your predicted masks of different models to folders then read these masks and do majority voting:\n\n    model_prediction_list = [glob.glob(model+'/*.gif') for model in model_list]\n    result = np.zeros((chunk_size, H, W))\n    # iterate through each model\n    for model in model_predicton_list:\n      # iterate through predictions in each model\n      for i, m in enumerate(model):\n        m = imread(m)\n        result[i] = result[i] + m \n    result = (result &gt; (len(model_list // 2))).astype(np.uint8)\nor 2) you can decode the rle files and use the same technique. I do not know whether there are faster and more memory-saving ways to do it.\n\nI tried averaging the probabilities(before thresholding) of each model, it does not give me a better result.",
    "224055": "yash_r's suggestion worked for me. I disabled gpu with \n`export CUDA_VISIBLE_DEVICES=\"\"` and reloaded the Python kernel. It's taking ~5 hours to predict and ensemble two models."
  },
  "source": "meta"
}