{
  "id": 73640,
  "title": "How to ensemble output from two models",
  "url": "/competitions/PLAsTiCC-2018/discussion/73640",
  "author_name": "",
  "post_date": "2018-12-04T14:19:43.667971900Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have two models with following confusion matrix </p>\n\n<p>Model 1</p>\n\n<p><img src=\"https://www.kaggleusercontent.com/kf/8064923/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..yfRrF1UXqpkBCfZqurLrhg.NbHo7XQ5sLpvLxiraWUoBV-QoZrhQa5yhxmT9OZoDTEFuHpaX7s1MLJ3K9EyehEKCWXJCg92j6f-E16P45g8ulejIWJBWWxPuN65LcLeB_BIj-qtU08AioCSetFRDmDGMpTRwLnzoXsQlvUPu_vq_TbQH2-vedpmOu16wd2ll_dCaolZkxI0SlnV653dpLwF.SzF7m_YPFbVh_eTkyUFayQ/__results___files/__results___17_17.png\" alt=\"Model1 Confusion matrix\"></p>\n\n<p>Model 2\n<img src=\"https://www.kaggleusercontent.com/kf/8028170/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..t5Lw4O2iOOkPhEzD3Epvaw.HNMZI3Te7dZcmZqK-thpOOCw53Jbj3T6yN8msZlyJQVQIVL8EaLmnfOEucbF4crW7e5a5E-rozsCsBrJEEOc6UoXcwkJxCcEQDMDnnl_ly9tYBCT64vcXiFf2au3gncZrQ48SHu-h5Rp8Vl_j1Uls_AyjMaQ5-PWcmIgE_JMcqzfukGNm5jQtQCTKEH_BUtM.U5RW2ieGf7uM8lGG07zhqg/__results___files/__results___12_2.png\" alt=\"Model2 Confusion matrix\"></p>\n\n<p>Model1 is performing better on classes 52, 67\n model2 is performing better on classes 42, 62, 90\nand both are performing similarly on classes 6, 15, 16, 53, 64, 65, 88, 92, 95</p>\n\n<p>I want to create an ensemble that takes into account these biases.\nAny pointers on how I could proceed? Also, both these models are taking around 6 hours to complete, so I want to just work with the test output files rather than on the models in a single kernel.</p>\n\n<p>I am not sure if I could ask this question here as this is a competetion, but I am just looking for some directions. Thanks in advance.</p>",
  "messages": [
    {
      "id": "432962",
      "postDate": "12/04/2018 14:19:43",
      "content": "<p>I have two models with following confusion matrix </p>\n\n<p>Model 1</p>\n\n<p><img src=\"https://www.kaggleusercontent.com/kf/8064923/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..yfRrF1UXqpkBCfZqurLrhg.NbHo7XQ5sLpvLxiraWUoBV-QoZrhQa5yhxmT9OZoDTEFuHpaX7s1MLJ3K9EyehEKCWXJCg92j6f-E16P45g8ulejIWJBWWxPuN65LcLeB_BIj-qtU08AioCSetFRDmDGMpTRwLnzoXsQlvUPu_vq_TbQH2-vedpmOu16wd2ll_dCaolZkxI0SlnV653dpLwF.SzF7m_YPFbVh_eTkyUFayQ/__results___files/__results___17_17.png\" alt=\"Model1 Confusion matrix\"></p>\n\n<p>Model 2\n<img src=\"https://www.kaggleusercontent.com/kf/8028170/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..t5Lw4O2iOOkPhEzD3Epvaw.HNMZI3Te7dZcmZqK-thpOOCw53Jbj3T6yN8msZlyJQVQIVL8EaLmnfOEucbF4crW7e5a5E-rozsCsBrJEEOc6UoXcwkJxCcEQDMDnnl_ly9tYBCT64vcXiFf2au3gncZrQ48SHu-h5Rp8Vl_j1Uls_AyjMaQ5-PWcmIgE_JMcqzfukGNm5jQtQCTKEH_BUtM.U5RW2ieGf7uM8lGG07zhqg/__results___files/__results___12_2.png\" alt=\"Model2 Confusion matrix\"></p>\n\n<p>Model1 is performing better on classes 52, 67\n model2 is performing better on classes 42, 62, 90\nand both are performing similarly on classes 6, 15, 16, 53, 64, 65, 88, 92, 95</p>\n\n<p>I want to create an ensemble that takes into account these biases.\nAny pointers on how I could proceed? Also, both these models are taking around 6 hours to complete, so I want to just work with the test output files rather than on the models in a single kernel.</p>\n\n<p>I am not sure if I could ask this question here as this is a competetion, but I am just looking for some directions. Thanks in advance.</p>",
      "rawMarkdown": "I have two models with following confusion matrix \n\nModel 1\n\n![Model1 Confusion matrix][1]\n\n  [1]: https://www.kaggleusercontent.com/kf/8064923/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..yfRrF1UXqpkBCfZqurLrhg.NbHo7XQ5sLpvLxiraWUoBV-QoZrhQa5yhxmT9OZoDTEFuHpaX7s1MLJ3K9EyehEKCWXJCg92j6f-E16P45g8ulejIWJBWWxPuN65LcLeB_BIj-qtU08AioCSetFRDmDGMpTRwLnzoXsQlvUPu_vq_TbQH2-vedpmOu16wd2ll_dCaolZkxI0SlnV653dpLwF.SzF7m_YPFbVh_eTkyUFayQ/__results___files/__results___17_17.png\n\nModel 2\n![Model2 Confusion matrix][2]\n\n\n\n  [2]: https://www.kaggleusercontent.com/kf/8028170/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..t5Lw4O2iOOkPhEzD3Epvaw.HNMZI3Te7dZcmZqK-thpOOCw53Jbj3T6yN8msZlyJQVQIVL8EaLmnfOEucbF4crW7e5a5E-rozsCsBrJEEOc6UoXcwkJxCcEQDMDnnl_ly9tYBCT64vcXiFf2au3gncZrQ48SHu-h5Rp8Vl_j1Uls_AyjMaQ5-PWcmIgE_JMcqzfukGNm5jQtQCTKEH_BUtM.U5RW2ieGf7uM8lGG07zhqg/__results___files/__results___12_2.png\n\n\nModel1 is performing better on classes 52, 67\n model2 is performing better on classes 42, 62, 90\nand both are performing similarly on classes 6, 15, 16, 53, 64, 65, 88, 92, 95\n\nI want to create an ensemble that takes into account these biases.\nAny pointers on how I could proceed? Also, both these models are taking around 6 hours to complete, so I want to just work with the test output files rather than on the models in a single kernel.\n\nI am not sure if I could ask this question here as this is a competetion, but I am just looking for some directions. Thanks in advance.",
      "votes": null
    },
    {
      "id": "432995",
      "postDate": "12/04/2018 14:51:57",
      "content": "<p>I would take the mean of their predictions to start with.</p>",
      "rawMarkdown": "I would take the mean of their predictions to start with.",
      "votes": null
    },
    {
      "id": "433046",
      "postDate": "12/04/2018 15:54:50",
      "content": "<p>Example of blending multiple submissions from different models <a href=\"https://www.kaggle.com/bulatza/blend-them-all-best-public-score\">https://www.kaggle.com/bulatza/blend-them-all-best-public-score</a></p>",
      "rawMarkdown": "Example of blending multiple submissions from different models https://www.kaggle.com/bulatza/blend-them-all-best-public-score",
      "votes": null
    },
    {
      "id": "433262",
      "postDate": "12/04/2018 23:38:15",
      "content": "<p>This link is very useful, you can learn some ensemble techniques here. (like stacking)\n<a href=\"https://mlwave.com/kaggle-ensembling-guide/\">https://mlwave.com/kaggle-ensembling-guide/</a></p>",
      "rawMarkdown": "This link is very useful, you can learn some ensemble techniques here. (like stacking)\nhttps://mlwave.com/kaggle-ensembling-guide/",
      "votes": null
    },
    {
      "id": "433279",
      "postDate": "12/05/2018 00:18:42",
      "content": "<p>@Jack, @mamas Thanks for the useful links. Lot of useful information.\n@CPMP good suggestion, indeed gives me good thought on how to apply the weights. Applying mean is like starting with 0.5 weight, then I can tweak it for some prediction classes.</p>",
      "rawMarkdown": "Jack, @mamas Thanks for the useful links. Lot of useful information.\n@CPMP good suggestion, indeed gives me good thought on how to apply the weights. Applying mean is like starting with 0.5 weight, then I can tweak it for some prediction classes.",
      "votes": null
    },
    {
      "id": "433295",
      "postDate": "12/05/2018 00:42:01",
      "content": "<p>Some more info on ensemble methods  <a href=\"https://www.coursera.org/learn/competitive-data-science/lecture/MJKCi/introduction-into-ensemble-methods\">https://www.coursera.org/learn/competitive-data-science/lecture/MJKCi/introduction-into-ensemble-methods</a> . You might need to log into coursera and choose to audit the course to view the content. </p>",
      "rawMarkdown": "Some more info on ensemble methods  https://www.coursera.org/learn/competitive-data-science/lecture/MJKCi/introduction-into-ensemble-methods . You might need to log into coursera and choose to audit the course to view the content.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 432995,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/04/2018 14:51:57",
      "content": "<p>I would take the mean of their predictions to start with.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 433046,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/04/2018 15:54:50",
      "content": "<p>Example of blending multiple submissions from different models <a href=\"https://www.kaggle.com/bulatza/blend-them-all-best-public-score\">https://www.kaggle.com/bulatza/blend-them-all-best-public-score</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 433262,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/04/2018 23:38:15",
      "content": "<p>This link is very useful, you can learn some ensemble techniques here. (like stacking)\n<a href=\"https://mlwave.com/kaggle-ensembling-guide/\">https://mlwave.com/kaggle-ensembling-guide/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 433279,
      "author_name": "subrahmanyamv",
      "author_url": "",
      "post_date": "12/05/2018 00:18:42",
      "content": "<p>@Jack, @mamas Thanks for the useful links. Lot of useful information.\n@CPMP good suggestion, indeed gives me good thought on how to apply the weights. Applying mean is like starting with 0.5 weight, then I can tweak it for some prediction classes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 433295,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/05/2018 00:42:01",
      "content": "<p>Some more info on ensemble methods  <a href=\"https://www.coursera.org/learn/competitive-data-science/lecture/MJKCi/introduction-into-ensemble-methods\">https://www.coursera.org/learn/competitive-data-science/lecture/MJKCi/introduction-into-ensemble-methods</a> . You might need to log into coursera and choose to audit the course to view the content. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "432962": "I have two models with following confusion matrix \n\nModel 1\n\n![Model1 Confusion matrix][1]\n\n  [1]: https://www.kaggleusercontent.com/kf/8064923/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..yfRrF1UXqpkBCfZqurLrhg.NbHo7XQ5sLpvLxiraWUoBV-QoZrhQa5yhxmT9OZoDTEFuHpaX7s1MLJ3K9EyehEKCWXJCg92j6f-E16P45g8ulejIWJBWWxPuN65LcLeB_BIj-qtU08AioCSetFRDmDGMpTRwLnzoXsQlvUPu_vq_TbQH2-vedpmOu16wd2ll_dCaolZkxI0SlnV653dpLwF.SzF7m_YPFbVh_eTkyUFayQ/__results___files/__results___17_17.png\n\nModel 2\n![Model2 Confusion matrix][2]\n\n\n\n  [2]: https://www.kaggleusercontent.com/kf/8028170/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..t5Lw4O2iOOkPhEzD3Epvaw.HNMZI3Te7dZcmZqK-thpOOCw53Jbj3T6yN8msZlyJQVQIVL8EaLmnfOEucbF4crW7e5a5E-rozsCsBrJEEOc6UoXcwkJxCcEQDMDnnl_ly9tYBCT64vcXiFf2au3gncZrQ48SHu-h5Rp8Vl_j1Uls_AyjMaQ5-PWcmIgE_JMcqzfukGNm5jQtQCTKEH_BUtM.U5RW2ieGf7uM8lGG07zhqg/__results___files/__results___12_2.png\n\n\nModel1 is performing better on classes 52, 67\n model2 is performing better on classes 42, 62, 90\nand both are performing similarly on classes 6, 15, 16, 53, 64, 65, 88, 92, 95\n\nI want to create an ensemble that takes into account these biases.\nAny pointers on how I could proceed? Also, both these models are taking around 6 hours to complete, so I want to just work with the test output files rather than on the models in a single kernel.\n\nI am not sure if I could ask this question here as this is a competetion, but I am just looking for some directions. Thanks in advance.",
    "432995": "I would take the mean of their predictions to start with.",
    "433046": "Example of blending multiple submissions from different models https://www.kaggle.com/bulatza/blend-them-all-best-public-score",
    "433262": "This link is very useful, you can learn some ensemble techniques here. (like stacking)\nhttps://mlwave.com/kaggle-ensembling-guide/",
    "433279": "Jack, @mamas Thanks for the useful links. Lot of useful information.\n@CPMP good suggestion, indeed gives me good thought on how to apply the weights. Applying mean is like starting with 0.5 weight, then I can tweak it for some prediction classes.",
    "433295": "Some more info on ensemble methods  https://www.coursera.org/learn/competitive-data-science/lecture/MJKCi/introduction-into-ensemble-methods . You might need to log into coursera and choose to audit the course to view the content."
  },
  "source": "meta"
}