{
  "id": 59963,
  "title": "Optimal linear ensemble with LB RMSEs",
  "url": "/competitions/avito-demand-prediction/discussion/59963",
  "author_name": "",
  "post_date": "2018-06-28T23:28:27.561628700Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>For a linear ensemble, we can calculate the optimal weights of input model predictions from RMSEs of predictions and all-zero prediction. </p>\n\n<p>Netflix Grand Prize (2006-2009) winners (including one of my favorite Kagglers, <a href=\"https://www.kaggle.com/mjahrer\">Michael Jahrer</a>) shared this approach in Section 7.1 of <a href=\"https://www.netflixprize.com/assets/GrandPrize2009_BPC_BigChaos.pdf\">their paper</a>. </p>\n\n<p>To make it simple, I added a function, <code>ensemble.netflix()</code> to the Python <code>Kaggler</code> package. You can check out the example here at <a href=\"https://github.com/jeongyoonlee/kaggler#netflix-blending\">the Kaggler Github page</a>.</p>\n\n<p>Hope this helps prevent last minute surprise of the \"public blending\" solution - because now we all know how to do it. ;)</p>",
  "messages": [
    {
      "id": "349945",
      "postDate": "06/28/2018 23:28:27",
      "content": "<p>For a linear ensemble, we can calculate the optimal weights of input model predictions from RMSEs of predictions and all-zero prediction. </p>\n\n<p>Netflix Grand Prize (2006-2009) winners (including one of my favorite Kagglers, <a href=\"https://www.kaggle.com/mjahrer\">Michael Jahrer</a>) shared this approach in Section 7.1 of <a href=\"https://www.netflixprize.com/assets/GrandPrize2009_BPC_BigChaos.pdf\">their paper</a>. </p>\n\n<p>To make it simple, I added a function, <code>ensemble.netflix()</code> to the Python <code>Kaggler</code> package. You can check out the example here at <a href=\"https://github.com/jeongyoonlee/kaggler#netflix-blending\">the Kaggler Github page</a>.</p>\n\n<p>Hope this helps prevent last minute surprise of the \"public blending\" solution - because now we all know how to do it. ;)</p>",
      "rawMarkdown": "For a linear ensemble, we can calculate the optimal weights of input model predictions from RMSEs of predictions and all-zero prediction. \n\nNetflix Grand Prize (2006-2009) winners (including one of my favorite Kagglers, [Michael Jahrer][1]) shared this approach in Section 7.1 of [their paper][2]. \n\nTo make it simple, I added a function, `ensemble.netflix()` to the Python `Kaggler` package. You can check out the example here at [the Kaggler Github page](https://github.com/jeongyoonlee/kaggler#netflix-blending).\n\nHope this helps prevent last minute surprise of the \"public blending\" solution - because now we all know how to do it. ;)\n\n\n  [1]: https://www.kaggle.com/mjahrer\n  [2]: https://www.netflixprize.com/assets/GrandPrize2009_BPC_BigChaos.pdf",
      "votes": null
    },
    {
      "id": "350246",
      "postDate": "06/29/2018 12:57:49",
      "content": "<p>Sorry for question, but in case of \"public blending solution\" where we have only set of \"test_y\" and not \"oof_y\" , what is</p>\n\n<pre><code>y = np.loadtxt('target.txt') \n</code></pre>\n\n<p>in your example?</p>",
      "rawMarkdown": "Sorry for question, but in case of \"public blending solution\" where we have only set of \"test_y\" and not \"oof_y\" , what is\n\n\n    y = np.loadtxt('target.txt') \n\n\nin your example?",
      "votes": null
    },
    {
      "id": "350348",
      "postDate": "06/29/2018 15:36:15",
      "content": "<p>In the example, the target values were used only to calculate RMSEs. </p>\n\n<p>As I put in the comment, at a competition, RMSEs (or RMLSEs) of submissions in the public leaderboard can be used, and you don't need the target values.</p>\n\n<p>BTW, no need to be sorry. Thanks for the question! This is the very purpose of the forum. :)</p>",
      "rawMarkdown": "In the example, the target values were used only to calculate RMSEs. \n\nAs I put in the comment, at a competition, RMSEs (or RMLSEs) of submissions in the public leaderboard can be used, and you don't need the target values.\n\nBTW, no need to be sorry. Thanks for the question! This is the very purpose of the forum. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 350246,
      "author_name": "kruegger",
      "author_url": "",
      "post_date": "06/29/2018 12:57:49",
      "content": "<p>Sorry for question, but in case of \"public blending solution\" where we have only set of \"test_y\" and not \"oof_y\" , what is</p>\n\n<pre><code>y = np.loadtxt('target.txt') \n</code></pre>\n\n<p>in your example?</p>",
      "votes": null,
      "replies": [
        {
          "id": 350348,
          "author_name": "jeongyoonlee",
          "author_url": "",
          "post_date": "06/29/2018 15:36:15",
          "content": "<p>In the example, the target values were used only to calculate RMSEs. </p>\n\n<p>As I put in the comment, at a competition, RMSEs (or RMLSEs) of submissions in the public leaderboard can be used, and you don't need the target values.</p>\n\n<p>BTW, no need to be sorry. Thanks for the question! This is the very purpose of the forum. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "349945": "For a linear ensemble, we can calculate the optimal weights of input model predictions from RMSEs of predictions and all-zero prediction. \n\nNetflix Grand Prize (2006-2009) winners (including one of my favorite Kagglers, [Michael Jahrer][1]) shared this approach in Section 7.1 of [their paper][2]. \n\nTo make it simple, I added a function, `ensemble.netflix()` to the Python `Kaggler` package. You can check out the example here at [the Kaggler Github page](https://github.com/jeongyoonlee/kaggler#netflix-blending).\n\nHope this helps prevent last minute surprise of the \"public blending\" solution - because now we all know how to do it. ;)\n\n\n  [1]: https://www.kaggle.com/mjahrer\n  [2]: https://www.netflixprize.com/assets/GrandPrize2009_BPC_BigChaos.pdf",
    "350246": "Sorry for question, but in case of \"public blending solution\" where we have only set of \"test_y\" and not \"oof_y\" , what is\n\n\n    y = np.loadtxt('target.txt') \n\n\nin your example?",
    "350348": "In the example, the target values were used only to calculate RMSEs. \n\nAs I put in the comment, at a competition, RMSEs (or RMLSEs) of submissions in the public leaderboard can be used, and you don't need the target values.\n\nBTW, no need to be sorry. Thanks for the question! This is the very purpose of the forum. :)"
  },
  "source": "meta"
}