{
  "id": 59528,
  "title": "Rules on Blending",
  "url": "/competitions/avito-demand-prediction/discussion/59528",
  "author_name": "",
  "post_date": "2018-06-23T10:47:26.085727100Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Are there any rules on blending a model? A recommended minimum/maximum number of model/s? Any special rule and/or rule of exception? Thanks! :-)</p>",
  "messages": [
    {
      "id": "347137",
      "postDate": "06/23/2018 10:47:26",
      "content": "<p>Are there any rules on blending a model? A recommended minimum/maximum number of model/s? Any special rule and/or rule of exception? Thanks! :-)</p>",
      "rawMarkdown": "Are there any rules on blending a model? A recommended minimum/maximum number of model/s? Any special rule and/or rule of exception? Thanks! :-)",
      "votes": null
    },
    {
      "id": "347178",
      "postDate": "06/23/2018 14:04:55",
      "content": "<p>You'd want to make sure that the predictions used in the blend aren't correlated. You can use <code>np. corrcoef (...)</code> for a numpy array or <code>df.corr()</code> for a pandas DataFrame. </p>\n\n<p>If you have a lot of good models which are highly correlated, just take their mean (or <a href=\"https://docs.scipy.org/doc/scipy-0.13.0/reference/generated/scipy.stats.mstats.gmean.html\">gmean</a> or <a href=\"https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.stats.hmean.html\">hmean</a>) before adding it to the blender.</p>",
      "rawMarkdown": "You'd want to make sure that the predictions used in the blend aren't correlated. You can use `np. corrcoef (...)` for a numpy array or `df.corr()` for a pandas DataFrame. \n\nIf you have a lot of good models which are highly correlated, just take their mean (or [gmean][1] or [hmean][2]) before adding it to the blender.\n\n  [1]: https://docs.scipy.org/doc/scipy-0.13.0/reference/generated/scipy.stats.mstats.gmean.html\n  [2]: https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.stats.hmean.html",
      "votes": null
    },
    {
      "id": "347231",
      "postDate": "06/23/2018 16:42:35",
      "content": "<p><a href=\"/quantumgeek\">@quantumgeek</a>: Thanks for your feedback! much appreciated! :-)</p>",
      "rawMarkdown": "quantumgeek: Thanks for your feedback! much appreciated! :-)",
      "votes": null
    },
    {
      "id": "347365",
      "postDate": "06/24/2018 04:17:35",
      "content": "<p>More Nice Models and More Diversity = High LB  Score</p>",
      "rawMarkdown": "More Nice Models and More Diversity = High LB  Score",
      "votes": null
    },
    {
      "id": "347592",
      "postDate": "06/24/2018 19:31:33",
      "content": "<p>Pretty nice tutorial <a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\">https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python</a> (really big guide)\n<a href=\"https://mlwave.com/kaggle-ensembling-guide/\">https://mlwave.com/kaggle-ensembling-guide/</a> (must to read)</p>",
      "rawMarkdown": "Pretty nice tutorial https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python (really big guide)\nhttps://mlwave.com/kaggle-ensembling-guide/ (must to read)",
      "votes": null
    },
    {
      "id": "347696",
      "postDate": "06/25/2018 04:53:56",
      "content": "<p>Hi, Thanks for your feedback! much appreciated! :-)</p>",
      "rawMarkdown": "Hi, Thanks for your feedback! much appreciated! :-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 347178,
      "author_name": "yk1598",
      "author_url": "",
      "post_date": "06/23/2018 14:04:55",
      "content": "<p>You'd want to make sure that the predictions used in the blend aren't correlated. You can use <code>np. corrcoef (...)</code> for a numpy array or <code>df.corr()</code> for a pandas DataFrame. </p>\n\n<p>If you have a lot of good models which are highly correlated, just take their mean (or <a href=\"https://docs.scipy.org/doc/scipy-0.13.0/reference/generated/scipy.stats.mstats.gmean.html\">gmean</a> or <a href=\"https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.stats.hmean.html\">hmean</a>) before adding it to the blender.</p>",
      "votes": null,
      "replies": [
        {
          "id": 347231,
          "author_name": "",
          "author_url": "",
          "post_date": "06/23/2018 16:42:35",
          "content": "<p><a href=\"/quantumgeek\">@quantumgeek</a>: Thanks for your feedback! much appreciated! :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 347365,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "06/24/2018 04:17:35",
      "content": "<p>More Nice Models and More Diversity = High LB  Score</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 347592,
      "author_name": "insaff",
      "author_url": "",
      "post_date": "06/24/2018 19:31:33",
      "content": "<p>Pretty nice tutorial <a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\">https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python</a> (really big guide)\n<a href=\"https://mlwave.com/kaggle-ensembling-guide/\">https://mlwave.com/kaggle-ensembling-guide/</a> (must to read)</p>",
      "votes": null,
      "replies": [
        {
          "id": 347696,
          "author_name": "",
          "author_url": "",
          "post_date": "06/25/2018 04:53:56",
          "content": "<p>Hi, Thanks for your feedback! much appreciated! :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "347137": "Are there any rules on blending a model? A recommended minimum/maximum number of model/s? Any special rule and/or rule of exception? Thanks! :-)",
    "347178": "You'd want to make sure that the predictions used in the blend aren't correlated. You can use `np. corrcoef (...)` for a numpy array or `df.corr()` for a pandas DataFrame. \n\nIf you have a lot of good models which are highly correlated, just take their mean (or [gmean][1] or [hmean][2]) before adding it to the blender.\n\n  [1]: https://docs.scipy.org/doc/scipy-0.13.0/reference/generated/scipy.stats.mstats.gmean.html\n  [2]: https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.stats.hmean.html",
    "347231": "quantumgeek: Thanks for your feedback! much appreciated! :-)",
    "347365": "More Nice Models and More Diversity = High LB  Score",
    "347592": "Pretty nice tutorial https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python (really big guide)\nhttps://mlwave.com/kaggle-ensembling-guide/ (must to read)",
    "347696": "Hi, Thanks for your feedback! much appreciated! :-)"
  },
  "source": "meta"
}