{
  "id": 362751,
  "title": "To win - blend 1000 models ? Recipe to make them.  ",
  "url": "/competitions/open-problems-multimodal/discussion/362751",
  "author_name": "Alexander Chervov",
  "post_date": "2022-10-28T21:43:37.552000",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Briefly - idea -</strong>  how to convert your single model to many-many similar models, such that their blend might be better than any single. </p>\n<p><strong>Context/motivation:</strong> What is surprising (for me) that blending seems to give more improvement than for other similar competitions.<br>\n(Even \"relatively weak\" models like Ridge and in general blending in that competition gives improvement which are somewhat bigger than expected).   Why it happens and is it really true ( or I am wrong) - that is another interesting question. </p>\n<p>So one can try to use several tricks to produce from a  model  many similar, typically weaker than original, but with some diversity (and then - blend them all): </p>\n<p><strong>Tricks:</strong></p>\n<p>1) For NN - if you optimize correlation metric - try to use standard MSE or another </p>\n<p>2) For NN - predict not ALL targets but choose group of them and predict only them, <br>\ntake many subgroups - then blend. (It requires metrics like MSE, probably, correlation will not work). <br>\nAs subgroups it might be wise to take correlated groups of targets.<br>\n(Correlated groups of targets can be found here: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02?scriptVersionId=106444735&amp;cellId=54\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02?scriptVersionId=106444735&amp;cellId=54</a> )</p>\n<p>3) Instead of predicting targets directly one may try to use PCA and predict PCA[: , :N] for some N and <br>\nthen use inverse transform to restore targets. One can try to use other dimensional reduction algorithms <br>\neven non-linear like UMAP - we just need inverse_transform. </p>\n<p>4) Feature engineering - use various feature engineering constructions - to add different sets of features<br>\nto your basic models. <br>\n(For that particular competition - there are plenty datasets with ready to use features:<br>\nsome along the ways described here: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361025\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361025</a> )</p>\n<p>You may try as a first step to select some important features by some feature importance algorithm,<br>\nand use them. But one can do it on some subset of the data - thus changing subsets of the data - <br>\nwill most probably will give another set of important features and so one can try to use thus obtained different geatures for your models and thus obtain plenty of the models. </p>\n<p>5) Use different random seeds - but it may not work - not enough diversity. </p>\n<p>PS<br>\nSo in part it is motivated by Random Forest approach - you take one simple model like decision tree,<br>\nbut then you make many-many models based on it - using different feature subsets, different sample subsets, etc…<br>\nand then - blend them all.</p>\n<p>Any comments are welcome. </p>",
  "messages": [
    {
      "id": 2008271,
      "postDate": "2022-10-28T21:43:37.553Z",
      "content": "<p><strong>Briefly - idea -</strong>  how to convert your single model to many-many similar models, such that their blend might be better than any single. </p>\n<p><strong>Context/motivation:</strong> What is surprising (for me) that blending seems to give more improvement than for other similar competitions.<br>\n(Even \"relatively weak\" models like Ridge and in general blending in that competition gives improvement which are somewhat bigger than expected).   Why it happens and is it really true ( or I am wrong) - that is another interesting question. </p>\n<p>So one can try to use several tricks to produce from a  model  many similar, typically weaker than original, but with some diversity (and then - blend them all): </p>\n<p><strong>Tricks:</strong></p>\n<p>1) For NN - if you optimize correlation metric - try to use standard MSE or another </p>\n<p>2) For NN - predict not ALL targets but choose group of them and predict only them, <br>\ntake many subgroups - then blend. (It requires metrics like MSE, probably, correlation will not work). <br>\nAs subgroups it might be wise to take correlated groups of targets.<br>\n(Correlated groups of targets can be found here: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02?scriptVersionId=106444735&amp;cellId=54\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02?scriptVersionId=106444735&amp;cellId=54</a> )</p>\n<p>3) Instead of predicting targets directly one may try to use PCA and predict PCA[: , :N] for some N and <br>\nthen use inverse transform to restore targets. One can try to use other dimensional reduction algorithms <br>\neven non-linear like UMAP - we just need inverse_transform. </p>\n<p>4) Feature engineering - use various feature engineering constructions - to add different sets of features<br>\nto your basic models. <br>\n(For that particular competition - there are plenty datasets with ready to use features:<br>\nsome along the ways described here: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361025\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/361025</a> )</p>\n<p>You may try as a first step to select some important features by some feature importance algorithm,<br>\nand use them. But one can do it on some subset of the data - thus changing subsets of the data - <br>\nwill most probably will give another set of important features and so one can try to use thus obtained different geatures for your models and thus obtain plenty of the models. </p>\n<p>5) Use different random seeds - but it may not work - not enough diversity. </p>\n<p>PS<br>\nSo in part it is motivated by Random Forest approach - you take one simple model like decision tree,<br>\nbut then you make many-many models based on it - using different feature subsets, different sample subsets, etc…<br>\nand then - blend them all.</p>\n<p>Any comments are welcome. </p>",
      "rawMarkdown": "**Briefly - idea -**  how to convert your single model to many-many similar models, such that their blend might be better than any single. \n\n**Context/motivation:** What is surprising (for me) that blending seems to give more improvement than for other similar competitions.\n(Even \"relatively weak\" models like Ridge and in general blending in that competition gives improvement which are somewhat bigger than expected).   Why it happens and is it really true ( or I am wrong) - that is another interesting question. \n\nSo one can try to use several tricks to produce from a  model  many similar, typically weaker than original, but with some diversity (and then - blend them all): \n\n**Tricks:**\n\n1) For NN - if you optimize correlation metric - try to use standard MSE or another \n\n2) For NN - predict not ALL targets but choose group of them and predict only them, \ntake many subgroups - then blend. (It requires metrics like MSE, probably, correlation will not work). \nAs subgroups it might be wise to take correlated groups of targets.\n(Correlated groups of targets can be found here: \nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02?scriptVersionId=106444735&cellId=54 )\n\n3) Instead of predicting targets directly one may try to use PCA and predict PCA[: , :N] for some N and \nthen use inverse transform to restore targets. One can try to use other dimensional reduction algorithms \neven non-linear like UMAP - we just need inverse_transform. \n\n4) Feature engineering - use various feature engineering constructions - to add different sets of features\nto your basic models. \n(For that particular competition - there are plenty datasets with ready to use features:\nsome along the ways described here: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/361025 )\n\nYou may try as a first step to select some important features by some feature importance algorithm,\nand use them. But one can do it on some subset of the data - thus changing subsets of the data - \nwill most probably will give another set of important features and so one can try to use thus obtained different geatures for your models and thus obtain plenty of the models. \n\n5) Use different random seeds - but it may not work - not enough diversity. \n\nPS\nSo in part it is motivated by Random Forest approach - you take one simple model like decision tree,\nbut then you make many-many models based on it - using different feature subsets, different sample subsets, etc...\nand then - blend them all.\n\nAny comments are welcome. ",
      "votes": 19
    },
    {
      "id": 2027623,
      "postDate": "2022-11-13T03:02:14.890Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2027625,
          "postDate": "2022-11-13T03:04:27.260Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2027623,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-13T03:02:14.890000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2027625,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-13T03:04:27.260000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2008271": "**Briefly - idea -**  how to convert your single model to many-many similar models, such that their blend might be better than any single. \n\n**Context/motivation:** What is surprising (for me) that blending seems to give more improvement than for other similar competitions.\n(Even \"relatively weak\" models like Ridge and in general blending in that competition gives improvement which are somewhat bigger than expected).   Why it happens and is it really true ( or I am wrong) - that is another interesting question. \n\nSo one can try to use several tricks to produce from a  model  many similar, typically weaker than original, but with some diversity (and then - blend them all): \n\n**Tricks:**\n\n1) For NN - if you optimize correlation metric - try to use standard MSE or another \n\n2) For NN - predict not ALL targets but choose group of them and predict only them, \ntake many subgroups - then blend. (It requires metrics like MSE, probably, correlation will not work). \nAs subgroups it might be wise to take correlated groups of targets.\n(Correlated groups of targets can be found here: \nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-targets-citeseq-02?scriptVersionId=106444735&cellId=54 )\n\n3) Instead of predicting targets directly one may try to use PCA and predict PCA[: , :N] for some N and \nthen use inverse transform to restore targets. One can try to use other dimensional reduction algorithms \neven non-linear like UMAP - we just need inverse_transform. \n\n4) Feature engineering - use various feature engineering constructions - to add different sets of features\nto your basic models. \n(For that particular competition - there are plenty datasets with ready to use features:\nsome along the ways described here: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/361025 )\n\nYou may try as a first step to select some important features by some feature importance algorithm,\nand use them. But one can do it on some subset of the data - thus changing subsets of the data - \nwill most probably will give another set of important features and so one can try to use thus obtained different geatures for your models and thus obtain plenty of the models. \n\n5) Use different random seeds - but it may not work - not enough diversity. \n\nPS\nSo in part it is motivated by Random Forest approach - you take one simple model like decision tree,\nbut then you make many-many models based on it - using different feature subsets, different sample subsets, etc...\nand then - blend them all.\n\nAny comments are welcome. ",
    "2027623": ""
  }
}