{
  "id": 363230,
  "title": "Terrible model -  might be useful in blend ? (KernelRidge - last year winner )",
  "url": "/competitions/open-problems-multimodal/discussion/363230",
  "author_name": "",
  "post_date": "2022-10-31T16:10:40.022436700Z",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Tried to create a somewhat tutorial notebook, on different models, and effect of blend.<br>\nAnd unexpectedly  observed the strange thing -  the TERRIBLE model (r2&lt;0, mse - enourmous) was useful in blend.<br>\nand improved scores even on completely unseen data, which is 10 times large than data where everything was trained or blended. </p>\n<p><strong>Question:</strong> Is it common phenomena ? How can one exploit it  ? Typically we search for the best models and blend them,<br>\nmay be we should also on bad models ? but how bad ? What can be criteria ? </p>\n<p>It might be - I made some  mistake , if so sorry for bothering. </p>\n<p>The notebook, section \"Blend\" cell \"In 398\", here is direct link to cell: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109657383&amp;cellId=85\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109657383&amp;cellId=85</a></p>\n<p>The terrible model is Kernel Ridge (by the way winner of the last year competition! - but seems cannot be made working this year.)  <br>\nThe scores of blend on different parts<br>\n(Here we blend) R2: 0.550912372396922 MSE: 18.294795712224616<br>\nUnseen Test2 R2: 0.4421005286800104 MSE: 17.280883838459452<br>\nUnseen and 10 times bigger: OOP R2: 0.47510925408331395 OOP MSE: 20.026047575114244<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F547eb5af0c6c1d591d85d5bfefe79128%2Fphoto_2022-10-31_16-58-16.jpg?generation=1667232386028992&amp;alt=media\" alt=\"\"></p>\n<p>The only reasonable thing - that it is highly uncorrelated with the other solutions - that is good for blend of course:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F3790b4794a26d57c006479351fffeaf4%2Fphoto_2022-10-31_16-58-28.jpg?generation=1667232604164340&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2011519",
      "postDate": "10/31/2022 16:10:40",
      "content": "<p>Tried to create a somewhat tutorial notebook, on different models, and effect of blend.<br>\nAnd unexpectedly  observed the strange thing -  the TERRIBLE model (r2&lt;0, mse - enourmous) was useful in blend.<br>\nand improved scores even on completely unseen data, which is 10 times large than data where everything was trained or blended. </p>\n<p><strong>Question:</strong> Is it common phenomena ? How can one exploit it  ? Typically we search for the best models and blend them,<br>\nmay be we should also on bad models ? but how bad ? What can be criteria ? </p>\n<p>It might be - I made some  mistake , if so sorry for bothering. </p>\n<p>The notebook, section \"Blend\" cell \"In 398\", here is direct link to cell: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109657383&amp;cellId=85\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109657383&amp;cellId=85</a></p>\n<p>The terrible model is Kernel Ridge (by the way winner of the last year competition! - but seems cannot be made working this year.)  <br>\nThe scores of blend on different parts<br>\n(Here we blend) R2: 0.550912372396922 MSE: 18.294795712224616<br>\nUnseen Test2 R2: 0.4421005286800104 MSE: 17.280883838459452<br>\nUnseen and 10 times bigger: OOP R2: 0.47510925408331395 OOP MSE: 20.026047575114244<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F547eb5af0c6c1d591d85d5bfefe79128%2Fphoto_2022-10-31_16-58-16.jpg?generation=1667232386028992&amp;alt=media\" alt=\"\"></p>\n<p>The only reasonable thing - that it is highly uncorrelated with the other solutions - that is good for blend of course:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F3790b4794a26d57c006479351fffeaf4%2Fphoto_2022-10-31_16-58-28.jpg?generation=1667232604164340&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Tried to create a somewhat tutorial notebook, on different models, and effect of blend.\nAnd unexpectedly  observed the strange thing -  the TERRIBLE model (r2<0, mse - enourmous) was useful in blend.\nand improved scores even on completely unseen data, which is 10 times large than data where everything was trained or blended. \n\n**Question:** Is it common phenomena ? How can one exploit it  ? Typically we search for the best models and blend them,\nmay be we should also on bad models ? but how bad ? What can be criteria ? \n\nIt might be - I made some  mistake , if so sorry for bothering. \n\nThe notebook, section \"Blend\" cell \"In 398\", here is direct link to cell: \nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109657383&cellId=85\n\nThe terrible model is Kernel Ridge (by the way winner of the last year competition! - but seems cannot be made working this year.)  \nThe scores of blend on different parts\n(Here we blend) R2: 0.550912372396922 MSE: 18.294795712224616\nUnseen Test2 R2: 0.4421005286800104 MSE: 17.280883838459452\nUnseen and 10 times bigger: OOP R2: 0.47510925408331395 OOP MSE: 20.026047575114244\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F547eb5af0c6c1d591d85d5bfefe79128%2Fphoto_2022-10-31_16-58-16.jpg?generation=1667232386028992&alt=media)\n\nThe only reasonable thing - that it is highly uncorrelated with the other solutions - that is good for blend of course:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F3790b4794a26d57c006479351fffeaf4%2Fphoto_2022-10-31_16-58-28.jpg?generation=1667232604164340&alt=media)",
      "votes": null
    },
    {
      "id": "2011891",
      "postDate": "10/31/2022 23:12:21",
      "content": "<p>Years ago when Random Forests first came to be, there was a lot of research around diversity of combined classifiers.  <a href=\"https://link.springer.com/article/10.1023/A:1022859003006\" target=\"_blank\">This paper</a> is particularly good.</p>\n<p>If I recall correctly, the premise of ensemble models is to combine a collection of weak, diverse base learners.  Random Forest achieves that diversity by boostrapping the samples and selecting a subset of the features for each tree.  If there was little or no diversity, adding new classifiers to the ensemble would not provide any additional information, however, if the classifier is too weak, it may simply provide noise.</p>\n<p>I've often wondered if you could train specifically for diversity.  Again, Random Forest created diversity by bootstrapping samples and subsetting features, but it does not have a diversity objective function.  XGBoost functions to identify the gradient of the errors, so perhaps diversity is an emergent property.  </p>\n<p>I would anticipate that the best combination of classifiers would mostly model the \"easy\" samples equally well.  For difficult regions of the sample space, different classifiers would specialize in those regions with some redundancy between classifiers, but not at the risk of becoming completely redundant with other classifiers.  I picture this as a sort of Venn diagram in which each classifier's circle overlaps with others in the easy sample space, but has unique coverage in areas in the difficult to model spaces.</p>\n<p>The diagram below from <a href=\"https://nl.pinterest.com/pin/90283167510252861/\" target=\"_blank\">vecteezy.com</a> captures my thoughts about classifier diversity and specialization really well.  Each classifier overlaps in the center, with additional overlap with other classifiers in around the periphery and some unique specializations.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff77ec00d07c152e43fb38c5ccf8cac72%2FScreen%20Shot%202022-10-31%20at%205.13.03%20PM.png?generation=1667257995680991&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Years ago when Random Forests first came to be, there was a lot of research around diversity of combined classifiers.  [This paper](https://link.springer.com/article/10.1023/A:1022859003006) is particularly good.\n\nIf I recall correctly, the premise of ensemble models is to combine a collection of weak, diverse base learners.  Random Forest achieves that diversity by boostrapping the samples and selecting a subset of the features for each tree.  If there was little or no diversity, adding new classifiers to the ensemble would not provide any additional information, however, if the classifier is too weak, it may simply provide noise.\n\nI've often wondered if you could train specifically for diversity.  Again, Random Forest created diversity by bootstrapping samples and subsetting features, but it does not have a diversity objective function.  XGBoost functions to identify the gradient of the errors, so perhaps diversity is an emergent property.  \n\nI would anticipate that the best combination of classifiers would mostly model the \"easy\" samples equally well.  For difficult regions of the sample space, different classifiers would specialize in those regions with some redundancy between classifiers, but not at the risk of becoming completely redundant with other classifiers.  I picture this as a sort of Venn diagram in which each classifier's circle overlaps with others in the easy sample space, but has unique coverage in areas in the difficult to model spaces.\n\nThe diagram below from [vecteezy.com](https://nl.pinterest.com/pin/90283167510252861/) captures my thoughts about classifier diversity and specialization really well.  Each classifier overlaps in the center, with additional overlap with other classifiers in around the periphery and some unique specializations.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff77ec00d07c152e43fb38c5ccf8cac72%2FScreen%20Shot%202022-10-31%20at%205.13.03%20PM.png?generation=1667257995680991&alt=media)",
      "votes": null
    },
    {
      "id": "2011911",
      "postDate": "11/01/2022 00:09:40",
      "content": "<p>In my experience, the blending works in following order;<br>\n①good and highly different models &gt; ②not good but highly different models ≒ ③good and slightly different models &gt;&gt; ④not good and not different models. <br>\nThe kernel ridge example is ② pattern. But that is because your base model is not strong itself. I think if the base model is the published strong nn, kernel ridge falls into ④ pattern, and the blending doesn't work.</p>",
      "rawMarkdown": "In my experience, the blending works in following order;\n①good and highly different models > ②not good but highly different models ≒ ③good and slightly different models >> ④not good and not different models. \nThe kernel ridge example is ② pattern. But that is because your base model is not strong itself. I think if the base model is the published strong nn, kernel ridge falls into ④ pattern, and the blending doesn't work.",
      "votes": null
    },
    {
      "id": "2012632",
      "postDate": "11/01/2022 10:09:39",
      "content": "<p>Thanks for your comment ! I appreciate it ! <br>\nHowever are there some quantitative measures of good and not good ? <br>\nI am suprised that KernelRidge is not simply \"not good\", but it is really terrible - mse - is 10 times higher than for other models, r2&lt;0 ! <br>\nPS<br>\nConcerning Ridge is not good enough  - pay attention that I downsampled the dataset and considered JUST ONE Target - CD31, not really sure one can do better than Ridge , by any model, not only NN, at least playing with boostings - cannot overcome Ridge. And NN has almost never advantage over boostings  for tabular data , except the multi-target case. But here is ONE Target example. </p>",
      "rawMarkdown": "Thanks for your comment ! I appreciate it ! \nHowever are there some quantitative measures of good and not good ? \nI am suprised that KernelRidge is not simply \"not good\", but it is really terrible - mse - is 10 times higher than for other models, r2<0 ! \nPS\nConcerning Ridge is not good enough  - pay attention that I downsampled the dataset and considered JUST ONE Target - CD31, not really sure one can do better than Ridge , by any model, not only NN, at least playing with boostings - cannot overcome Ridge. And NN has almost never advantage over boostings  for tabular data , except the multi-target case. But here is ONE Target example.",
      "votes": null
    },
    {
      "id": "2020715",
      "postDate": "11/07/2022 18:53:48",
      "content": "<p>May be this is somehow similar to TensorFlow's Dropout layer, when we \"damage\" the model but get some better performance because the model cannot learn too perfect to learn the noise from the train set.</p>",
      "rawMarkdown": "May be this is somehow similar to TensorFlow's Dropout layer, when we \"damage\" the model but get some better performance because the model cannot learn too perfect to learn the noise from the train set.",
      "votes": null
    },
    {
      "id": "2020816",
      "postDate": "11/07/2022 19:52:37",
      "content": "<p>Yes, I've read that dropout is equivalent to creating an ensemble of neural networks consisting of all possible dropout networks.  I'll try to find the paper and post here if I find it.</p>",
      "rawMarkdown": "Yes, I've read that dropout is equivalent to creating an ensemble of neural networks consisting of all possible dropout networks.  I'll try to find the paper and post here if I find it.",
      "votes": null
    },
    {
      "id": "2020913",
      "postDate": "11/07/2022 21:04:26",
      "content": "<p>I also tried KernelRidge but I cannot make it outperforms even simple Ridge Regression… </p>\n<p>Can I ask based on your work, what is the overall best models?</p>",
      "rawMarkdown": "I also tried KernelRidge but I cannot make it outperforms even simple Ridge Regression... \n\nCan I ask based on your work, what is the overall best models?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2011891,
      "author_name": "kirkdco",
      "author_url": "",
      "post_date": "10/31/2022 23:12:21",
      "content": "<p>Years ago when Random Forests first came to be, there was a lot of research around diversity of combined classifiers.  <a href=\"https://link.springer.com/article/10.1023/A:1022859003006\" target=\"_blank\">This paper</a> is particularly good.</p>\n<p>If I recall correctly, the premise of ensemble models is to combine a collection of weak, diverse base learners.  Random Forest achieves that diversity by boostrapping the samples and selecting a subset of the features for each tree.  If there was little or no diversity, adding new classifiers to the ensemble would not provide any additional information, however, if the classifier is too weak, it may simply provide noise.</p>\n<p>I've often wondered if you could train specifically for diversity.  Again, Random Forest created diversity by bootstrapping samples and subsetting features, but it does not have a diversity objective function.  XGBoost functions to identify the gradient of the errors, so perhaps diversity is an emergent property.  </p>\n<p>I would anticipate that the best combination of classifiers would mostly model the \"easy\" samples equally well.  For difficult regions of the sample space, different classifiers would specialize in those regions with some redundancy between classifiers, but not at the risk of becoming completely redundant with other classifiers.  I picture this as a sort of Venn diagram in which each classifier's circle overlaps with others in the easy sample space, but has unique coverage in areas in the difficult to model spaces.</p>\n<p>The diagram below from <a href=\"https://nl.pinterest.com/pin/90283167510252861/\" target=\"_blank\">vecteezy.com</a> captures my thoughts about classifier diversity and specialization really well.  Each classifier overlaps in the center, with additional overlap with other classifiers in around the periphery and some unique specializations.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff77ec00d07c152e43fb38c5ccf8cac72%2FScreen%20Shot%202022-10-31%20at%205.13.03%20PM.png?generation=1667257995680991&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2011911,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "11/01/2022 00:09:40",
      "content": "<p>In my experience, the blending works in following order;<br>\n①good and highly different models &gt; ②not good but highly different models ≒ ③good and slightly different models &gt;&gt; ④not good and not different models. <br>\nThe kernel ridge example is ② pattern. But that is because your base model is not strong itself. I think if the base model is the published strong nn, kernel ridge falls into ④ pattern, and the blending doesn't work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2012632,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "11/01/2022 10:09:39",
          "content": "<p>Thanks for your comment ! I appreciate it ! <br>\nHowever are there some quantitative measures of good and not good ? <br>\nI am suprised that KernelRidge is not simply \"not good\", but it is really terrible - mse - is 10 times higher than for other models, r2&lt;0 ! <br>\nPS<br>\nConcerning Ridge is not good enough  - pay attention that I downsampled the dataset and considered JUST ONE Target - CD31, not really sure one can do better than Ridge , by any model, not only NN, at least playing with boostings - cannot overcome Ridge. And NN has almost never advantage over boostings  for tabular data , except the multi-target case. But here is ONE Target example. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2020715,
      "author_name": "alexandrgusev",
      "author_url": "",
      "post_date": "11/07/2022 18:53:48",
      "content": "<p>May be this is somehow similar to TensorFlow's Dropout layer, when we \"damage\" the model but get some better performance because the model cannot learn too perfect to learn the noise from the train set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2020816,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "11/07/2022 19:52:37",
          "content": "<p>Yes, I've read that dropout is equivalent to creating an ensemble of neural networks consisting of all possible dropout networks.  I'll try to find the paper and post here if I find it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2020913,
      "author_name": "jinyang18",
      "author_url": "",
      "post_date": "11/07/2022 21:04:26",
      "content": "<p>I also tried KernelRidge but I cannot make it outperforms even simple Ridge Regression… </p>\n<p>Can I ask based on your work, what is the overall best models?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2011519": "Tried to create a somewhat tutorial notebook, on different models, and effect of blend.\nAnd unexpectedly  observed the strange thing -  the TERRIBLE model (r2<0, mse - enourmous) was useful in blend.\nand improved scores even on completely unseen data, which is 10 times large than data where everything was trained or blended. \n\n**Question:** Is it common phenomena ? How can one exploit it  ? Typically we search for the best models and blend them,\nmay be we should also on bad models ? but how bad ? What can be criteria ? \n\nIt might be - I made some  mistake , if so sorry for bothering. \n\nThe notebook, section \"Blend\" cell \"In 398\", here is direct link to cell: \nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-with-playground-fold?scriptVersionId=109657383&cellId=85\n\nThe terrible model is Kernel Ridge (by the way winner of the last year competition! - but seems cannot be made working this year.)  \nThe scores of blend on different parts\n(Here we blend) R2: 0.550912372396922 MSE: 18.294795712224616\nUnseen Test2 R2: 0.4421005286800104 MSE: 17.280883838459452\nUnseen and 10 times bigger: OOP R2: 0.47510925408331395 OOP MSE: 20.026047575114244\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F547eb5af0c6c1d591d85d5bfefe79128%2Fphoto_2022-10-31_16-58-16.jpg?generation=1667232386028992&alt=media)\n\nThe only reasonable thing - that it is highly uncorrelated with the other solutions - that is good for blend of course:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F3790b4794a26d57c006479351fffeaf4%2Fphoto_2022-10-31_16-58-28.jpg?generation=1667232604164340&alt=media)",
    "2011891": "Years ago when Random Forests first came to be, there was a lot of research around diversity of combined classifiers.  [This paper](https://link.springer.com/article/10.1023/A:1022859003006) is particularly good.\n\nIf I recall correctly, the premise of ensemble models is to combine a collection of weak, diverse base learners.  Random Forest achieves that diversity by boostrapping the samples and selecting a subset of the features for each tree.  If there was little or no diversity, adding new classifiers to the ensemble would not provide any additional information, however, if the classifier is too weak, it may simply provide noise.\n\nI've often wondered if you could train specifically for diversity.  Again, Random Forest created diversity by bootstrapping samples and subsetting features, but it does not have a diversity objective function.  XGBoost functions to identify the gradient of the errors, so perhaps diversity is an emergent property.  \n\nI would anticipate that the best combination of classifiers would mostly model the \"easy\" samples equally well.  For difficult regions of the sample space, different classifiers would specialize in those regions with some redundancy between classifiers, but not at the risk of becoming completely redundant with other classifiers.  I picture this as a sort of Venn diagram in which each classifier's circle overlaps with others in the easy sample space, but has unique coverage in areas in the difficult to model spaces.\n\nThe diagram below from [vecteezy.com](https://nl.pinterest.com/pin/90283167510252861/) captures my thoughts about classifier diversity and specialization really well.  Each classifier overlaps in the center, with additional overlap with other classifiers in around the periphery and some unique specializations.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff77ec00d07c152e43fb38c5ccf8cac72%2FScreen%20Shot%202022-10-31%20at%205.13.03%20PM.png?generation=1667257995680991&alt=media)",
    "2011911": "In my experience, the blending works in following order;\n①good and highly different models > ②not good but highly different models ≒ ③good and slightly different models >> ④not good and not different models. \nThe kernel ridge example is ② pattern. But that is because your base model is not strong itself. I think if the base model is the published strong nn, kernel ridge falls into ④ pattern, and the blending doesn't work.",
    "2012632": "Thanks for your comment ! I appreciate it ! \nHowever are there some quantitative measures of good and not good ? \nI am suprised that KernelRidge is not simply \"not good\", but it is really terrible - mse - is 10 times higher than for other models, r2<0 ! \nPS\nConcerning Ridge is not good enough  - pay attention that I downsampled the dataset and considered JUST ONE Target - CD31, not really sure one can do better than Ridge , by any model, not only NN, at least playing with boostings - cannot overcome Ridge. And NN has almost never advantage over boostings  for tabular data , except the multi-target case. But here is ONE Target example.",
    "2020715": "May be this is somehow similar to TensorFlow's Dropout layer, when we \"damage\" the model but get some better performance because the model cannot learn too perfect to learn the noise from the train set.",
    "2020816": "Yes, I've read that dropout is equivalent to creating an ensemble of neural networks consisting of all possible dropout networks.  I'll try to find the paper and post here if I find it.",
    "2020913": "I also tried KernelRidge but I cannot make it outperforms even simple Ridge Regression... \n\nCan I ask based on your work, what is the overall best models?"
  },
  "source": "meta"
}