{
  "id": 337610,
  "title": "What models ensemble well?",
  "url": "/competitions/amex-default-prediction/discussion/337610",
  "author_name": "",
  "post_date": "2022-07-16T19:51:02.321066400Z",
  "votes": 46,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I posted <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/50827\" target=\"_blank\"><strong>a script</strong></a> long time ago that calculates correlations between models. It was for a different competition and will need to be adjusted slightly, but it should work.</p>\n<p>Generally speaking, models that have Kolmogorov-Smirnov statistic greater than 0.02 could be useful for ensembling. They should ensemble well if KS-stat &gt; 0.05, and should give a really nice score boost when KS-stat &gt; 0.1. Models don't even have to be made by different methods, although it will likely be helpful to add neural network and linear models to LGB/XGB models. Still, even two LGB models made from ~900 and ~700 features can be different enough, as shown below.</p>\n<pre><code> Column to be measured: prediction\n Pearson's correlation score: 0.993078\n Kendall's correlation score: 0.918780\n Spearman's correlation score: 0.990127\n Kolmogorov-Smirnov test:    KS-stat = 0.060161    p-value = 0.000e+00\n</code></pre>\n<p><strong>EDIT:</strong> A longer explanation for why diverse models ensemble better is given <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>here</strong></a>, also by yours truly.</p>",
  "messages": [
    {
      "id": "1858309",
      "postDate": "07/16/2022 19:51:02",
      "content": "<p>I posted <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/50827\" target=\"_blank\"><strong>a script</strong></a> long time ago that calculates correlations between models. It was for a different competition and will need to be adjusted slightly, but it should work.</p>\n<p>Generally speaking, models that have Kolmogorov-Smirnov statistic greater than 0.02 could be useful for ensembling. They should ensemble well if KS-stat &gt; 0.05, and should give a really nice score boost when KS-stat &gt; 0.1. Models don't even have to be made by different methods, although it will likely be helpful to add neural network and linear models to LGB/XGB models. Still, even two LGB models made from ~900 and ~700 features can be different enough, as shown below.</p>\n<pre><code> Column to be measured: prediction\n Pearson's correlation score: 0.993078\n Kendall's correlation score: 0.918780\n Spearman's correlation score: 0.990127\n Kolmogorov-Smirnov test:    KS-stat = 0.060161    p-value = 0.000e+00\n</code></pre>\n<p><strong>EDIT:</strong> A longer explanation for why diverse models ensemble better is given <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>here</strong></a>, also by yours truly.</p>",
      "rawMarkdown": "I posted [**a script**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/50827) long time ago that calculates correlations between models. It was for a different competition and will need to be adjusted slightly, but it should work.\n\nGenerally speaking, models that have Kolmogorov-Smirnov statistic greater than 0.02 could be useful for ensembling. They should ensemble well if KS-stat > 0.05, and should give a really nice score boost when KS-stat > 0.1. Models don't even have to be made by different methods, although it will likely be helpful to add neural network and linear models to LGB/XGB models. Still, even two LGB models made from ~900 and ~700 features can be different enough, as shown below.\n\n```\n Column to be measured: prediction\n Pearson's correlation score: 0.993078\n Kendall's correlation score: 0.918780\n Spearman's correlation score: 0.990127\n Kolmogorov-Smirnov test:    KS-stat = 0.060161    p-value = 0.000e+00\n\n```\n**EDIT:** A longer explanation for why diverse models ensemble better is given [**here**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058), also by yours truly.",
      "votes": null
    },
    {
      "id": "1858661",
      "postDate": "07/17/2022 05:43:29",
      "content": "<p>Very good post, thanks for sharing. This is a new learning for me. I hope to ensemble better with this approach.</p>",
      "rawMarkdown": "Very good post, thanks for sharing. This is a new learning for me. I hope to ensemble better with this approach.",
      "votes": null
    },
    {
      "id": "1858834",
      "postDate": "07/17/2022 08:41:36",
      "content": "<p>Thanks for posting! Will try this out.</p>",
      "rawMarkdown": "Thanks for posting! Will try this out.",
      "votes": null
    },
    {
      "id": "1858969",
      "postDate": "07/17/2022 09:49:39",
      "content": "<p>Thanks for sharing . One question (since the blog is not longer available), what exactly is passed in function <code>corr</code>? It says model1.csv and model2.csv, what exactly does these files have? </p>",
      "rawMarkdown": "Thanks for sharing . One question (since the blog is not longer available), what exactly is passed in function `corr`? It says model1.csv and model2.csv, what exactly does these files have?",
      "votes": null
    },
    {
      "id": "1859547",
      "postDate": "07/17/2022 17:49:31",
      "content": "<p>Those are your submission models in .csv format, where the first column (in this case <code>customer_ID</code>) is ignored. The original script compares several columns named <code>['toxic', 'severe_toxic', 'obscene', 'threat', 'insult', 'identity_hate']</code> which were targets in that competition. For this competition the only column to compare should be named <code>prediction</code>. </p>\n<p>It determines various correlation measures between <code>prediction</code> columns of the two models, followed by a KS 2-sample test which tries to ascertain whether they came from the same distribution. Models that are highly correlated will have a KS-stat value close to 0.</p>",
      "rawMarkdown": "Those are your submission models in .csv format, where the first column (in this case `customer_ID`) is ignored. The original script compares several columns named `['toxic', 'severe_toxic', 'obscene', 'threat', 'insult', 'identity_hate']` which were targets in that competition. For this competition the only column to compare should be named `prediction`. \n\nIt determines various correlation measures between `prediction` columns of the two models, followed by a KS 2-sample test which tries to ascertain whether they came from the same distribution. Models that are highly correlated will have a KS-stat value close to 0.",
      "votes": null
    },
    {
      "id": "1859566",
      "postDate": "07/17/2022 18:02:09",
      "content": "<p>So we are comparing correlation between the predictions of models. Understood. Thank you !</p>",
      "rawMarkdown": "So we are comparing correlation between the predictions of models. Understood. Thank you !",
      "votes": null
    },
    {
      "id": "1860890",
      "postDate": "07/18/2022 16:25:36",
      "content": "<p>Do use this and pearsons but does get tricky if you have lot of models and then u need to do permutation and combinations to mix them all together . In these case u need some iterative code that tries combinations and picks out the best based on these statistics .</p>",
      "rawMarkdown": "Do use this and pearsons but does get tricky if you have lot of models and then u need to do permutation and combinations to mix them all together . In these case u need some iterative code that tries combinations and picks out the best based on these statistics .",
      "votes": null
    },
    {
      "id": "1860927",
      "postDate": "07/18/2022 16:50:34",
      "content": "<p>Assuming that all models follow the naming convention <code>sub*.csv</code> and that python code is saved in <code>correlations.py</code>, we can make a <code>bash</code> script by pasting the following lines into text file <code>combinations.sh</code>:</p>\n<pre><code>#!/bin/bash\nfor i in sub*.csv\ndo\n  for j in sub*.csv\n  do\n    if [ \"$i\" \\&lt; \"$j\" ]\n    then\n     echo correlations.py $i $j &gt;&gt; correlations.txt\n    fi\n  done\ndone\n</code></pre>\n<p>Then make the script executive and run from a directory containing all models:</p>\n<pre><code>chmod +x combinations.sh\n./combinations.sh\n</code></pre>\n<p>The results will be saved in <code>correlations.txt</code>.</p>\n<p>The whole thing can be done within the python script by finding a list of files and putting them through a double loop.</p>",
      "rawMarkdown": "Assuming that all models follow the naming convention `sub*.csv` and that python code is saved in `correlations.py`, we can make a `bash` script by pasting the following lines into text file `combinations.sh`:\n\n```\n#!/bin/bash\nfor i in sub*.csv\ndo\n  for j in sub*.csv\n  do\n    if [ \"$i\" \\< \"$j\" ]\n    then\n     echo correlations.py $i $j >> correlations.txt\n    fi\n  done\ndone\n```\n\nThen make the script executive and run from a directory containing all models:\n\n```\nchmod +x combinations.sh\n./combinations.sh\n\n```\nThe results will be saved in `correlations.txt`.\n\nThe whole thing can be done within the python script by finding a list of files and putting them through a double loop.",
      "votes": null
    },
    {
      "id": "1860935",
      "postDate": "07/18/2022 16:53:41",
      "content": "<p>Yup I do it via something like <code>list(itertools.combinations(features, 2))</code></p>",
      "rawMarkdown": "Yup I do it via something like `list(itertools.combinations(features, 2)) `",
      "votes": null
    },
    {
      "id": "1861061",
      "postDate": "07/18/2022 18:41:54",
      "content": "<p>make a similar design pattern grid search.</p>",
      "rawMarkdown": "make a similar design pattern grid search.",
      "votes": null
    },
    {
      "id": "1861578",
      "postDate": "07/19/2022 05:27:24",
      "content": "<p>Great info!! Learned something new. Thanks.</p>",
      "rawMarkdown": "Great info!! Learned something new. Thanks.",
      "votes": null
    },
    {
      "id": "1862023",
      "postDate": "07/19/2022 12:01:19",
      "content": "<p>thanks sir <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
      "rawMarkdown": "thanks sir @tilii7",
      "votes": null
    },
    {
      "id": "1864141",
      "postDate": "07/20/2022 18:54:22",
      "content": "<p>If I understand correctly the reason you use KS testing between the models is to ensure they aren't overlapping in their outputs? Would these outputs be features? I am very new to ensemble learning.</p>",
      "rawMarkdown": "If I understand correctly the reason you use KS testing between the models is to ensure they aren't overlapping in their outputs? Would these outputs be features? I am very new to ensemble learning.",
      "votes": null
    },
    {
      "id": "1864163",
      "postDate": "07/20/2022 19:29:09",
      "content": "<p>We are testing model files that are being submitted for Kaggle evaluations - those 924621 predictions at the end of each modeling cycle. The idea is that if two models correlate very well - if their distributions are very similar - they will not give a boost after ensembling.</p>",
      "rawMarkdown": "We are testing model files that are being submitted for Kaggle evaluations - those 924621 predictions at the end of each modeling cycle. The idea is that if two models correlate very well - if their distributions are very similar - they will not give a boost after ensembling.",
      "votes": null
    },
    {
      "id": "1864176",
      "postDate": "07/20/2022 19:49:13",
      "content": "<p>This is because they are explaining the same information? There is no benefit to having two models that are always in agreement?</p>",
      "rawMarkdown": "This is because they are explaining the same information? There is no benefit to having two models that are always in agreement?",
      "votes": null
    },
    {
      "id": "1864201",
      "postDate": "07/20/2022 20:24:40",
      "content": "<p>All good models are in general agreement, but they shouldn't be too similar for the ensembling to work well.</p>\n<p>Here is a toy example with two models that have 3 predictions each:</p>\n<pre><code>[0.01, 0.3, 0.9]\n[0.011, 0.305, 0.895]\n</code></pre>\n<p>Let's say that targets we are trying to predict are <code>[0, 0, 1]</code>. No matter how you combine the two models, whether by simple average or some kind of weighting (say, <code>0.9*model1 + 0.1*model2</code>), you will always end up with a model that is hardly any better because the original two predictions are near-identical.</p>\n<p>Let's say that instead we have these two quite different models to combine:</p>\n<pre><code>[0.015, 0.31, 0.935]\n[0.005, 0.255, 0.91]\n</code></pre>\n<p>In this case the first model is better at predicting target 1, while the second model is better at predicting target 0. When combined they will give a model that should be significantly better than either one alone. So we are trying to find those models that differ enough as to be able to assemble well.</p>\n<p>I wrote a longer explanation about the need for model diversity <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>here</strong></a>.</p>",
      "rawMarkdown": "All good models are in general agreement, but they shouldn't be too similar for the ensembling to work well.\n\nHere is a toy example with two models that have 3 predictions each:\n\n    [0.01, 0.3, 0.9]\n    [0.011, 0.305, 0.895]\n\nLet's say that targets we are trying to predict are `[0, 0, 1]`. No matter how you combine the two models, whether by simple average or some kind of weighting (say, `0.9*model1 + 0.1*model2`), you will always end up with a model that is hardly any better because the original two predictions are near-identical.\n\nLet's say that instead we have these two quite different models to combine:\n\n    [0.015, 0.31, 0.935]\n    [0.005, 0.255, 0.91]\n\nIn this case the first model is better at predicting target 1, while the second model is better at predicting target 0. When combined they will give a model that should be significantly better than either one alone. So we are trying to find those models that differ enough as to be able to assemble well.\n\nI wrote a longer explanation about the need for model diversity [**here**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1858661,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/17/2022 05:43:29",
      "content": "<p>Very good post, thanks for sharing. This is a new learning for me. I hope to ensemble better with this approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1858834,
      "author_name": "nijianzhang",
      "author_url": "",
      "post_date": "07/17/2022 08:41:36",
      "content": "<p>Thanks for posting! Will try this out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1858969,
      "author_name": "nitishraj",
      "author_url": "",
      "post_date": "07/17/2022 09:49:39",
      "content": "<p>Thanks for sharing . One question (since the blog is not longer available), what exactly is passed in function <code>corr</code>? It says model1.csv and model2.csv, what exactly does these files have? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1859547,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/17/2022 17:49:31",
          "content": "<p>Those are your submission models in .csv format, where the first column (in this case <code>customer_ID</code>) is ignored. The original script compares several columns named <code>['toxic', 'severe_toxic', 'obscene', 'threat', 'insult', 'identity_hate']</code> which were targets in that competition. For this competition the only column to compare should be named <code>prediction</code>. </p>\n<p>It determines various correlation measures between <code>prediction</code> columns of the two models, followed by a KS 2-sample test which tries to ascertain whether they came from the same distribution. Models that are highly correlated will have a KS-stat value close to 0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1859566,
          "author_name": "nitishraj",
          "author_url": "",
          "post_date": "07/17/2022 18:02:09",
          "content": "<p>So we are comparing correlation between the predictions of models. Understood. Thank you !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1860890,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/18/2022 16:25:36",
      "content": "<p>Do use this and pearsons but does get tricky if you have lot of models and then u need to do permutation and combinations to mix them all together . In these case u need some iterative code that tries combinations and picks out the best based on these statistics .</p>",
      "votes": null,
      "replies": [
        {
          "id": 1860927,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/18/2022 16:50:34",
          "content": "<p>Assuming that all models follow the naming convention <code>sub*.csv</code> and that python code is saved in <code>correlations.py</code>, we can make a <code>bash</code> script by pasting the following lines into text file <code>combinations.sh</code>:</p>\n<pre><code>#!/bin/bash\nfor i in sub*.csv\ndo\n  for j in sub*.csv\n  do\n    if [ \"$i\" \\&lt; \"$j\" ]\n    then\n     echo correlations.py $i $j &gt;&gt; correlations.txt\n    fi\n  done\ndone\n</code></pre>\n<p>Then make the script executive and run from a directory containing all models:</p>\n<pre><code>chmod +x combinations.sh\n./combinations.sh\n</code></pre>\n<p>The results will be saved in <code>correlations.txt</code>.</p>\n<p>The whole thing can be done within the python script by finding a list of files and putting them through a double loop.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1860935,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "07/18/2022 16:53:41",
          "content": "<p>Yup I do it via something like <code>list(itertools.combinations(features, 2))</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1861061,
      "author_name": "zhehaoliang",
      "author_url": "",
      "post_date": "07/18/2022 18:41:54",
      "content": "<p>make a similar design pattern grid search.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1861578,
      "author_name": "cid007",
      "author_url": "",
      "post_date": "07/19/2022 05:27:24",
      "content": "<p>Great info!! Learned something new. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1862023,
      "author_name": "bilalsuppal",
      "author_url": "",
      "post_date": "07/19/2022 12:01:19",
      "content": "<p>thanks sir <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1864141,
      "author_name": "ianolmstead",
      "author_url": "",
      "post_date": "07/20/2022 18:54:22",
      "content": "<p>If I understand correctly the reason you use KS testing between the models is to ensure they aren't overlapping in their outputs? Would these outputs be features? I am very new to ensemble learning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1864163,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 19:29:09",
          "content": "<p>We are testing model files that are being submitted for Kaggle evaluations - those 924621 predictions at the end of each modeling cycle. The idea is that if two models correlate very well - if their distributions are very similar - they will not give a boost after ensembling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1864176,
          "author_name": "ianolmstead",
          "author_url": "",
          "post_date": "07/20/2022 19:49:13",
          "content": "<p>This is because they are explaining the same information? There is no benefit to having two models that are always in agreement?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1864201,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 20:24:40",
          "content": "<p>All good models are in general agreement, but they shouldn't be too similar for the ensembling to work well.</p>\n<p>Here is a toy example with two models that have 3 predictions each:</p>\n<pre><code>[0.01, 0.3, 0.9]\n[0.011, 0.305, 0.895]\n</code></pre>\n<p>Let's say that targets we are trying to predict are <code>[0, 0, 1]</code>. No matter how you combine the two models, whether by simple average or some kind of weighting (say, <code>0.9*model1 + 0.1*model2</code>), you will always end up with a model that is hardly any better because the original two predictions are near-identical.</p>\n<p>Let's say that instead we have these two quite different models to combine:</p>\n<pre><code>[0.015, 0.31, 0.935]\n[0.005, 0.255, 0.91]\n</code></pre>\n<p>In this case the first model is better at predicting target 1, while the second model is better at predicting target 0. When combined they will give a model that should be significantly better than either one alone. So we are trying to find those models that differ enough as to be able to assemble well.</p>\n<p>I wrote a longer explanation about the need for model diversity <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>here</strong></a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1858309": "I posted [**a script**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/50827) long time ago that calculates correlations between models. It was for a different competition and will need to be adjusted slightly, but it should work.\n\nGenerally speaking, models that have Kolmogorov-Smirnov statistic greater than 0.02 could be useful for ensembling. They should ensemble well if KS-stat > 0.05, and should give a really nice score boost when KS-stat > 0.1. Models don't even have to be made by different methods, although it will likely be helpful to add neural network and linear models to LGB/XGB models. Still, even two LGB models made from ~900 and ~700 features can be different enough, as shown below.\n\n```\n Column to be measured: prediction\n Pearson's correlation score: 0.993078\n Kendall's correlation score: 0.918780\n Spearman's correlation score: 0.990127\n Kolmogorov-Smirnov test:    KS-stat = 0.060161    p-value = 0.000e+00\n\n```\n**EDIT:** A longer explanation for why diverse models ensemble better is given [**here**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058), also by yours truly.",
    "1858661": "Very good post, thanks for sharing. This is a new learning for me. I hope to ensemble better with this approach.",
    "1858834": "Thanks for posting! Will try this out.",
    "1858969": "Thanks for sharing . One question (since the blog is not longer available), what exactly is passed in function `corr`? It says model1.csv and model2.csv, what exactly does these files have?",
    "1859547": "Those are your submission models in .csv format, where the first column (in this case `customer_ID`) is ignored. The original script compares several columns named `['toxic', 'severe_toxic', 'obscene', 'threat', 'insult', 'identity_hate']` which were targets in that competition. For this competition the only column to compare should be named `prediction`. \n\nIt determines various correlation measures between `prediction` columns of the two models, followed by a KS 2-sample test which tries to ascertain whether they came from the same distribution. Models that are highly correlated will have a KS-stat value close to 0.",
    "1859566": "So we are comparing correlation between the predictions of models. Understood. Thank you !",
    "1860890": "Do use this and pearsons but does get tricky if you have lot of models and then u need to do permutation and combinations to mix them all together . In these case u need some iterative code that tries combinations and picks out the best based on these statistics .",
    "1860927": "Assuming that all models follow the naming convention `sub*.csv` and that python code is saved in `correlations.py`, we can make a `bash` script by pasting the following lines into text file `combinations.sh`:\n\n```\n#!/bin/bash\nfor i in sub*.csv\ndo\n  for j in sub*.csv\n  do\n    if [ \"$i\" \\< \"$j\" ]\n    then\n     echo correlations.py $i $j >> correlations.txt\n    fi\n  done\ndone\n```\n\nThen make the script executive and run from a directory containing all models:\n\n```\nchmod +x combinations.sh\n./combinations.sh\n\n```\nThe results will be saved in `correlations.txt`.\n\nThe whole thing can be done within the python script by finding a list of files and putting them through a double loop.",
    "1860935": "Yup I do it via something like `list(itertools.combinations(features, 2)) `",
    "1861061": "make a similar design pattern grid search.",
    "1861578": "Great info!! Learned something new. Thanks.",
    "1862023": "thanks sir @tilii7",
    "1864141": "If I understand correctly the reason you use KS testing between the models is to ensure they aren't overlapping in their outputs? Would these outputs be features? I am very new to ensemble learning.",
    "1864163": "We are testing model files that are being submitted for Kaggle evaluations - those 924621 predictions at the end of each modeling cycle. The idea is that if two models correlate very well - if their distributions are very similar - they will not give a boost after ensembling.",
    "1864176": "This is because they are explaining the same information? There is no benefit to having two models that are always in agreement?",
    "1864201": "All good models are in general agreement, but they shouldn't be too similar for the ensembling to work well.\n\nHere is a toy example with two models that have 3 predictions each:\n\n    [0.01, 0.3, 0.9]\n    [0.011, 0.305, 0.895]\n\nLet's say that targets we are trying to predict are `[0, 0, 1]`. No matter how you combine the two models, whether by simple average or some kind of weighting (say, `0.9*model1 + 0.1*model2`), you will always end up with a model that is hardly any better because the original two predictions are near-identical.\n\nLet's say that instead we have these two quite different models to combine:\n\n    [0.015, 0.31, 0.935]\n    [0.005, 0.255, 0.91]\n\nIn this case the first model is better at predicting target 1, while the second model is better at predicting target 0. When combined they will give a model that should be significantly better than either one alone. So we are trying to find those models that differ enough as to be able to assemble well.\n\nI wrote a longer explanation about the need for model diversity [**here**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058)."
  },
  "source": "meta"
}