{
  "id": 199458,
  "title": "CV vs LB Comparison",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/199458",
  "author_name": "",
  "post_date": "2020-11-25T19:23:43.983226200Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Anyone interested in sharing CV and LB score?</p>\n<p>Following is my CV vs LB comparison:</p>\n<p>I am using 5-fold LGB model.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n<th>features</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>388529</td>\n<td>4827594</td>\n<td>223</td>\n<td></td>\n</tr>\n<tr>\n<td>387870</td>\n<td>4759488</td>\n<td>183</td>\n<td></td>\n</tr>\n<tr>\n<td>399560</td>\n<td>4833849</td>\n<td>143</td>\n<td></td>\n</tr>\n<tr>\n<td>401977</td>\n<td>4942119</td>\n<td>123</td>\n<td></td>\n</tr>\n<tr>\n<td>480388</td>\n<td>5208812</td>\n<td>100</td>\n<td>using <a href=\"https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a></td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "1091120",
      "postDate": "11/25/2020 19:23:43",
      "content": "<p>Hi all,</p>\n<p>Anyone interested in sharing CV and LB score?</p>\n<p>Following is my CV vs LB comparison:</p>\n<p>I am using 5-fold LGB model.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n<th>features</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>388529</td>\n<td>4827594</td>\n<td>223</td>\n<td></td>\n</tr>\n<tr>\n<td>387870</td>\n<td>4759488</td>\n<td>183</td>\n<td></td>\n</tr>\n<tr>\n<td>399560</td>\n<td>4833849</td>\n<td>143</td>\n<td></td>\n</tr>\n<tr>\n<td>401977</td>\n<td>4942119</td>\n<td>123</td>\n<td></td>\n</tr>\n<tr>\n<td>480388</td>\n<td>5208812</td>\n<td>100</td>\n<td>using <a href=\"https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft\" target=\"_blank\">https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft</a></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Hi all,\n\nAnyone interested in sharing CV and LB score?\n\nFollowing is my CV vs LB comparison:\n\nI am using 5-fold LGB model.\n| CV | LB | features | Note\n| --- | --- | --- | --- |\n| 388529 | 4827594  | 223 |\n| 387870 | 4759488  | 183 |\n| 399560 | 4833849  | 143 |\n| 401977 | 4942119  | 123 |\n| 480388 | 5208812  | 100 | using https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft",
      "votes": null
    },
    {
      "id": "1091771",
      "postDate": "11/26/2020 09:14:00",
      "content": "<p>I have CV: 2306182 and LB: 4797783 at the 15th position currently, I use 5-fold CV and LGBM as well. I experience huge train/validation gaps most of the time that translates to a big gap between CV and LB errors, though the errors correlate well. I haven't figured out how to reduce overfitting as much as you did, I tried to remove features and currently use only 120.<br>\nDo you use a simple shuffled 5-fold CV as validation strategy? </p>",
      "rawMarkdown": "I have CV: 2306182 and LB: 4797783 at the 15th position currently, I use 5-fold CV and LGBM as well. I experience huge train/validation gaps most of the time that translates to a big gap between CV and LB errors, though the errors correlate well. I haven't figured out how to reduce overfitting as much as you did, I tried to remove features and currently use only 120.\nDo you use a simple shuffled 5-fold CV as validation strategy?",
      "votes": null
    },
    {
      "id": "1093494",
      "postDate": "11/27/2020 19:20:57",
      "content": "<p>yeah i am using simple 5 fold CV as validation strategy and using covariate shift to check for features.</p>",
      "rawMarkdown": "yeah i am using simple 5 fold CV as validation strategy and using covariate shift to check for features.",
      "votes": null
    },
    {
      "id": "1095867",
      "postDate": "11/30/2020 02:49:55",
      "content": "<p>CV:3190586.8421 LB: 5688433 by using LGBM and CV:2425092.204 LB: 4836750 by using CatBoost (5-fold) <br>\nI think test and train are very different. but I don't have any idea…. </p>",
      "rawMarkdown": "CV:3190586.8421 LB: 5688433 by using LGBM and CV:2425092.204 LB: 4836750 by using CatBoost (5-fold) \nI think test and train are very different. but I don't have any idea....",
      "votes": null
    },
    {
      "id": "1095901",
      "postDate": "11/30/2020 03:47:10",
      "content": "<p>Nice to see that catboost gave you better result, i tried but results are not better than LGB</p>",
      "rawMarkdown": "Nice to see that catboost gave you better result, i tried but results are not better than LGB",
      "votes": null
    },
    {
      "id": "1100770",
      "postDate": "12/03/2020 10:22:54",
      "content": "<p>Could you elaborate on using covariate shift for selecting features? I'm not that familiar with that method.</p>",
      "rawMarkdown": "Could you elaborate on using covariate shift for selecting features? I'm not that familiar with that method.",
      "votes": null
    },
    {
      "id": "1101387",
      "postDate": "12/03/2020 21:40:24",
      "content": "<p><a href=\"https://www.kaggle.com/tunguz/elo-adversarial-validation\" target=\"_blank\">https://www.kaggle.com/tunguz/elo-adversarial-validation</a> this might help. </p>\n<p>In nutshell In Adversarial validation, we will predict whether the data point is from train or test set. We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score.</p>",
      "rawMarkdown": "https://www.kaggle.com/tunguz/elo-adversarial-validation this might help. \n\nIn nutshell In Adversarial validation, we will predict whether the data point is from train or test set. We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score.",
      "votes": null
    },
    {
      "id": "1101519",
      "postDate": "12/04/2020 01:15:03",
      "content": "<p>Here is something I think I have noticed (maybe I am wrong)</p>\n<p>There are some specific 'large' events at very low frequency at some higher distances from eruption (something like 25%, 50% , 75% of the maximum time from eruption).</p>\n<p>I <em>think</em> that these are tending to cause CV predictions to cluster around these points in time. However removing them does not seem to help train or test predictions.</p>\n<p>When I tried a linear regression (SGD), after a lot of messing around with input data I managed to get a semi passable score (but still very poor compared to LGBM), which actually gave CV &gt; LB (something like 8million LB vs 9million CV).</p>\n<p>So it seemed to me that maybe some of the larger events at mid points along the time axis are causing the tree models to overfit to those specific data points…which a linear regression wouldn't do I think, explaining why my linear regression CV doesn't overfit. However as I've not seen any LB improvement from removing them, it doesn't seem like getting rid of them completely helps overall, so I feel like that suggests that they have some reflection in the test data.</p>\n<p>Any thoughts welcome, I may well be wrong on this.</p>\n<p>Ps. the segment IDs I have filtered to with high readings / low frequencies and &gt;25% of max time to eruption are:</p>\n<p>[43303881, 179584121, 196942129, 302498114, 310568336, 567814647, 633652522, 957807611, 1111649370, 1423567749, 1424510231, 1443158011, 1623655934, 1634919976, 1642532502, 1778817952, 1848578834, 2071108961]</p>",
      "rawMarkdown": "Here is something I think I have noticed (maybe I am wrong)\n\nThere are some specific 'large' events at very low frequency at some higher distances from eruption (something like 25%, 50% , 75% of the maximum time from eruption).\n\nI *think* that these are tending to cause CV predictions to cluster around these points in time. However removing them does not seem to help train or test predictions.\n\nWhen I tried a linear regression (SGD), after a lot of messing around with input data I managed to get a semi passable score (but still very poor compared to LGBM), which actually gave CV > LB (something like 8million LB vs 9million CV).\n\nSo it seemed to me that maybe some of the larger events at mid points along the time axis are causing the tree models to overfit to those specific data points...which a linear regression wouldn't do I think, explaining why my linear regression CV doesn't overfit. However as I've not seen any LB improvement from removing them, it doesn't seem like getting rid of them completely helps overall, so I feel like that suggests that they have some reflection in the test data.\n\nAny thoughts welcome, I may well be wrong on this.\n\nPs. the segment IDs I have filtered to with high readings / low frequencies and >25% of max time to eruption are:\n\n[43303881, 179584121, 196942129, 302498114, 310568336, 567814647, 633652522, 957807611, 1111649370, 1423567749, 1424510231, 1443158011, 1623655934, 1634919976, 1642532502, 1778817952, 1848578834, 2071108961]",
      "votes": null
    },
    {
      "id": "1101745",
      "postDate": "12/04/2020 07:43:58",
      "content": "<p>Awesome, thanks for the explanation and the link!</p>",
      "rawMarkdown": "Awesome, thanks for the explanation and the link!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091771,
      "author_name": "leventelippenszky",
      "author_url": "",
      "post_date": "11/26/2020 09:14:00",
      "content": "<p>I have CV: 2306182 and LB: 4797783 at the 15th position currently, I use 5-fold CV and LGBM as well. I experience huge train/validation gaps most of the time that translates to a big gap between CV and LB errors, though the errors correlate well. I haven't figured out how to reduce overfitting as much as you did, I tried to remove features and currently use only 120.<br>\nDo you use a simple shuffled 5-fold CV as validation strategy? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1093494,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "11/27/2020 19:20:57",
          "content": "<p>yeah i am using simple 5 fold CV as validation strategy and using covariate shift to check for features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100770,
          "author_name": "leventelippenszky",
          "author_url": "",
          "post_date": "12/03/2020 10:22:54",
          "content": "<p>Could you elaborate on using covariate shift for selecting features? I'm not that familiar with that method.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101387,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "12/03/2020 21:40:24",
          "content": "<p><a href=\"https://www.kaggle.com/tunguz/elo-adversarial-validation\" target=\"_blank\">https://www.kaggle.com/tunguz/elo-adversarial-validation</a> this might help. </p>\n<p>In nutshell In Adversarial validation, we will predict whether the data point is from train or test set. We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101745,
          "author_name": "leventelippenszky",
          "author_url": "",
          "post_date": "12/04/2020 07:43:58",
          "content": "<p>Awesome, thanks for the explanation and the link!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1095867,
      "author_name": "windowbyryeol",
      "author_url": "",
      "post_date": "11/30/2020 02:49:55",
      "content": "<p>CV:3190586.8421 LB: 5688433 by using LGBM and CV:2425092.204 LB: 4836750 by using CatBoost (5-fold) <br>\nI think test and train are very different. but I don't have any idea…. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1095901,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "11/30/2020 03:47:10",
          "content": "<p>Nice to see that catboost gave you better result, i tried but results are not better than LGB</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1101519,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "12/04/2020 01:15:03",
      "content": "<p>Here is something I think I have noticed (maybe I am wrong)</p>\n<p>There are some specific 'large' events at very low frequency at some higher distances from eruption (something like 25%, 50% , 75% of the maximum time from eruption).</p>\n<p>I <em>think</em> that these are tending to cause CV predictions to cluster around these points in time. However removing them does not seem to help train or test predictions.</p>\n<p>When I tried a linear regression (SGD), after a lot of messing around with input data I managed to get a semi passable score (but still very poor compared to LGBM), which actually gave CV &gt; LB (something like 8million LB vs 9million CV).</p>\n<p>So it seemed to me that maybe some of the larger events at mid points along the time axis are causing the tree models to overfit to those specific data points…which a linear regression wouldn't do I think, explaining why my linear regression CV doesn't overfit. However as I've not seen any LB improvement from removing them, it doesn't seem like getting rid of them completely helps overall, so I feel like that suggests that they have some reflection in the test data.</p>\n<p>Any thoughts welcome, I may well be wrong on this.</p>\n<p>Ps. the segment IDs I have filtered to with high readings / low frequencies and &gt;25% of max time to eruption are:</p>\n<p>[43303881, 179584121, 196942129, 302498114, 310568336, 567814647, 633652522, 957807611, 1111649370, 1423567749, 1424510231, 1443158011, 1623655934, 1634919976, 1642532502, 1778817952, 1848578834, 2071108961]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091120": "Hi all,\n\nAnyone interested in sharing CV and LB score?\n\nFollowing is my CV vs LB comparison:\n\nI am using 5-fold LGB model.\n| CV | LB | features | Note\n| --- | --- | --- | --- |\n| 388529 | 4827594  | 223 |\n| 387870 | 4759488  | 183 |\n| 399560 | 4833849  | 143 |\n| 401977 | 4942119  | 123 |\n| 480388 | 5208812  | 100 | using https://www.kaggle.com/amanooo/ingv-volcanic-basic-solution-stft",
    "1091771": "I have CV: 2306182 and LB: 4797783 at the 15th position currently, I use 5-fold CV and LGBM as well. I experience huge train/validation gaps most of the time that translates to a big gap between CV and LB errors, though the errors correlate well. I haven't figured out how to reduce overfitting as much as you did, I tried to remove features and currently use only 120.\nDo you use a simple shuffled 5-fold CV as validation strategy?",
    "1093494": "yeah i am using simple 5 fold CV as validation strategy and using covariate shift to check for features.",
    "1095867": "CV:3190586.8421 LB: 5688433 by using LGBM and CV:2425092.204 LB: 4836750 by using CatBoost (5-fold) \nI think test and train are very different. but I don't have any idea....",
    "1095901": "Nice to see that catboost gave you better result, i tried but results are not better than LGB",
    "1100770": "Could you elaborate on using covariate shift for selecting features? I'm not that familiar with that method.",
    "1101387": "https://www.kaggle.com/tunguz/elo-adversarial-validation this might help. \n\nIn nutshell In Adversarial validation, we will predict whether the data point is from train or test set. We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score.",
    "1101519": "Here is something I think I have noticed (maybe I am wrong)\n\nThere are some specific 'large' events at very low frequency at some higher distances from eruption (something like 25%, 50% , 75% of the maximum time from eruption).\n\nI *think* that these are tending to cause CV predictions to cluster around these points in time. However removing them does not seem to help train or test predictions.\n\nWhen I tried a linear regression (SGD), after a lot of messing around with input data I managed to get a semi passable score (but still very poor compared to LGBM), which actually gave CV > LB (something like 8million LB vs 9million CV).\n\nSo it seemed to me that maybe some of the larger events at mid points along the time axis are causing the tree models to overfit to those specific data points...which a linear regression wouldn't do I think, explaining why my linear regression CV doesn't overfit. However as I've not seen any LB improvement from removing them, it doesn't seem like getting rid of them completely helps overall, so I feel like that suggests that they have some reflection in the test data.\n\nAny thoughts welcome, I may well be wrong on this.\n\nPs. the segment IDs I have filtered to with high readings / low frequencies and >25% of max time to eruption are:\n\n[43303881, 179584121, 196942129, 302498114, 310568336, 567814647, 633652522, 957807611, 1111649370, 1423567749, 1424510231, 1443158011, 1623655934, 1634919976, 1642532502, 1778817952, 1848578834, 2071108961]",
    "1101745": "Awesome, thanks for the explanation and the link!"
  },
  "source": "meta"
}