{
  "id": 327295,
  "title": "CV vs LB Scores",
  "url": "/competitions/amex-default-prediction/discussion/327295",
  "author_name": "",
  "post_date": "2022-05-26T16:07:39.281178600Z",
  "votes": 29,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Although it might seem too early to start, no Kaggle competition is complete without this discussion, so let's keep it rolling.</p>\n<p>My current CV: 0.79 and LB: 0.791</p>\n<p>What's yours ?</p>",
  "messages": [
    {
      "id": "1802286",
      "postDate": "05/26/2022 16:07:39",
      "content": "<p>Although it might seem too early to start, no Kaggle competition is complete without this discussion, so let's keep it rolling.</p>\n<p>My current CV: 0.79 and LB: 0.791</p>\n<p>What's yours ?</p>",
      "rawMarkdown": "Although it might seem too early to start, no Kaggle competition is complete without this discussion, so let's keep it rolling.\n\nMy current CV: 0.79 and LB: 0.791\n\nWhat's yours ?",
      "votes": null
    },
    {
      "id": "1802687",
      "postDate": "05/27/2022 03:30:21",
      "content": "<p>StratifiedKFold (To equalize the target ratio between train and valid)<br>\ncv : 0.792<br>\nlb : 0.793</p>\n<p>There seems to be an almost perfect correlation between cv and lb.</p>",
      "rawMarkdown": "StratifiedKFold (To equalize the target ratio between train and valid)\ncv : 0.792\nlb : 0.793\n\nThere seems to be an almost perfect correlation between cv and lb.",
      "votes": null
    },
    {
      "id": "1803104",
      "postDate": "05/27/2022 13:27:31",
      "content": "<p>exactly lb=cv+0.001</p>",
      "rawMarkdown": "exactly lb=cv+0.001",
      "votes": null
    },
    {
      "id": "1804754",
      "postDate": "05/29/2022 12:15:53",
      "content": "<p>Mine is lb = cv + 0.002</p>",
      "rawMarkdown": "Mine is lb = cv + 0.002",
      "votes": null
    },
    {
      "id": "1804882",
      "postDate": "05/29/2022 14:17:34",
      "content": "<p>CV : 0.7926<br>\nLB : 0.793</p>\n<p>Not much of a difference.</p>",
      "rawMarkdown": "CV : 0.7926\nLB : 0.793\n\nNot much of a difference.",
      "votes": null
    },
    {
      "id": "1805043",
      "postDate": "05/29/2022 17:26:14",
      "content": "<p>0.7875 -&gt; 0.789<br>\n0.7883 -&gt; 0.790</p>",
      "rawMarkdown": "0.7875 -> 0.789\n0.7883 -> 0.790",
      "votes": null
    },
    {
      "id": "1805246",
      "postDate": "05/30/2022 02:01:52",
      "content": "<p>Nice to see you in here.</p>",
      "rawMarkdown": "Nice to see you in here.",
      "votes": null
    },
    {
      "id": "1806366",
      "postDate": "05/31/2022 05:09:09",
      "content": "<p>cv: 0.79100<br>\nlb: 0.794</p>",
      "rawMarkdown": "cv: 0.79100\nlb: 0.794",
      "votes": null
    },
    {
      "id": "1814513",
      "postDate": "06/08/2022 01:59:20",
      "content": "<p>slight difference</p>",
      "rawMarkdown": "slight difference",
      "votes": null
    },
    {
      "id": "1815030",
      "postDate": "06/08/2022 15:31:29",
      "content": "<p>CV: 0.7979 LB: 0.798</p>",
      "rawMarkdown": "CV: 0.7979 LB: 0.798",
      "votes": null
    },
    {
      "id": "1818643",
      "postDate": "06/13/2022 02:58:34",
      "content": "<p>update CV：7997, LB 800</p>",
      "rawMarkdown": "update CV：7997, LB 800",
      "votes": null
    },
    {
      "id": "1819706",
      "postDate": "06/14/2022 05:06:46",
      "content": "<p>First dictionary is the mean scores between folds and second one is the standard deviations of them. Third dictionary is the oof scores. Test predictions score 0.020 on leaderboard for some reason. I have to debug and find out why. Any ideas?</p>\n<pre><code>2022-06-13 23:13:29 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier Mean Validation Scores: {\n  \"accuracy\": 0.9042694357097327,\n  \"roc_auc\": 0.9619140777122137,\n  \"precision\": 0.8161049576096374,\n  \"recall\": 0.8136382237510731,\n  \"specificity\": 0.9359366040842729,\n  \"f1\": 0.8148650687635344,\n  \"evaluation_metric\": 0.794684423335588\n} (±{\n  \"accuracy\": 0.0007686215641755108,\n  \"roc_auc\": 0.0005379466081383408,\n  \"precision\": 0.002399818905144646,\n  \"recall\": 0.0028754265534823,\n  \"specificity\": 0.0011257321434339884,\n  \"f1\": 0.0015034944970760782,\n  \"evaluation_metric\": 0.003465897684270095\n})\n2022-06-13 23:13:30 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier OOF Scores: {\n  \"accuracy\": 0.9042694366906145,\n  \"roc_auc\": 0.9618961932602416,\n  \"precision\": 0.8160969021693256,\n  \"recall\": 0.8136381997509005,\n  \"specificity\": 0.935936604084273,\n  \"f1\": 0.8148656962974825,\n  \"evaluation_metric\": 0.7942420003857089\n}\n</code></pre>",
      "rawMarkdown": "First dictionary is the mean scores between folds and second one is the standard deviations of them. Third dictionary is the oof scores. Test predictions score 0.020 on leaderboard for some reason. I have to debug and find out why. Any ideas?\n\n```\n2022-06-13 23:13:29 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier Mean Validation Scores: {\n  \"accuracy\": 0.9042694357097327,\n  \"roc_auc\": 0.9619140777122137,\n  \"precision\": 0.8161049576096374,\n  \"recall\": 0.8136382237510731,\n  \"specificity\": 0.9359366040842729,\n  \"f1\": 0.8148650687635344,\n  \"evaluation_metric\": 0.794684423335588\n} (±{\n  \"accuracy\": 0.0007686215641755108,\n  \"roc_auc\": 0.0005379466081383408,\n  \"precision\": 0.002399818905144646,\n  \"recall\": 0.0028754265534823,\n  \"specificity\": 0.0011257321434339884,\n  \"f1\": 0.0015034944970760782,\n  \"evaluation_metric\": 0.003465897684270095\n})\n2022-06-13 23:13:30 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier OOF Scores: {\n  \"accuracy\": 0.9042694366906145,\n  \"roc_auc\": 0.9618961932602416,\n  \"precision\": 0.8160969021693256,\n  \"recall\": 0.8136381997509005,\n  \"specificity\": 0.935936604084273,\n  \"f1\": 0.8148656962974825,\n  \"evaluation_metric\": 0.7942420003857089\n}\n```",
      "votes": null
    },
    {
      "id": "1820625",
      "postDate": "06/14/2022 19:40:56",
      "content": "<p>Maybe you lost the records' order.</p>",
      "rawMarkdown": "Maybe you lost the records' order.",
      "votes": null
    },
    {
      "id": "1820910",
      "postDate": "06/15/2022 05:12:00",
      "content": "<p>That makes sense. Groupby might be messing up with the order of customer_IDs. I'll try left join while assigning the predictions.</p>",
      "rawMarkdown": "That makes sense. Groupby might be messing up with the order of customer_IDs. I'll try left join while assigning the predictions.",
      "votes": null
    },
    {
      "id": "1852273",
      "postDate": "07/12/2022 01:03:31",
      "content": "<p>CV : 7991<br>\nLB : 800</p>",
      "rawMarkdown": "CV : 7991\nLB : 800",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1802687,
      "author_name": "ryotak12",
      "author_url": "",
      "post_date": "05/27/2022 03:30:21",
      "content": "<p>StratifiedKFold (To equalize the target ratio between train and valid)<br>\ncv : 0.792<br>\nlb : 0.793</p>\n<p>There seems to be an almost perfect correlation between cv and lb.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1803104,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "05/27/2022 13:27:31",
          "content": "<p>exactly lb=cv+0.001</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1804754,
      "author_name": "brandonhu0215",
      "author_url": "",
      "post_date": "05/29/2022 12:15:53",
      "content": "<p>Mine is lb = cv + 0.002</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1804882,
      "author_name": "devkhant24",
      "author_url": "",
      "post_date": "05/29/2022 14:17:34",
      "content": "<p>CV : 0.7926<br>\nLB : 0.793</p>\n<p>Not much of a difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1805043,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "05/29/2022 17:26:14",
      "content": "<p>0.7875 -&gt; 0.789<br>\n0.7883 -&gt; 0.790</p>",
      "votes": null,
      "replies": [
        {
          "id": 1805246,
          "author_name": "duykhanh99",
          "author_url": "",
          "post_date": "05/30/2022 02:01:52",
          "content": "<p>Nice to see you in here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1806366,
      "author_name": "kutsenkodmitriy",
      "author_url": "",
      "post_date": "05/31/2022 05:09:09",
      "content": "<p>cv: 0.79100<br>\nlb: 0.794</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1814513,
      "author_name": "shanggangli",
      "author_url": "",
      "post_date": "06/08/2022 01:59:20",
      "content": "<p>slight difference</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1815030,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "06/08/2022 15:31:29",
      "content": "<p>CV: 0.7979 LB: 0.798</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1818643,
      "author_name": "mahluo",
      "author_url": "",
      "post_date": "06/13/2022 02:58:34",
      "content": "<p>update CV：7997, LB 800</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1819706,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "06/14/2022 05:06:46",
      "content": "<p>First dictionary is the mean scores between folds and second one is the standard deviations of them. Third dictionary is the oof scores. Test predictions score 0.020 on leaderboard for some reason. I have to debug and find out why. Any ideas?</p>\n<pre><code>2022-06-13 23:13:29 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier Mean Validation Scores: {\n  \"accuracy\": 0.9042694357097327,\n  \"roc_auc\": 0.9619140777122137,\n  \"precision\": 0.8161049576096374,\n  \"recall\": 0.8136382237510731,\n  \"specificity\": 0.9359366040842729,\n  \"f1\": 0.8148650687635344,\n  \"evaluation_metric\": 0.794684423335588\n} (±{\n  \"accuracy\": 0.0007686215641755108,\n  \"roc_auc\": 0.0005379466081383408,\n  \"precision\": 0.002399818905144646,\n  \"recall\": 0.0028754265534823,\n  \"specificity\": 0.0011257321434339884,\n  \"f1\": 0.0015034944970760782,\n  \"evaluation_metric\": 0.003465897684270095\n})\n2022-06-13 23:13:30 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier OOF Scores: {\n  \"accuracy\": 0.9042694366906145,\n  \"roc_auc\": 0.9618961932602416,\n  \"precision\": 0.8160969021693256,\n  \"recall\": 0.8136381997509005,\n  \"specificity\": 0.935936604084273,\n  \"f1\": 0.8148656962974825,\n  \"evaluation_metric\": 0.7942420003857089\n}\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1820625,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "06/14/2022 19:40:56",
          "content": "<p>Maybe you lost the records' order.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1820910,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "06/15/2022 05:12:00",
          "content": "<p>That makes sense. Groupby might be messing up with the order of customer_IDs. I'll try left join while assigning the predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1852273,
      "author_name": "zakopur0",
      "author_url": "",
      "post_date": "07/12/2022 01:03:31",
      "content": "<p>CV : 7991<br>\nLB : 800</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802286": "Although it might seem too early to start, no Kaggle competition is complete without this discussion, so let's keep it rolling.\n\nMy current CV: 0.79 and LB: 0.791\n\nWhat's yours ?",
    "1802687": "StratifiedKFold (To equalize the target ratio between train and valid)\ncv : 0.792\nlb : 0.793\n\nThere seems to be an almost perfect correlation between cv and lb.",
    "1803104": "exactly lb=cv+0.001",
    "1804754": "Mine is lb = cv + 0.002",
    "1804882": "CV : 0.7926\nLB : 0.793\n\nNot much of a difference.",
    "1805043": "0.7875 -> 0.789\n0.7883 -> 0.790",
    "1805246": "Nice to see you in here.",
    "1806366": "cv: 0.79100\nlb: 0.794",
    "1814513": "slight difference",
    "1815030": "CV: 0.7979 LB: 0.798",
    "1818643": "update CV：7997, LB 800",
    "1819706": "First dictionary is the mean scores between folds and second one is the standard deviations of them. Third dictionary is the oof scores. Test predictions score 0.020 on leaderboard for some reason. I have to debug and find out why. Any ideas?\n\n```\n2022-06-13 23:13:29 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier Mean Validation Scores: {\n  \"accuracy\": 0.9042694357097327,\n  \"roc_auc\": 0.9619140777122137,\n  \"precision\": 0.8161049576096374,\n  \"recall\": 0.8136382237510731,\n  \"specificity\": 0.9359366040842729,\n  \"f1\": 0.8148650687635344,\n  \"evaluation_metric\": 0.794684423335588\n} (±{\n  \"accuracy\": 0.0007686215641755108,\n  \"roc_auc\": 0.0005379466081383408,\n  \"precision\": 0.002399818905144646,\n  \"recall\": 0.0028754265534823,\n  \"specificity\": 0.0011257321434339884,\n  \"f1\": 0.0015034944970760782,\n  \"evaluation_metric\": 0.003465897684270095\n})\n2022-06-13 23:13:30 INFO lightgbm_model_trainer - train_and_validate:\nlightgbm_classifier OOF Scores: {\n  \"accuracy\": 0.9042694366906145,\n  \"roc_auc\": 0.9618961932602416,\n  \"precision\": 0.8160969021693256,\n  \"recall\": 0.8136381997509005,\n  \"specificity\": 0.935936604084273,\n  \"f1\": 0.8148656962974825,\n  \"evaluation_metric\": 0.7942420003857089\n}\n```",
    "1820625": "Maybe you lost the records' order.",
    "1820910": "That makes sense. Groupby might be messing up with the order of customer_IDs. I'll try left join while assigning the predictions.",
    "1852273": "CV : 7991\nLB : 800"
  },
  "source": "meta"
}