{
  "id": 476832,
  "title": "Train vs Test (Public LB) Data Distribution",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/476832",
  "author_name": "Vikram Sandu",
  "post_date": "2024-02-13T16:59:21.396000",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I did NOT find a good correlation between CV and LB. To figure out why I did this small experiment, I compared each class's CV separately across 3 different models.</p>\n<p>To understand this, Let's consider Model-1 as a reference model since it is the best LB and the worst CV model among the 3 models. If we compare Each Class's CV Score of this model with Model-2 and 3. we find out the CV Score (KLD-Loss) of the Model-2 and 3 are better for every class except LPD and GRDA. The table below illustrates these findings, with the 'D/I' column indicating whether the score decreased or increased w.r.t Model-1.</p>\n<p>Can we infer from this experiment that the Public LB contains a higher proportion of LPD and GRDA examples?</p>\n\n<table>\n    <tbody><tr>\n        <th></th>\n        <th></th>\n        <th></th>\n        <th></th>\n        <th></th>\n        <th></th>\n    </tr>\n    <tr>\n        <td>Seizure</td>\n        <td> 0.8386 </td>\n        <td>0.7313</td>\n        <td>D</td>\n        <td>0.6749</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LPD</td>\n        <td> 0.6017 </td>\n        <td>0.6805</td>\n        <td>I</td>\n        <td>0.7092</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>GPD</td>\n        <td> 0.6080 </td>\n        <td>0.5625</td>\n        <td>D</td>\n        <td>0.5502</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LRDA</td>\n        <td> 1.2434 </td>\n        <td>1.1401</td>\n        <td>D</td>\n        <td>1.1756</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>GRDA</td>\n        <td> 0.8362 </td>\n        <td>0.8805</td>\n        <td>I</td>\n        <td>0.8796</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>Other</td>\n        <td> 0.4649 </td>\n        <td>0.4569</td>\n        <td>D</td>\n        <td>0.4585</td>\n        <td>D</td>\n    </tr>\n        <tr>\n        <td>OOF-CV</td>\n        <td> 0.6551 </td>\n        <td>0.6377</td>\n        <td>D</td>\n        <td>0.6326</td>\n        <td>D</td>\n    </tr>\n        <tr>\n        <td>LB</td>\n        <td> 0.39 </td>\n        <td>0.41</td>\n        <td>I</td>\n        <td>0.42</td>\n        <td>I</td>\n    </tr>\n\n\n</tbody></table>",
  "messages": [
    {
      "id": 2650715,
      "postDate": "2024-02-13T16:59:21.397Z",
      "content": "<p>I did NOT find a good correlation between CV and LB. To figure out why I did this small experiment, I compared each class's CV separately across 3 different models.</p>\n<p>To understand this, Let's consider Model-1 as a reference model since it is the best LB and the worst CV model among the 3 models. If we compare Each Class's CV Score of this model with Model-2 and 3. we find out the CV Score (KLD-Loss) of the Model-2 and 3 are better for every class except LPD and GRDA. The table below illustrates these findings, with the 'D/I' column indicating whether the score decreased or increased w.r.t Model-1.</p>\n<p>Can we infer from this experiment that the Public LB contains a higher proportion of LPD and GRDA examples?</p>\n\n<table>\n    <tbody><tr>\n        <th></th>\n        <th></th>\n        <th></th>\n        <th></th>\n        <th></th>\n        <th></th>\n    </tr>\n    <tr>\n        <td>Seizure</td>\n        <td> 0.8386 </td>\n        <td>0.7313</td>\n        <td>D</td>\n        <td>0.6749</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LPD</td>\n        <td> 0.6017 </td>\n        <td>0.6805</td>\n        <td>I</td>\n        <td>0.7092</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>GPD</td>\n        <td> 0.6080 </td>\n        <td>0.5625</td>\n        <td>D</td>\n        <td>0.5502</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LRDA</td>\n        <td> 1.2434 </td>\n        <td>1.1401</td>\n        <td>D</td>\n        <td>1.1756</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>GRDA</td>\n        <td> 0.8362 </td>\n        <td>0.8805</td>\n        <td>I</td>\n        <td>0.8796</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>Other</td>\n        <td> 0.4649 </td>\n        <td>0.4569</td>\n        <td>D</td>\n        <td>0.4585</td>\n        <td>D</td>\n    </tr>\n        <tr>\n        <td>OOF-CV</td>\n        <td> 0.6551 </td>\n        <td>0.6377</td>\n        <td>D</td>\n        <td>0.6326</td>\n        <td>D</td>\n    </tr>\n        <tr>\n        <td>LB</td>\n        <td> 0.39 </td>\n        <td>0.41</td>\n        <td>I</td>\n        <td>0.42</td>\n        <td>I</td>\n    </tr>\n\n\n</tbody></table>",
      "rawMarkdown": "I did NOT find a good correlation between CV and LB. To figure out why I did this small experiment, I compared each class's CV separately across 3 different models.\n\nTo understand this, Let's consider Model-1 as a reference model since it is the best LB and the worst CV model among the 3 models. If we compare Each Class's CV Score of this model with Model-2 and 3. we find out the CV Score (KLD-Loss) of the Model-2 and 3 are better for every class except LPD and GRDA. The table below illustrates these findings, with the 'D/I' column indicating whether the score decreased or increased w.r.t Model-1.\n\nCan we infer from this experiment that the Public LB contains a higher proportion of LPD and GRDA examples?\n\n\n<style>\ntable {\n    width: 100%;\n    border-collapse: collapse;\n}\n\nth, td {\n    padding: 8px;\n    text-align: left;\n}\n\nth {\n    background-color: #FF0000; /* Red color */\n    color: white;\n}\n</style>\n\n<table>\n    <tr>\n        <th><span style=\"color:red\">Class</span></th>\n        <th><span style=\"color:red\">Model-1</span></th>\n        <th><span style=\"color:red\">Model-2</span></th>\n        <th><span style=\"color:red\">D/I</span></th>\n        <th><span style=\"color:red\">Model-3</span></th>\n        <th><span style=\"color:red\">D/I</span></th>\n    </tr>\n    <tr>\n        <td>Seizure</td>\n        <td> 0.8386 </td>\n        <td>0.7313</td>\n        <td>D</td>\n        <td>0.6749</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LPD</td>\n        <td> 0.6017 </td>\n        <td>0.6805</td>\n        <td>I</td>\n        <td>0.7092</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>GPD</td>\n        <td> 0.6080 </td>\n        <td>0.5625</td>\n        <td>D</td>\n        <td>0.5502</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LRDA</td>\n        <td> 1.2434 </td>\n        <td>1.1401</td>\n        <td>D</td>\n        <td>1.1756</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>GRDA</td>\n        <td> 0.8362 </td>\n        <td>0.8805</td>\n        <td>I</td>\n        <td>0.8796</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>Other</td>\n        <td> 0.4649 </td>\n        <td>0.4569</td>\n        <td>D</td>\n        <td>0.4585</td>\n        <td>D</td>\n    </tr>\n        <tr style=\"background-color: lightgreen;\">\n        <td>OOF-CV</td>\n        <td> 0.6551 </td>\n        <td>0.6377</td>\n        <td>D</td>\n        <td>0.6326</td>\n        <td>D</td>\n    </tr>\n        <tr style=\"background-color: lightgreen;\">\n        <td>LB</td>\n        <td> 0.39 </td>\n        <td>0.41</td>\n        <td>I</td>\n        <td>0.42</td>\n        <td>I</td>\n    </tr>\n    \n    \n</table>",
      "votes": 7
    },
    {
      "id": 2651045,
      "postDate": "2024-02-13T21:09:35.650Z",
      "content": "<p>Great Analysis  .. Funny thing is how you guys are getting .39 LB with .655 CV  :D .. I am not able to break .40 with .550 CV :P</p>",
      "rawMarkdown": "Great Analysis  .. Funny thing is how you guys are getting .39 LB with .655 CV  :D .. I am not able to break .40 with .550 CV :P",
      "votes": 4,
      "replies": [
        {
          "id": 2651186,
          "postDate": "2024-02-14T02:49:58.027Z",
          "content": "<p>I am in the same boat 🛶. My best single model is CV=0.59 and LB=0.40 and I haven’t been able to breach the barrier yet. </p>",
          "rawMarkdown": "I am in the same boat 🛶. My best single model is CV=0.59 and LB=0.40 and I haven’t been able to breach the barrier yet. ",
          "votes": 1
        },
        {
          "id": 2653953,
          "postDate": "2024-02-15T18:34:10.290Z",
          "content": "<p>Could be a signal of a better designed CV? I mean. If the folds are not different enough wouldn't be like train same model n times?</p>",
          "rawMarkdown": "Could be a signal of a better designed CV? I mean. If the folds are not different enough wouldn't be like train same model n times?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2656076,
      "postDate": "2024-02-17T12:45:49.913Z",
      "content": "<p><a href=\"https://www.kaggle.com/vikramsandu\" target=\"_blank\">@vikramsandu</a>  is possible to share the code to calculate the KL score per class? <br>\nI'm getting some strange results </p>\n<p>ps: also I noticed that the average of all class scores is not equal to OOF-CV score </p>",
      "rawMarkdown": "@vikramsandu  is possible to share the code to calculate the KL score per class? \nI'm getting some strange results \n\nps: also I noticed that the average of all class scores is not equal to OOF-CV score ",
      "votes": 1,
      "replies": [
        {
          "id": 2656156,
          "postDate": "2024-02-17T13:56:36.347Z",
          "content": "<pre><code> sys\nsys.path.append()\n kaggle_kl_div  score\n\ntrue = train_df[[] + TARGETS].copy()\ntrue_copy = true.copy()\noof = pd.DataFrame(oof_pred_arr, columns=TARGETS)\noof.insert(, , train_df[])\noof_copy = oof.copy()\n\ncv_score = score(solution=true, submission=oof, row_id_column_name=)\n(,cv_score)\n\n\n\n\n t  train_df[].unique():\n    t_labels = train_df[train_df[]==t][].values\n    true_filtered_df = true_copy[true_copy[].isin(t_labels)].reset_index(drop=)\n    oof_filtered_df = oof_copy[oof_copy[].isin(t_labels)].reset_index(drop=)\n    cv_score = score(solution=true_filtered_df, submission=oof_filtered_df, row_id_column_name=)\n    ()\n</code></pre>\n<p>Your getting different score because every fold doesn't consists of same number of examples in my case.</p>",
          "rawMarkdown": "```python\n\nimport sys\nsys.path.append('/kaggle/input/kaggle-kl-div')\nfrom kaggle_kl_div import score\n\ntrue = train_df[[\"label_id\"] + TARGETS].copy()\ntrue_copy = true.copy()\noof = pd.DataFrame(oof_pred_arr, columns=TARGETS)\noof.insert(0, \"label_id\", train_df[\"label_id\"])\noof_copy = oof.copy()\n\ncv_score = score(solution=true, submission=oof, row_id_column_name='label_id')\nprint('CV Score KL-Div for EfficientNet-B0',cv_score)\n\n#--------------------------------------------------------------------------\n\n\nfor t in train_df['target'].unique():\n    t_labels = train_df[train_df['target']==t]['label_id'].values\n    true_filtered_df = true_copy[true_copy['label_id'].isin(t_labels)].reset_index(drop=True)\n    oof_filtered_df = oof_copy[oof_copy['label_id'].isin(t_labels)].reset_index(drop=True)\n    cv_score = score(solution=true_filtered_df, submission=oof_filtered_df, row_id_column_name='label_id')\n    print(f'Target: {t} CV-Score: {cv_score}')\n```\n\nYour getting different score because every fold doesn't consists of same number of examples in my case.",
          "votes": 2,
          "replies": [
            {
              "id": 2656157,
              "postDate": "2024-02-17T13:57:10.620Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2654636,
      "postDate": "2024-02-16T10:23:29.047Z",
      "content": "<p>Nice work!!<br>\nAnyway, do you have any insights about private dataset?</p>",
      "rawMarkdown": "Nice work!!\nAnyway, do you have any insights about private dataset?",
      "votes": 1,
      "replies": [
        {
          "id": 2654668,
          "postDate": "2024-02-16T11:05:40.090Z",
          "content": "<p>No idea, man!! Just trying to figure out the correlation between CV and Public LB. and hoping it would work on Private set too.</p>",
          "rawMarkdown": "No idea, man!! Just trying to figure out the correlation between CV and Public LB. and hoping it would work on Private set too.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2651045,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2024-02-13T21:09:35.650000",
      "content": "<p>Great Analysis  .. Funny thing is how you guys are getting .39 LB with .655 CV  :D .. I am not able to break .40 with .550 CV :P</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2651186,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2024-02-14T02:49:58.027000",
          "content": "<p>I am in the same boat 🛶. My best single model is CV=0.59 and LB=0.40 and I haven’t been able to breach the barrier yet. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2653953,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2024-02-15T18:34:10.290000",
          "content": "<p>Could be a signal of a better designed CV? I mean. If the folds are not different enough wouldn't be like train same model n times?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2656076,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2024-02-17T12:45:49.913000",
      "content": "<p><a href=\"https://www.kaggle.com/vikramsandu\" target=\"_blank\">@vikramsandu</a>  is possible to share the code to calculate the KL score per class? <br>\nI'm getting some strange results </p>\n<p>ps: also I noticed that the average of all class scores is not equal to OOF-CV score </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2656156,
          "author_name": "Vikram Sandu",
          "author_url": "",
          "post_date": "2024-02-17T13:56:36.347000",
          "content": "<pre><code> sys\nsys.path.append()\n kaggle_kl_div  score\n\ntrue = train_df[[] + TARGETS].copy()\ntrue_copy = true.copy()\noof = pd.DataFrame(oof_pred_arr, columns=TARGETS)\noof.insert(, , train_df[])\noof_copy = oof.copy()\n\ncv_score = score(solution=true, submission=oof, row_id_column_name=)\n(,cv_score)\n\n\n\n\n t  train_df[].unique():\n    t_labels = train_df[train_df[]==t][].values\n    true_filtered_df = true_copy[true_copy[].isin(t_labels)].reset_index(drop=)\n    oof_filtered_df = oof_copy[oof_copy[].isin(t_labels)].reset_index(drop=)\n    cv_score = score(solution=true_filtered_df, submission=oof_filtered_df, row_id_column_name=)\n    ()\n</code></pre>\n<p>Your getting different score because every fold doesn't consists of same number of examples in my case.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2656157,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-17T13:57:10.620000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2654636,
      "author_name": "HideBu",
      "author_url": "",
      "post_date": "2024-02-16T10:23:29.047000",
      "content": "<p>Nice work!!<br>\nAnyway, do you have any insights about private dataset?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2654668,
          "author_name": "Vikram Sandu",
          "author_url": "",
          "post_date": "2024-02-16T11:05:40.090000",
          "content": "<p>No idea, man!! Just trying to figure out the correlation between CV and Public LB. and hoping it would work on Private set too.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2650715": "I did NOT find a good correlation between CV and LB. To figure out why I did this small experiment, I compared each class's CV separately across 3 different models.\n\nTo understand this, Let's consider Model-1 as a reference model since it is the best LB and the worst CV model among the 3 models. If we compare Each Class's CV Score of this model with Model-2 and 3. we find out the CV Score (KLD-Loss) of the Model-2 and 3 are better for every class except LPD and GRDA. The table below illustrates these findings, with the 'D/I' column indicating whether the score decreased or increased w.r.t Model-1.\n\nCan we infer from this experiment that the Public LB contains a higher proportion of LPD and GRDA examples?\n\n\n<style>\ntable {\n    width: 100%;\n    border-collapse: collapse;\n}\n\nth, td {\n    padding: 8px;\n    text-align: left;\n}\n\nth {\n    background-color: #FF0000; /* Red color */\n    color: white;\n}\n</style>\n\n<table>\n    <tr>\n        <th><span style=\"color:red\">Class</span></th>\n        <th><span style=\"color:red\">Model-1</span></th>\n        <th><span style=\"color:red\">Model-2</span></th>\n        <th><span style=\"color:red\">D/I</span></th>\n        <th><span style=\"color:red\">Model-3</span></th>\n        <th><span style=\"color:red\">D/I</span></th>\n    </tr>\n    <tr>\n        <td>Seizure</td>\n        <td> 0.8386 </td>\n        <td>0.7313</td>\n        <td>D</td>\n        <td>0.6749</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LPD</td>\n        <td> 0.6017 </td>\n        <td>0.6805</td>\n        <td>I</td>\n        <td>0.7092</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>GPD</td>\n        <td> 0.6080 </td>\n        <td>0.5625</td>\n        <td>D</td>\n        <td>0.5502</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>LRDA</td>\n        <td> 1.2434 </td>\n        <td>1.1401</td>\n        <td>D</td>\n        <td>1.1756</td>\n        <td>D</td>\n    </tr>\n    <tr>\n        <td>GRDA</td>\n        <td> 0.8362 </td>\n        <td>0.8805</td>\n        <td>I</td>\n        <td>0.8796</td>\n        <td>I</td>\n    </tr>\n    <tr>\n        <td>Other</td>\n        <td> 0.4649 </td>\n        <td>0.4569</td>\n        <td>D</td>\n        <td>0.4585</td>\n        <td>D</td>\n    </tr>\n        <tr style=\"background-color: lightgreen;\">\n        <td>OOF-CV</td>\n        <td> 0.6551 </td>\n        <td>0.6377</td>\n        <td>D</td>\n        <td>0.6326</td>\n        <td>D</td>\n    </tr>\n        <tr style=\"background-color: lightgreen;\">\n        <td>LB</td>\n        <td> 0.39 </td>\n        <td>0.41</td>\n        <td>I</td>\n        <td>0.42</td>\n        <td>I</td>\n    </tr>\n    \n    \n</table>",
    "2651045": "Great Analysis  .. Funny thing is how you guys are getting .39 LB with .655 CV  :D .. I am not able to break .40 with .550 CV :P",
    "2656076": "@vikramsandu  is possible to share the code to calculate the KL score per class? \nI'm getting some strange results \n\nps: also I noticed that the average of all class scores is not equal to OOF-CV score ",
    "2654636": "Nice work!!\nAnyway, do you have any insights about private dataset?"
  }
}