{
  "id": 485697,
  "title": "CV vs LB gap shortens?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/485697",
  "author_name": "Cody_Null",
  "post_date": "2024-03-21T20:11:13.391000",
  "votes": 10,
  "comment_count": 23,
  "views": 0,
  "content": "<p>As most have found by now, there are two groups in our data set. Those who had many votes and those who did not. Calculating CV for the subsets separately we can see that the CV LB gap is actually quite small if we consider that the public LB is mostly high vote data. Of course, there are many assumptions people are making about what the private dataset may hold for us. I thought I would share some of my numbers in hope other may share as well.</p>\n<p>High votes group CV: 0.33 <br>\nLow votes group CV: 0.6<br>\nLB: 0.35</p>\n<p>This should be a fun competition at the end!</p>",
  "messages": [
    {
      "id": 2709761,
      "postDate": "2024-03-21T20:11:13.390Z",
      "content": "<p>As most have found by now, there are two groups in our data set. Those who had many votes and those who did not. Calculating CV for the subsets separately we can see that the CV LB gap is actually quite small if we consider that the public LB is mostly high vote data. Of course, there are many assumptions people are making about what the private dataset may hold for us. I thought I would share some of my numbers in hope other may share as well.</p>\n<p>High votes group CV: 0.33 <br>\nLow votes group CV: 0.6<br>\nLB: 0.35</p>\n<p>This should be a fun competition at the end!</p>",
      "rawMarkdown": "As most have found by now, there are two groups in our data set. Those who had many votes and those who did not. Calculating CV for the subsets separately we can see that the CV LB gap is actually quite small if we consider that the public LB is mostly high vote data. Of course, there are many assumptions people are making about what the private dataset may hold for us. I thought I would share some of my numbers in hope other may share as well.\n\nHigh votes group CV: 0.33 \nLow votes group CV: 0.6\nLB: 0.35\n\nThis should be a fun competition at the end!",
      "votes": 10
    },
    {
      "id": 2710168,
      "postDate": "2024-03-22T04:31:59.207Z",
      "content": "<p>Our vote count &gt;= 10 OOF score is 0.2174. Top 5 teams probably reached beyond 0.2.</p>",
      "rawMarkdown": "Our vote count >= 10 OOF score is 0.2174. Top 5 teams probably reached beyond 0.2.",
      "votes": 8,
      "replies": [
        {
          "id": 2710462,
          "postDate": "2024-03-22T09:21:18.293Z",
          "content": "<p>is this cv on Chris's created training data filtered out with votes &gt;= 10</p>",
          "rawMarkdown": "is this cv on Chris's created training data filtered out with votes >= 10",
          "replies": [
            {
              "id": 2710463,
              "postDate": "2024-03-22T09:22:15.347Z",
              "content": "<p>Nope, we are using a custom cv scheme.</p>",
              "rawMarkdown": "Nope, we are using a custom cv scheme.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2710549,
      "postDate": "2024-03-22T10:54:24.913Z",
      "content": "<p>Here is the prediction distribution by <code>total_evaluators</code> in oof_df.<br>\nY: kl divergence from label<br>\nX: total_evaluators</p>\n<p>This prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.<br>\nThe baseline model is almost same as <a href=\"https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train\" target=\"_blank\">moth's published notebook</a>.</p>\n<p>As you can see, kl divergence is lager in data with total_evaluators &lt;= 3.<br>\nIt seems that data with total_evaluators &lt;= 3 have difficult sample or noizy sample.<br>\nIn my experiments, this distribution often happens when training baseline efficient net model with all data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F9f2d180d88de96d03a41fa119a0703c2%2Fviz_.png?generation=1711104172826486&amp;alt=media\" alt=\"viz_of_kl_div_by_total_evaluators\"></p>\n<p>I will publish code soon.</p>",
      "rawMarkdown": "Here is the prediction distribution by `total_evaluators` in oof_df.\nY: kl divergence from label\nX: total_evaluators\n\nThis prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.\nThe baseline model is almost same as [moth's published notebook](https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train).\n\nAs you can see, kl divergence is lager in data with total_evaluators <= 3.\nIt seems that data with total_evaluators <= 3 have difficult sample or noizy sample.\nIn my experiments, this distribution often happens when training baseline efficient net model with all data. \n\n![viz_of_kl_div_by_total_evaluators](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F9f2d180d88de96d03a41fa119a0703c2%2Fviz_.png?generation=1711104172826486&alt=media)\n\nI will publish code soon.",
      "votes": 5,
      "replies": [
        {
          "id": 2710824,
          "postDate": "2024-03-22T15:23:37.120Z",
          "content": "<p>Nice looking plot, but I think the<code>total_evaluators</code> axis might be incorrect. I don't see any columns with 8 or 9 votes.</p>\n<pre><code> pandas  pd\n\ndf= pd.read_csv()\ntotal_votes= df[[, , , , , ]].(axis=)\n(((total_votes.unique())))\n</code></pre>",
          "rawMarkdown": "Nice looking plot, but I think the`total_evaluators` axis might be incorrect. I don't see any columns with 8 or 9 votes.\n\n```\nimport pandas as pd\n\ndf= pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\ntotal_votes= df[['gpd_vote', 'grda_vote', 'lpd_vote', 'lrda_vote', 'other_vote', 'seizure_vote']].sum(axis=1)\nprint(str(sorted(total_votes.unique())))\n```",
          "replies": [
            {
              "id": 2710834,
              "postDate": "2024-03-22T15:37:17.317Z",
              "content": "<p>This plot is based on <em><code>eeg-unique-df</code></em>, not train.csv<br>\nThe eeg-unique-df is created  in the same way as  in <a href=\"https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train?scriptVersionId=161067167&amp;cellId=11\" target=\"_blank\">chris or moth published notebook</a><br>\nThe eeg-unique-df is made by aggregation, so \"total_evaluators \" column contains float value.<br>\nI use  <code>df.astype(int)</code> to plot this, so this is rough plot of result.</p>",
              "rawMarkdown": "This plot is based on *`eeg-unique-df`*, not train.csv\nThe eeg-unique-df is created  in the same way as  in [chris or moth published notebook](https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train?scriptVersionId=161067167&cellId=11)\nThe eeg-unique-df is made by aggregation, so \"total_evaluators \" column contains float value.\nI use  `df.astype(int)` to plot this, so this is rough plot of result.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2710871,
          "postDate": "2024-03-22T16:17:41.940Z",
          "content": "<p>2stage learning looks like this.<br>\nThe model is trained in this way, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">EffNetB0 model trained twice, once for each of two training populations - [LB 0.39]</a></p>\n<ul>\n<li>Blue: stage1, CV(total_evaluators &gt; 9)=0.40</li>\n<li>Orange: srage2, CV(total_evaluators &gt; 9)=0.33</li>\n</ul>\n<p>As you can see, in data with total_evaluators &gt; 9, score get better. <br>\nThis 2stage model get around 0.35 in public LB.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F950351391c5b1c9226c18a501a690dce%2Foutput_2.png?generation=1711123884044601&amp;alt=media\" alt=\"viz_compare_2stage\"></p>",
          "rawMarkdown": "2stage learning looks like this.\nThe model is trained in this way, [EffNetB0 model trained twice, once for each of two training populations - [LB 0.39]](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135)\n\n- Blue: stage1, CV(total_evaluators > 9)=0.40\n- Orange: srage2, CV(total_evaluators > 9)=0.33\n\nAs you can see, in data with total_evaluators > 9, score get better. \nThis 2stage model get around 0.35 in public LB.\n![viz_compare_2stage](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F950351391c5b1c9226c18a501a690dce%2Foutput_2.png?generation=1711123884044601&alt=media)\n"
        }
      ]
    },
    {
      "id": 2710373,
      "postDate": "2024-03-22T07:52:23.347Z",
      "content": "<p>4 folds votes &gt;= 10 cv 2366 with online maybe 2499 LB.<br>\nOne thing I've noticed is if you change your folds seed as well as other seeds, you might see different cv, sometimes as large as from 256 to 249.<br>\nNow I got 4 folds model LB 0.24 on LB with cv 0.244.</p>",
      "rawMarkdown": "4 folds votes >= 10 cv 2366 with online maybe 2499 LB.\nOne thing I've noticed is if you change your folds seed as well as other seeds, you might see different cv, sometimes as large as from 256 to 249.\nNow I got 4 folds model LB 0.24 on LB with cv 0.244.",
      "votes": 4,
      "replies": [
        {
          "id": 2710563,
          "postDate": "2024-03-22T11:05:13.500Z",
          "content": "<p>Super score! What do you mean \"with online\"?</p>",
          "rawMarkdown": "Super score! What do you mean \"with online\"?",
          "replies": [
            {
              "id": 2710622,
              "postDate": "2024-03-22T12:07:13.153Z",
              "content": "<p>Online just mean submission LB result.</p>",
              "rawMarkdown": "Online just mean submission LB result."
            }
          ]
        },
        {
          "id": 2710922,
          "postDate": "2024-03-22T16:47:05.083Z",
          "content": "<p>Thanks for sharing. <br>\nHow about cv of vote &lt; 10 ? Not calculated ?\nDo you calculate cv in custom way or just filtered oof with vote &gt;= 10 ?</p>",
          "rawMarkdown": "Thanks for sharing. \nHow about cv of vote < 10 ? Not calculated ?\nDo you calculate cv in custom way or just filtered oof with vote >= 10 ?",
          "replies": [
            {
              "id": 2713023,
              "postDate": "2024-03-23T23:04:31.333Z",
              "content": "<p>I removed vote &lt; 10 for eval part to speedup. I used custom cv.</p>",
              "rawMarkdown": "I removed vote < 10 for eval part to speedup. I used custom cv.",
              "votes": 2
            },
            {
              "id": 2713472,
              "postDate": "2024-03-24T08:18:45.323Z",
              "content": "<blockquote>\n  <p>I used custom cv</p>\n</blockquote>\n<p>thanks. If just filtered vote &gt;= 10, I got 0.27 in CV with vote &gt;= 10, but it's not good in public LB.</p>",
              "rawMarkdown": "> I used custom cv\n\nthanks. If just filtered vote >= 10, I got 0.27 in CV with vote >= 10, but it's not good in public LB."
            },
            {
              "id": 2713543,
              "postDate": "2024-03-24T09:02:51.520Z",
              "content": "<p>Have been looking at the breakdown of number of votes e.g., &lt; 3 or &gt;=10 etc. and at the high level expert consensus given by unique eeg id (17089) in train.  </p>\n<p>Other has the majority of entries 7196 and roughly 40% of these have &gt;=10 votes 37.5% &lt; 3\nso fits in well with the &lt;3 or &gt;=10 idea.  Possibly due to the majority class it has too much influence with &gt;=10 filters (or &lt; 3).  GPD and LPD are somewhat similar but have less entries by comparison</p>\n<p>Seizure has 2785 entries but only around 6.6% of these have &gt;= 10 votes and around 8% &lt;3.  Most ~85% are in the 3-9 votes.  GRDA and LRDA are less extreme but over 50% in the 3-9 votes range.</p>\n<p>It is possible public LB has greater representation of Other and Seizure since these are the majority in train. <br>\nNotably Seizure as the only pattern with votes make up nearly 80% of entries.  So losing these in splits by number of votes probably is bad for LB even if CV is good due to the influence of Other, etc. </p>",
              "rawMarkdown": "Have been looking at the breakdown of number of votes e.g., < 3 or >=10 etc. and at the high level expert consensus given by unique eeg id (17089) in train.  \n\nOther has the majority of entries 7196 and roughly 40% of these have >=10 votes 37.5% < 3\nso fits in well with the <3 or >=10 idea.  Possibly due to the majority class it has too much influence with >=10 filters (or < 3).  GPD and LPD are somewhat similar but have less entries by comparison\n\nSeizure has 2785 entries but only around 6.6% of these have >= 10 votes and around 8% <3.  Most ~85% are in the 3-9 votes.  GRDA and LRDA are less extreme but over 50% in the 3-9 votes range.\n\nIt is possible public LB has greater representation of Other and Seizure since these are the majority in train. \nNotably Seizure as the only pattern with votes make up nearly 80% of entries.  So losing these in splits by number of votes probably is bad for LB even if CV is good due to the influence of Other, etc. ",
              "votes": 9
            },
            {
              "id": 2713854,
              "postDate": "2024-03-24T13:13:33.420Z",
              "content": "<p>Thanks for insight !!<br>\nAggregating or filtering train.csv influences the \"expert_consensus\" distribution like following plot.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F1dc4366bcc96ca123d710e392df41bdd%2Flabel_dist.png?generation=1711285918774479&amp;alt=media\" alt=\"label_distribution\"></p>",
              "rawMarkdown": "Thanks for insight !!\nAggregating or filtering train.csv influences the \"expert_consensus\" distribution like following plot.\n\n![label_distribution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F1dc4366bcc96ca123d710e392df41bdd%2Flabel_dist.png?generation=1711285918774479&alt=media)",
              "votes": 8
            },
            {
              "id": 2713944,
              "postDate": "2024-03-24T14:42:28.440Z",
              "content": "<p>Thanks for the plot!  </p>",
              "rawMarkdown": "Thanks for the plot!  "
            },
            {
              "id": 2714037,
              "postDate": "2024-03-24T16:08:16.687Z",
              "content": "<p>Thanks for your plot, it is clear to see now the main distribution diff is on Seizure.</p>",
              "rawMarkdown": "Thanks for your plot, it is clear to see now the main distribution diff is on Seizure.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2711260,
          "postDate": "2024-03-22T20:09:10.123Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2725489,
          "postDate": "2024-03-31T15:26:12.253Z",
          "content": "<p>do you apply \"label+0.1666\"？to get cv0.244</p>",
          "rawMarkdown": "do you apply \"label+0.1666\"？to get cv0.244",
          "replies": [
            {
              "id": 2725508,
              "postDate": "2024-03-31T15:41:22.410Z",
              "content": "<p>No I do not use that.</p>",
              "rawMarkdown": "No I do not use that."
            }
          ]
        }
      ]
    },
    {
      "id": 2716441,
      "postDate": "2024-03-26T03:41:13.670Z",
      "content": "<p>actually ignore the kl divergence score and only consider e.g. top1, top2 classification accuracy, they are quite high.<br>\nanother wrong design of evaluation metric?</p>",
      "rawMarkdown": "actually ignore the kl divergence score and only consider e.g. top1, top2 classification accuracy, they are quite high.\nanother wrong design of evaluation metric?"
    },
    {
      "id": 2710298,
      "postDate": "2024-03-22T06:43:00.497Z",
      "content": "<p>In two-stage model, high quality(vote &gt;= 10) CV always same with LB  </p>",
      "rawMarkdown": "In two-stage model, high quality(vote >= 10) CV always same with LB  "
    },
    {
      "id": 2710521,
      "postDate": "2024-03-22T10:27:10.007Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2710168,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-03-22T04:31:59.207000",
      "content": "<p>Our vote count &gt;= 10 OOF score is 0.2174. Top 5 teams probably reached beyond 0.2.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2710462,
          "author_name": "NikhilMishra",
          "author_url": "",
          "post_date": "2024-03-22T09:21:18.293000",
          "content": "<p>is this cv on Chris's created training data filtered out with votes &gt;= 10</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2710463,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-22T09:22:15.347000",
              "content": "<p>Nope, we are using a custom cv scheme.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2710549,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-03-22T10:54:24.913000",
      "content": "<p>Here is the prediction distribution by <code>total_evaluators</code> in oof_df.<br>\nY: kl divergence from label<br>\nX: total_evaluators</p>\n<p>This prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.<br>\nThe baseline model is almost same as <a href=\"https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train\" target=\"_blank\">moth's published notebook</a>.</p>\n<p>As you can see, kl divergence is lager in data with total_evaluators &lt;= 3.<br>\nIt seems that data with total_evaluators &lt;= 3 have difficult sample or noizy sample.<br>\nIn my experiments, this distribution often happens when training baseline efficient net model with all data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F9f2d180d88de96d03a41fa119a0703c2%2Fviz_.png?generation=1711104172826486&amp;alt=media\" alt=\"viz_of_kl_div_by_total_evaluators\"></p>\n<p>I will publish code soon.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2710824,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-03-22T15:23:37.120000",
          "content": "<p>Nice looking plot, but I think the<code>total_evaluators</code> axis might be incorrect. I don't see any columns with 8 or 9 votes.</p>\n<pre><code> pandas  pd\n\ndf= pd.read_csv()\ntotal_votes= df[[, , , , , ]].(axis=)\n(((total_votes.unique())))\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2710834,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-22T15:37:17.317000",
              "content": "<p>This plot is based on <em><code>eeg-unique-df</code></em>, not train.csv<br>\nThe eeg-unique-df is created  in the same way as  in <a href=\"https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train?scriptVersionId=161067167&amp;cellId=11\" target=\"_blank\">chris or moth published notebook</a><br>\nThe eeg-unique-df is made by aggregation, so \"total_evaluators \" column contains float value.<br>\nI use  <code>df.astype(int)</code> to plot this, so this is rough plot of result.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2710871,
          "author_name": "Aurora_blue",
          "author_url": "",
          "post_date": "2024-03-22T16:17:41.940000",
          "content": "<p>2stage learning looks like this.<br>\nThe model is trained in this way, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">EffNetB0 model trained twice, once for each of two training populations - [LB 0.39]</a></p>\n<ul>\n<li>Blue: stage1, CV(total_evaluators &gt; 9)=0.40</li>\n<li>Orange: srage2, CV(total_evaluators &gt; 9)=0.33</li>\n</ul>\n<p>As you can see, in data with total_evaluators &gt; 9, score get better. <br>\nThis 2stage model get around 0.35 in public LB.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F950351391c5b1c9226c18a501a690dce%2Foutput_2.png?generation=1711123884044601&amp;alt=media\" alt=\"viz_compare_2stage\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2710373,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2024-03-22T07:52:23.347000",
      "content": "<p>4 folds votes &gt;= 10 cv 2366 with online maybe 2499 LB.<br>\nOne thing I've noticed is if you change your folds seed as well as other seeds, you might see different cv, sometimes as large as from 256 to 249.<br>\nNow I got 4 folds model LB 0.24 on LB with cv 0.244.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2710563,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2024-03-22T11:05:13.500000",
          "content": "<p>Super score! What do you mean \"with online\"?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2710622,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-03-22T12:07:13.153000",
              "content": "<p>Online just mean submission LB result.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2710922,
          "author_name": "Aurora_blue",
          "author_url": "",
          "post_date": "2024-03-22T16:47:05.083000",
          "content": "<p>Thanks for sharing. <br>\nHow about cv of vote &lt; 10 ? Not calculated ?\nDo you calculate cv in custom way or just filtered oof with vote &gt;= 10 ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2713023,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-03-23T23:04:31.333000",
              "content": "<p>I removed vote &lt; 10 for eval part to speedup. I used custom cv.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2713472,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-24T08:18:45.323000",
              "content": "<blockquote>\n  <p>I used custom cv</p>\n</blockquote>\n<p>thanks. If just filtered vote &gt;= 10, I got 0.27 in CV with vote &gt;= 10, but it's not good in public LB.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2713543,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2024-03-24T09:02:51.520000",
              "content": "<p>Have been looking at the breakdown of number of votes e.g., &lt; 3 or &gt;=10 etc. and at the high level expert consensus given by unique eeg id (17089) in train.  </p>\n<p>Other has the majority of entries 7196 and roughly 40% of these have &gt;=10 votes 37.5% &lt; 3\nso fits in well with the &lt;3 or &gt;=10 idea.  Possibly due to the majority class it has too much influence with &gt;=10 filters (or &lt; 3).  GPD and LPD are somewhat similar but have less entries by comparison</p>\n<p>Seizure has 2785 entries but only around 6.6% of these have &gt;= 10 votes and around 8% &lt;3.  Most ~85% are in the 3-9 votes.  GRDA and LRDA are less extreme but over 50% in the 3-9 votes range.</p>\n<p>It is possible public LB has greater representation of Other and Seizure since these are the majority in train. <br>\nNotably Seizure as the only pattern with votes make up nearly 80% of entries.  So losing these in splits by number of votes probably is bad for LB even if CV is good due to the influence of Other, etc. </p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 2713854,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-24T13:13:33.420000",
              "content": "<p>Thanks for insight !!<br>\nAggregating or filtering train.csv influences the \"expert_consensus\" distribution like following plot.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F1dc4366bcc96ca123d710e392df41bdd%2Flabel_dist.png?generation=1711285918774479&amp;alt=media\" alt=\"label_distribution\"></p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2713944,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2024-03-24T14:42:28.440000",
              "content": "<p>Thanks for the plot!  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2714037,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-03-24T16:08:16.687000",
              "content": "<p>Thanks for your plot, it is clear to see now the main distribution diff is on Seizure.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2711260,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-22T20:09:10.123000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2725489,
          "author_name": "pky",
          "author_url": "",
          "post_date": "2024-03-31T15:26:12.253000",
          "content": "<p>do you apply \"label+0.1666\"？to get cv0.244</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2725508,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-03-31T15:41:22.410000",
              "content": "<p>No I do not use that.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2716441,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-03-26T03:41:13.670000",
      "content": "<p>actually ignore the kl divergence score and only consider e.g. top1, top2 classification accuracy, they are quite high.<br>\nanother wrong design of evaluation metric?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2710298,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2024-03-22T06:43:00.497000",
      "content": "<p>In two-stage model, high quality(vote &gt;= 10) CV always same with LB  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2710521,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-22T10:27:10.007000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2709761": "As most have found by now, there are two groups in our data set. Those who had many votes and those who did not. Calculating CV for the subsets separately we can see that the CV LB gap is actually quite small if we consider that the public LB is mostly high vote data. Of course, there are many assumptions people are making about what the private dataset may hold for us. I thought I would share some of my numbers in hope other may share as well.\n\nHigh votes group CV: 0.33 \nLow votes group CV: 0.6\nLB: 0.35\n\nThis should be a fun competition at the end!",
    "2710168": "Our vote count >= 10 OOF score is 0.2174. Top 5 teams probably reached beyond 0.2.",
    "2710549": "Here is the prediction distribution by `total_evaluators` in oof_df.\nY: kl divergence from label\nX: total_evaluators\n\nThis prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.\nThe baseline model is almost same as [moth's published notebook](https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train).\n\nAs you can see, kl divergence is lager in data with total_evaluators <= 3.\nIt seems that data with total_evaluators <= 3 have difficult sample or noizy sample.\nIn my experiments, this distribution often happens when training baseline efficient net model with all data. \n\n![viz_of_kl_div_by_total_evaluators](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F9f2d180d88de96d03a41fa119a0703c2%2Fviz_.png?generation=1711104172826486&alt=media)\n\nI will publish code soon.",
    "2710373": "4 folds votes >= 10 cv 2366 with online maybe 2499 LB.\nOne thing I've noticed is if you change your folds seed as well as other seeds, you might see different cv, sometimes as large as from 256 to 249.\nNow I got 4 folds model LB 0.24 on LB with cv 0.244.",
    "2716441": "actually ignore the kl divergence score and only consider e.g. top1, top2 classification accuracy, they are quite high.\nanother wrong design of evaluation metric?",
    "2710298": "In two-stage model, high quality(vote >= 10) CV always same with LB  ",
    "2710521": ""
  }
}