{
  "id": 477610,
  "title": "Chris' WaveNet PyTorch Version",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477610",
  "author_name": "",
  "post_date": "2024-02-17T00:53:51.082685600Z",
  "votes": 47,
  "comment_count": 13,
  "views": 0,
  "content": "<p>This is my implementation of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter - [LB 0.52]</a>. I tried to resemble it as closely as possible.</p>\n<p>You can find the code here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train\" target=\"_blank\">HMS | WaveNet PyTorch [Train]</a></li>\n<li><a href=\"https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-inference\" target=\"_blank\">HMS | WaveNet PyTorch [Inference]</a></li>\n</ul>\n<p>This notebook has a little better CV and public LB scores. It achieves 0.77 in CV score and 0.50 score in public LB.</p>\n<p>Hope you like it 👍🏼</p>",
  "messages": [
    {
      "id": "2655513",
      "postDate": "02/17/2024 00:53:51",
      "content": "<p>This is my implementation of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter - [LB 0.52]</a>. I tried to resemble it as closely as possible.</p>\n<p>You can find the code here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train\" target=\"_blank\">HMS | WaveNet PyTorch [Train]</a></li>\n<li><a href=\"https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-inference\" target=\"_blank\">HMS | WaveNet PyTorch [Inference]</a></li>\n</ul>\n<p>This notebook has a little better CV and public LB scores. It achieves 0.77 in CV score and 0.50 score in public LB.</p>\n<p>Hope you like it 👍🏼</p>",
      "rawMarkdown": "This is my implementation of @cdeotte [WaveNet Starter - [LB 0.52]](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52). I tried to resemble it as closely as possible.\n\nYou can find the code here:\n\n- [HMS | WaveNet PyTorch [Train]](https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train)\n- [HMS | WaveNet PyTorch [Inference]](https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-inference)\n\nThis notebook has a little better CV and public LB scores. It achieves 0.77 in CV score and 0.50 score in public LB.\n\nHope you like it 👍🏼",
      "votes": null
    },
    {
      "id": "2655549",
      "postDate": "02/17/2024 01:59:29",
      "content": "<p>Great work! I was just going to look into doing a similar thing!</p>",
      "rawMarkdown": "Great work! I was just going to look into doing a similar thing!",
      "votes": null
    },
    {
      "id": "2655558",
      "postDate": "02/17/2024 02:35:17",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> I hope I saved you some time :)</p>",
      "rawMarkdown": "Thanks @cody11null I hope I saved you some time :)",
      "votes": null
    },
    {
      "id": "2655560",
      "postDate": "02/17/2024 02:45:49",
      "content": "<p>you are awesome .. I was feeling so lazy to do it .. You saved a lot of my time ..</p>",
      "rawMarkdown": "you are awesome .. I was feeling so lazy to do it .. You saved a lot of my time ..",
      "votes": null
    },
    {
      "id": "2655563",
      "postDate": "02/17/2024 02:57:18",
      "content": "<p>Thanks for your kind words, I am glad people find the code useful</p>",
      "rawMarkdown": "Thanks for your kind words, I am glad people find the code useful",
      "votes": null
    },
    {
      "id": "2655807",
      "postDate": "02/17/2024 08:04:10",
      "content": "<p>Amazing! Great work!</p>",
      "rawMarkdown": "Amazing! Great work!",
      "votes": null
    },
    {
      "id": "2655936",
      "postDate": "02/17/2024 10:07:47",
      "content": "<p>Great you are awesome brilliant work and helped us too much</p>",
      "rawMarkdown": "Great you are awesome brilliant work and helped us too much",
      "votes": null
    },
    {
      "id": "2656523",
      "postDate": "02/17/2024 19:19:24",
      "content": "<p>Great work! Saves a lot of time</p>",
      "rawMarkdown": "Great work! Saves a lot of time",
      "votes": null
    },
    {
      "id": "2666190",
      "postDate": "02/24/2024 07:52:04",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "2727770",
      "postDate": "04/01/2024 23:06:10",
      "content": "<p>Thank you for this! We are new to this and have some confusion about why it would achieve so much lower LB score than CV score (0.77 vs 0.50). Are they not both measuring KL divergence -- one on the validation set and one on the test set (which is hidden by Kaggle)? In which case we would expect them to be similar since the model wasn't trained on either of those sets right? </p>",
      "rawMarkdown": "Thank you for this! We are new to this and have some confusion about why it would achieve so much lower LB score than CV score (0.77 vs 0.50). Are they not both measuring KL divergence -- one on the validation set and one on the test set (which is hidden by Kaggle)? In which case we would expect them to be similar since the model wasn't trained on either of those sets right?",
      "votes": null
    },
    {
      "id": "2727790",
      "postDate": "04/01/2024 23:30:03",
      "content": "<p>Hi. See discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hi. See discussion [here][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697",
      "votes": null
    },
    {
      "id": "2727815",
      "postDate": "04/02/2024 00:13:19",
      "content": "<p>Thanks for the quick response! Can I check if I understand correctly?</p>\n<p>The training data from kaggle seems to have merged two training sets, one with a large number of labelers and one small. The test set on the other hand seems to only have samples from the set with the large number of labelers. (This has been inferred because when the training data is divided based on number of voters, models perform similarly on the high voter training set as on the hidden test set).</p>\n<p>In general the high-voter data is easier to learn than the low voter data (ie models generally perform better on the high-voter data). As a result leaderboard scores tend to be higher than scores on randomly sampled validation data.</p>\n<p>Is that the right conclusion?</p>",
      "rawMarkdown": "Thanks for the quick response! Can I check if I understand correctly?\n\nThe training data from kaggle seems to have merged two training sets, one with a large number of labelers and one small. The test set on the other hand seems to only have samples from the set with the large number of labelers. (This has been inferred because when the training data is divided based on number of voters, models perform similarly on the high voter training set as on the hidden test set).\n\nIn general the high-voter data is easier to learn than the low voter data (ie models generally perform better on the high-voter data). As a result leaderboard scores tend to be higher than scores on randomly sampled validation data.\n\nIs that the right conclusion?",
      "votes": null
    },
    {
      "id": "2727825",
      "postDate": "04/02/2024 00:24:10",
      "content": "<p>Yes your understanding is correct. </p>\n<p>And when we compute CV scores, it appears that CV score computed from train data subset <code>votes&gt;=10</code> correlates better with public LB score. And there is less difference between CV score and LB score when doing so.</p>\n<p>(And the original discussion post here computes CV score from <code>all data</code> which explains the large CV score LB score gap).</p>",
      "rawMarkdown": "Yes your understanding is correct. \n\nAnd when we compute CV scores, it appears that CV score computed from train data subset `votes>=10` correlates better with public LB score. And there is less difference between CV score and LB score when doing so.\n\n(And the original discussion post here computes CV score from `all data` which explains the large CV score LB score gap).",
      "votes": null
    },
    {
      "id": "2727828",
      "postDate": "04/02/2024 00:25:45",
      "content": "<p>Got it. Thanks so much!</p>",
      "rawMarkdown": "Got it. Thanks so much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2655549,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "02/17/2024 01:59:29",
      "content": "<p>Great work! I was just going to look into doing a similar thing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655558,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "02/17/2024 02:35:17",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> I hope I saved you some time :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2655560,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "02/17/2024 02:45:49",
      "content": "<p>you are awesome .. I was feeling so lazy to do it .. You saved a lot of my time ..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655563,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "02/17/2024 02:57:18",
          "content": "<p>Thanks for your kind words, I am glad people find the code useful</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2655807,
      "author_name": "aidlre001",
      "author_url": "",
      "post_date": "02/17/2024 08:04:10",
      "content": "<p>Amazing! Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2655936,
      "author_name": "tanishqdublish",
      "author_url": "",
      "post_date": "02/17/2024 10:07:47",
      "content": "<p>Great you are awesome brilliant work and helped us too much</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2656523,
      "author_name": "dabumts",
      "author_url": "",
      "post_date": "02/17/2024 19:19:24",
      "content": "<p>Great work! Saves a lot of time</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2666190,
      "author_name": "jasperpieterse",
      "author_url": "",
      "post_date": "02/24/2024 07:52:04",
      "content": "<p>Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2727770,
      "author_name": "karinana",
      "author_url": "",
      "post_date": "04/01/2024 23:06:10",
      "content": "<p>Thank you for this! We are new to this and have some confusion about why it would achieve so much lower LB score than CV score (0.77 vs 0.50). Are they not both measuring KL divergence -- one on the validation set and one on the test set (which is hidden by Kaggle)? In which case we would expect them to be similar since the model wasn't trained on either of those sets right? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2727790,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/01/2024 23:30:03",
          "content": "<p>Hi. See discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2727815,
              "author_name": "karinana",
              "author_url": "",
              "post_date": "04/02/2024 00:13:19",
              "content": "<p>Thanks for the quick response! Can I check if I understand correctly?</p>\n<p>The training data from kaggle seems to have merged two training sets, one with a large number of labelers and one small. The test set on the other hand seems to only have samples from the set with the large number of labelers. (This has been inferred because when the training data is divided based on number of voters, models perform similarly on the high voter training set as on the hidden test set).</p>\n<p>In general the high-voter data is easier to learn than the low voter data (ie models generally perform better on the high-voter data). As a result leaderboard scores tend to be higher than scores on randomly sampled validation data.</p>\n<p>Is that the right conclusion?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2727825,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "04/02/2024 00:24:10",
                  "content": "<p>Yes your understanding is correct. </p>\n<p>And when we compute CV scores, it appears that CV score computed from train data subset <code>votes&gt;=10</code> correlates better with public LB score. And there is less difference between CV score and LB score when doing so.</p>\n<p>(And the original discussion post here computes CV score from <code>all data</code> which explains the large CV score LB score gap).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2727828,
                      "author_name": "karinana",
                      "author_url": "",
                      "post_date": "04/02/2024 00:25:45",
                      "content": "<p>Got it. Thanks so much!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2655513": "This is my implementation of @cdeotte [WaveNet Starter - [LB 0.52]](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52). I tried to resemble it as closely as possible.\n\nYou can find the code here:\n\n- [HMS | WaveNet PyTorch [Train]](https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train)\n- [HMS | WaveNet PyTorch [Inference]](https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-inference)\n\nThis notebook has a little better CV and public LB scores. It achieves 0.77 in CV score and 0.50 score in public LB.\n\nHope you like it 👍🏼",
    "2655549": "Great work! I was just going to look into doing a similar thing!",
    "2655558": "Thanks @cody11null I hope I saved you some time :)",
    "2655560": "you are awesome .. I was feeling so lazy to do it .. You saved a lot of my time ..",
    "2655563": "Thanks for your kind words, I am glad people find the code useful",
    "2655807": "Amazing! Great work!",
    "2655936": "Great you are awesome brilliant work and helped us too much",
    "2656523": "Great work! Saves a lot of time",
    "2666190": "Thanks a lot!",
    "2727770": "Thank you for this! We are new to this and have some confusion about why it would achieve so much lower LB score than CV score (0.77 vs 0.50). Are they not both measuring KL divergence -- one on the validation set and one on the test set (which is hidden by Kaggle)? In which case we would expect them to be similar since the model wasn't trained on either of those sets right?",
    "2727790": "Hi. See discussion [here][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697",
    "2727815": "Thanks for the quick response! Can I check if I understand correctly?\n\nThe training data from kaggle seems to have merged two training sets, one with a large number of labelers and one small. The test set on the other hand seems to only have samples from the set with the large number of labelers. (This has been inferred because when the training data is divided based on number of voters, models perform similarly on the high voter training set as on the hidden test set).\n\nIn general the high-voter data is easier to learn than the low voter data (ie models generally perform better on the high-voter data). As a result leaderboard scores tend to be higher than scores on randomly sampled validation data.\n\nIs that the right conclusion?",
    "2727825": "Yes your understanding is correct. \n\nAnd when we compute CV scores, it appears that CV score computed from train data subset `votes>=10` correlates better with public LB score. And there is less difference between CV score and LB score when doing so.\n\n(And the original discussion post here computes CV score from `all data` which explains the large CV score LB score gap).",
    "2727828": "Got it. Thanks so much!"
  },
  "source": "meta"
}