{
  "id": 491070,
  "title": "The Importance of Sharing: A Bug-to-LB 0.31 ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/491070",
  "author_name": "Rafael Zimmermann",
  "post_date": "2024-04-04T14:28:20.326000",
  "votes": 30,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hey, everyone!</p>\n<p>I wanted to share something interesting that happened. A while back, I posted here a model that hit an LB of 0.37. It was a pretty cool model, using some label refinement techniques and custom spectrograms that I had put together.</p>\n<p>After I posted it, I ended up exploring other ideas and sort of left that one aside. But then, recently, a colleague from the forum pointed out a mistake in my code, specifically in how I was splitting the data. And he was right - there was indeed a problem there.</p>\n<p>I had used this:<br>\n<code>'sum_votes': 'sum'</code></p>\n<p>And the correct way was this:<br>\n<code>'sum_votes': 'mean',</code></p>\n<p>So, I went back, fixed it, and reran the code to see what would happen. And guess what? It improved! The LB went to 0.31!</p>\n<p>This kind of showed me how important it is for us to share what we do. There's always someone who will appreciate the work and help us improve or see where we're going wrong.</p>\n<p>If I hadn't posted, I might never have noticed this mistake. I want to especially thank <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>, who found the bug. Thanks a lot!</p>\n<p>Below is the updated post, now with an LB of 0.31. In the link, I explain the pipeline in detail:<br>\n<a href=\"https://www.kaggle.com/code/rafaelzimmermann1/no-ensemble-new-spectrograms-label-refine?scriptVersionId=170305079\" target=\"_blank\">https://www.kaggle.com/code/rafaelzimmermann1/no-ensemble-new-spectrograms-label-refine?scriptVersionId=170305079</a></p>",
  "messages": [
    {
      "id": 2735100,
      "postDate": "2024-04-04T14:28:20.327Z",
      "content": "<p>Hey, everyone!</p>\n<p>I wanted to share something interesting that happened. A while back, I posted here a model that hit an LB of 0.37. It was a pretty cool model, using some label refinement techniques and custom spectrograms that I had put together.</p>\n<p>After I posted it, I ended up exploring other ideas and sort of left that one aside. But then, recently, a colleague from the forum pointed out a mistake in my code, specifically in how I was splitting the data. And he was right - there was indeed a problem there.</p>\n<p>I had used this:<br>\n<code>'sum_votes': 'sum'</code></p>\n<p>And the correct way was this:<br>\n<code>'sum_votes': 'mean',</code></p>\n<p>So, I went back, fixed it, and reran the code to see what would happen. And guess what? It improved! The LB went to 0.31!</p>\n<p>This kind of showed me how important it is for us to share what we do. There's always someone who will appreciate the work and help us improve or see where we're going wrong.</p>\n<p>If I hadn't posted, I might never have noticed this mistake. I want to especially thank <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>, who found the bug. Thanks a lot!</p>\n<p>Below is the updated post, now with an LB of 0.31. In the link, I explain the pipeline in detail:<br>\n<a href=\"https://www.kaggle.com/code/rafaelzimmermann1/no-ensemble-new-spectrograms-label-refine?scriptVersionId=170305079\" target=\"_blank\">https://www.kaggle.com/code/rafaelzimmermann1/no-ensemble-new-spectrograms-label-refine?scriptVersionId=170305079</a></p>",
      "rawMarkdown": "Hey, everyone!\n\nI wanted to share something interesting that happened. A while back, I posted here a model that hit an LB of 0.37. It was a pretty cool model, using some label refinement techniques and custom spectrograms that I had put together.\n\nAfter I posted it, I ended up exploring other ideas and sort of left that one aside. But then, recently, a colleague from the forum pointed out a mistake in my code, specifically in how I was splitting the data. And he was right - there was indeed a problem there.\n\nI had used this:\n`'sum_votes': 'sum'`\n\nAnd the correct way was this:\n`'sum_votes': 'mean',`\n\nSo, I went back, fixed it, and reran the code to see what would happen. And guess what? It improved! The LB went to 0.31!\n\nThis kind of showed me how important it is for us to share what we do. There's always someone who will appreciate the work and help us improve or see where we're going wrong.\n\nIf I hadn't posted, I might never have noticed this mistake. I want to especially thank @clearwaterkzk, who found the bug. Thanks a lot!\n\nBelow is the updated post, now with an LB of 0.31. In the link, I explain the pipeline in detail:\nhttps://www.kaggle.com/code/rafaelzimmermann1/no-ensemble-new-spectrograms-label-refine?scriptVersionId=170305079",
      "votes": 29
    },
    {
      "id": 2739727,
      "postDate": "2024-04-07T08:17:15.980Z",
      "content": "<p>Amazing for share this idea!</p>",
      "rawMarkdown": "Amazing for share this idea!"
    },
    {
      "id": 2735116,
      "postDate": "2024-04-04T14:33:44.023Z",
      "content": "<p>That is nice news. I'm delighted that my advice was very helpful to you.</p>",
      "rawMarkdown": "That is nice news. I'm delighted that my advice was very helpful to you.",
      "replies": [
        {
          "id": 2735122,
          "postDate": "2024-04-04T14:41:44.903Z",
          "content": "<p>Yes, you were of great help!</p>",
          "rawMarkdown": "Yes, you were of great help!",
          "votes": -1,
          "replies": [
            {
              "id": 2735169,
              "postDate": "2024-04-04T15:09:14.780Z",
              "content": "<p><a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> <br>\nHowever, it is not recommended to publish high score notebook towards ends.<br>\nI think you should not share model dataset or should change to private, just discussion about improving results may be okay.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fe0ab017ac5145dbc4bb2c27a76c3f4d4%2Falert.png?generation=1712243304694393&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F4c52f98b647dacb7a56e55a74baceb7b%2Falert_2.png?generation=1712243383536214&amp;alt=media\"></p>",
              "rawMarkdown": "@rafaelzimmermann1 \nHowever, it is not recommended to publish high score notebook towards ends.\nI think you should not share model dataset or should change to private, just discussion about improving results may be okay.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fe0ab017ac5145dbc4bb2c27a76c3f4d4%2Falert.png?generation=1712243304694393&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F4c52f98b647dacb7a56e55a74baceb7b%2Falert_2.png?generation=1712243383536214&alt=media)",
              "votes": 7
            },
            {
              "id": 2735175,
              "postDate": "2024-04-04T15:19:32.210Z",
              "content": "<p>I don't see anything wrong with posting a correction for a code written a month ago, even if it improves your LB at the end of the competition. Many people have used it and will appreciate the improvement, and an LB of 0.31 is not at the top of the ranking—it's the ensemble at 0.29.</p>",
              "rawMarkdown": "I don't see anything wrong with posting a correction for a code written a month ago, even if it improves your LB at the end of the competition. Many people have used it and will appreciate the improvement, and an LB of 0.31 is not at the top of the ranking—it's the ensemble at 0.29."
            },
            {
              "id": 2735226,
              "postDate": "2024-04-04T15:47:32.580Z",
              "content": "<p>As in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485352\" target=\"_blank\">previous discussion</a>, publishing any notebook cannot be done now. It's locked like following img.<br>\nOnly updating published notebook can be done.<br>\nThis implies that we can technically do updating, but it also seems that providing any information with code is not recommended near the end.</p>\n<p>Indeed that your score is not top of shared notebook, but it may improve published ensemble notebook and give a big influence in public LB.<br>\nTherefore, it may be very difficult problem about this.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fae95d7c56553cf6f915a4222a8be873a%2F2024-04-05%20003000.png?generation=1712244785588133&amp;alt=media\"></p>",
              "rawMarkdown": "As in [previous discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485352), publishing any notebook cannot be done now. It's locked like following img.\nOnly updating published notebook can be done.\nThis implies that we can technically do updating, but it also seems that providing any information with code is not recommended near the end.\n\nIndeed that your score is not top of shared notebook, but it may improve published ensemble notebook and give a big influence in public LB.\nTherefore, it may be very difficult problem about this.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fae95d7c56553cf6f915a4222a8be873a%2F2024-04-05%20003000.png?generation=1712244785588133&alt=media)",
              "votes": 4
            },
            {
              "id": 2735279,
              "postDate": "2024-04-04T16:27:27.397Z",
              "content": "<p>honestly, the work on that notebook was done more than 1 month ago, it's not just about submitting and posting, you have to add it to a new ensemble, my opinion is that you saw the bug, adjusted the code and don't want it to be shared with your colleagues…. my opinion is that you are being very selfish, be honest, I was the one who created the code, if I shared it, it is because I wanted to share it at its real value…. if others are bothered, I will remove it. just ask, but a lot of people have already copied it, including you</p>\n<p>When you publish something and there is an error, the right thing is to correct it, to be reproducible and scientific, I did this… the crazy thing about all this is that I made this post to thank you for finding the bug, lol</p>",
              "rawMarkdown": "honestly, the work on that notebook was done more than 1 month ago, it's not just about submitting and posting, you have to add it to a new ensemble, my opinion is that you saw the bug, adjusted the code and don't want it to be shared with your colleagues.... my opinion is that you are being very selfish, be honest, I was the one who created the code, if I shared it, it is because I wanted to share it at its real value.... if others are bothered, I will remove it. just ask, but a lot of people have already copied it, including you\n\nWhen you publish something and there is an error, the right thing is to correct it, to be reproducible and scientific, I did this... the crazy thing about all this is that I made this post to thank you for finding the bug, lol",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2735658,
      "postDate": "2024-04-04T20:28:32.550Z",
      "content": "<p>Interesting BTW</p>",
      "rawMarkdown": "Interesting BTW",
      "votes": 1
    },
    {
      "id": 2735798,
      "postDate": "2024-04-04T22:50:03.717Z",
      "content": "<p>Thanks for sharing!<br>\nDoes the cv score also improve?</p>",
      "rawMarkdown": "Thanks for sharing!\nDoes the cv score also improve?",
      "replies": [
        {
          "id": 2735804,
          "postDate": "2024-04-04T23:00:06.457Z",
          "content": "<p>Yes, the cv improves, mainly because the data is separated correctly now, this means that the cv validation is actually based on data with &gt; 10 votes, greatly improving the cv.</p>\n<p>Before 0.53<br>\nnow 0.31 - particularly interesting because it is the same value obtained in LB</p>",
          "rawMarkdown": "Yes, the cv improves, mainly because the data is separated correctly now, this means that the cv validation is actually based on data with > 10 votes, greatly improving the cv.\n\nBefore 0.53\nnow 0.31 - particularly interesting because it is the same value obtained in LB",
          "votes": 1
        }
      ]
    },
    {
      "id": 2736675,
      "postDate": "2024-04-05T10:36:35.150Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 2737883,
      "postDate": "2024-04-06T02:30:25.487Z",
      "content": "<p>Thank you for sharing! </p>",
      "rawMarkdown": "Thank you for sharing! "
    },
    {
      "id": 2735174,
      "postDate": "2024-04-04T15:16:28.020Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2739727,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-07T08:17:15.980000",
      "content": "<p>Amazing for share this idea!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2735116,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-04-04T14:33:44.023000",
      "content": "<p>That is nice news. I'm delighted that my advice was very helpful to you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2735122,
          "author_name": "Rafael Zimmermann",
          "author_url": "",
          "post_date": "2024-04-04T14:41:44.903000",
          "content": "<p>Yes, you were of great help!</p>",
          "votes": -1,
          "replies": [
            {
              "id": 2735169,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-04T15:09:14.780000",
              "content": "<p><a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> <br>\nHowever, it is not recommended to publish high score notebook towards ends.<br>\nI think you should not share model dataset or should change to private, just discussion about improving results may be okay.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fe0ab017ac5145dbc4bb2c27a76c3f4d4%2Falert.png?generation=1712243304694393&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F4c52f98b647dacb7a56e55a74baceb7b%2Falert_2.png?generation=1712243383536214&amp;alt=media\"></p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2735175,
              "author_name": "Rafael Zimmermann",
              "author_url": "",
              "post_date": "2024-04-04T15:19:32.210000",
              "content": "<p>I don't see anything wrong with posting a correction for a code written a month ago, even if it improves your LB at the end of the competition. Many people have used it and will appreciate the improvement, and an LB of 0.31 is not at the top of the ranking—it's the ensemble at 0.29.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2735226,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-04T15:47:32.580000",
              "content": "<p>As in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485352\" target=\"_blank\">previous discussion</a>, publishing any notebook cannot be done now. It's locked like following img.<br>\nOnly updating published notebook can be done.<br>\nThis implies that we can technically do updating, but it also seems that providing any information with code is not recommended near the end.</p>\n<p>Indeed that your score is not top of shared notebook, but it may improve published ensemble notebook and give a big influence in public LB.<br>\nTherefore, it may be very difficult problem about this.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fae95d7c56553cf6f915a4222a8be873a%2F2024-04-05%20003000.png?generation=1712244785588133&amp;alt=media\"></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2735279,
              "author_name": "Rafael Zimmermann",
              "author_url": "",
              "post_date": "2024-04-04T16:27:27.397000",
              "content": "<p>honestly, the work on that notebook was done more than 1 month ago, it's not just about submitting and posting, you have to add it to a new ensemble, my opinion is that you saw the bug, adjusted the code and don't want it to be shared with your colleagues…. my opinion is that you are being very selfish, be honest, I was the one who created the code, if I shared it, it is because I wanted to share it at its real value…. if others are bothered, I will remove it. just ask, but a lot of people have already copied it, including you</p>\n<p>When you publish something and there is an error, the right thing is to correct it, to be reproducible and scientific, I did this… the crazy thing about all this is that I made this post to thank you for finding the bug, lol</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2735658,
      "author_name": "Md Raihan Khan",
      "author_url": "",
      "post_date": "2024-04-04T20:28:32.550000",
      "content": "<p>Interesting BTW</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2735798,
      "author_name": "MBOOK",
      "author_url": "",
      "post_date": "2024-04-04T22:50:03.717000",
      "content": "<p>Thanks for sharing!<br>\nDoes the cv score also improve?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2735804,
          "author_name": "Rafael Zimmermann",
          "author_url": "",
          "post_date": "2024-04-04T23:00:06.457000",
          "content": "<p>Yes, the cv improves, mainly because the data is separated correctly now, this means that the cv validation is actually based on data with &gt; 10 votes, greatly improving the cv.</p>\n<p>Before 0.53<br>\nnow 0.31 - particularly interesting because it is the same value obtained in LB</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2736675,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-05T10:36:35.150000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2737883,
      "author_name": "Penny_BChen",
      "author_url": "",
      "post_date": "2024-04-06T02:30:25.487000",
      "content": "<p>Thank you for sharing! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2735174,
      "author_name": "Danial Zakaria",
      "author_url": "",
      "post_date": "2024-04-04T15:16:28.020000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2735100": "Hey, everyone!\n\nI wanted to share something interesting that happened. A while back, I posted here a model that hit an LB of 0.37. It was a pretty cool model, using some label refinement techniques and custom spectrograms that I had put together.\n\nAfter I posted it, I ended up exploring other ideas and sort of left that one aside. But then, recently, a colleague from the forum pointed out a mistake in my code, specifically in how I was splitting the data. And he was right - there was indeed a problem there.\n\nI had used this:\n`'sum_votes': 'sum'`\n\nAnd the correct way was this:\n`'sum_votes': 'mean',`\n\nSo, I went back, fixed it, and reran the code to see what would happen. And guess what? It improved! The LB went to 0.31!\n\nThis kind of showed me how important it is for us to share what we do. There's always someone who will appreciate the work and help us improve or see where we're going wrong.\n\nIf I hadn't posted, I might never have noticed this mistake. I want to especially thank @clearwaterkzk, who found the bug. Thanks a lot!\n\nBelow is the updated post, now with an LB of 0.31. In the link, I explain the pipeline in detail:\nhttps://www.kaggle.com/code/rafaelzimmermann1/no-ensemble-new-spectrograms-label-refine?scriptVersionId=170305079",
    "2739727": "Amazing for share this idea!",
    "2735116": "That is nice news. I'm delighted that my advice was very helpful to you.",
    "2735658": "Interesting BTW",
    "2735798": "Thanks for sharing!\nDoes the cv score also improve?",
    "2736675": "",
    "2737883": "Thank you for sharing! ",
    "2735174": "Thank you for sharing!"
  }
}