{
  "id": 207003,
  "title": "Let's stop publishing easy to fork high scoring notebooks.",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207003",
  "author_name": "Manikanth Reddy",
  "post_date": "2020-12-27T16:22:28.569000",
  "votes": 63,
  "comment_count": 47,
  "views": 0,
  "content": "<p>Hi all, I think it's about time we stop posting high scoring notebooks (one's that are easy to fork and submit without adding any contribution). We have about only 10 days left and I don't think it's fair anymore to publish any high scoring notebooks. </p>\n<ul>\n<li>In LISH MoA competition, 1 high scoring public notebook released in the last few days ended up contributing to most of the silver and bronze medals. People who have worked hard got frustrated (check <a href=\"https://www.kaggle.com/c/lish-moa/discussion/200586\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/c/lish-moa/discussion/200535\" target=\"_blank\">this</a>). This competition is going through the same phase and might even become worse. (I myself lost 500+ positions last week since I was <strong>on vacation</strong>. I want to work on this one till the end but I do understand that others may not be able to work till the end. There are a lot of people who have worked on this from the very beginning and might be on vacation due to the year end).</li>\n<li>All the latest high scoring notebooks are just <strong>forks of forks of forks</strong> (Just look at the notebook names) with funda of <strong>\"Do hyper-parameter tuning and make the notebook public\"</strong>. I think this is <strong>not adding any value</strong>. </li>\n<li><strong>1200 forks</strong> from just 2 notebooks (900, 300), this is scary stuff. <strong>This has become rat race (endless, self-defeating, or pointless pursuit) to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else</strong>. Those who are just forking and submitting public notebooks without actually working on this competition, you guys have my pity (I don't enjoy playing my video games from someone else's <em>final checkpoint</em>, I want to beat the final boss by myself). This competition is one of the few gems for learning feature engineering, understanding how lgbm is too different from xgboost and catboost in its tree building, how transformers work and the most important - building efficient pipelines with limited RAM and TIME. If you haven't learnt these….</li>\n</ul>\n<hr>\n<p>For those of you who are misunderstanding this post:</p>\n<ul>\n<li> Don't share easy to fork notebooks in the last days of the competition. </li>\n<li>People should understand that this is not an easy competition. A lot of people have been working hard for months and it's not fair for them to loose to someone who didn't work hard but ended up submitting someone else's work in the last days. </li>\n<li>Kaggle is as much of competition site as that of a learning site for us. Please try to appreciate it. Learning can wait few days if people's hardwork is at stake. </li>\n</ul>\n<hr>\n<p>So please I request all of you who are considering making high scoring public notebooks, please wait till the end of the competition. Thanks 😃</p>",
  "messages": [
    {
      "id": 1128648,
      "postDate": "2020-12-27T16:22:28.570Z",
      "content": "<p>Hi all, I think it's about time we stop posting high scoring notebooks (one's that are easy to fork and submit without adding any contribution). We have about only 10 days left and I don't think it's fair anymore to publish any high scoring notebooks. </p>\n<ul>\n<li>In LISH MoA competition, 1 high scoring public notebook released in the last few days ended up contributing to most of the silver and bronze medals. People who have worked hard got frustrated (check <a href=\"https://www.kaggle.com/c/lish-moa/discussion/200586\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/c/lish-moa/discussion/200535\" target=\"_blank\">this</a>). This competition is going through the same phase and might even become worse. (I myself lost 500+ positions last week since I was <strong>on vacation</strong>. I want to work on this one till the end but I do understand that others may not be able to work till the end. There are a lot of people who have worked on this from the very beginning and might be on vacation due to the year end).</li>\n<li>All the latest high scoring notebooks are just <strong>forks of forks of forks</strong> (Just look at the notebook names) with funda of <strong>\"Do hyper-parameter tuning and make the notebook public\"</strong>. I think this is <strong>not adding any value</strong>. </li>\n<li><strong>1200 forks</strong> from just 2 notebooks (900, 300), this is scary stuff. <strong>This has become rat race (endless, self-defeating, or pointless pursuit) to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else</strong>. Those who are just forking and submitting public notebooks without actually working on this competition, you guys have my pity (I don't enjoy playing my video games from someone else's <em>final checkpoint</em>, I want to beat the final boss by myself). This competition is one of the few gems for learning feature engineering, understanding how lgbm is too different from xgboost and catboost in its tree building, how transformers work and the most important - building efficient pipelines with limited RAM and TIME. If you haven't learnt these….</li>\n</ul>\n<hr>\n<p>For those of you who are misunderstanding this post:</p>\n<ul>\n<li> Don't share easy to fork notebooks in the last days of the competition. </li>\n<li>People should understand that this is not an easy competition. A lot of people have been working hard for months and it's not fair for them to loose to someone who didn't work hard but ended up submitting someone else's work in the last days. </li>\n<li>Kaggle is as much of competition site as that of a learning site for us. Please try to appreciate it. Learning can wait few days if people's hardwork is at stake. </li>\n</ul>\n<hr>\n<p>So please I request all of you who are considering making high scoring public notebooks, please wait till the end of the competition. Thanks 😃</p>",
      "rawMarkdown": "Hi all, I think it's about time we stop posting high scoring notebooks (one's that are easy to fork and submit without adding any contribution). We have about only 10 days left and I don't think it's fair anymore to publish any high scoring notebooks. \n- In LISH MoA competition, 1 high scoring public notebook released in the last few days ended up contributing to most of the silver and bronze medals. People who have worked hard got frustrated (check [this](https://www.kaggle.com/c/lish-moa/discussion/200586) and [this](https://www.kaggle.com/c/lish-moa/discussion/200535)). This competition is going through the same phase and might even become worse. (I myself lost 500+ positions last week since I was **on vacation**. I want to work on this one till the end but I do understand that others may not be able to work till the end. There are a lot of people who have worked on this from the very beginning and might be on vacation due to the year end).\n- All the latest high scoring notebooks are just **forks of forks of forks** (Just look at the notebook names) with funda of **\"Do hyper-parameter tuning and make the notebook public\"**. I think this is **not adding any value**. \n- **1200 forks** from just 2 notebooks (900, 300), this is scary stuff. **This has become rat race (endless, self-defeating, or pointless pursuit) to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else**. Those who are just forking and submitting public notebooks without actually working on this competition, you guys have my pity (I don't enjoy playing my video games from someone else's *final checkpoint*, I want to beat the final boss by myself). This competition is one of the few gems for learning feature engineering, understanding how lgbm is too different from xgboost and catboost in its tree building, how transformers work and the most important - building efficient pipelines with limited RAM and TIME. If you haven't learnt these....\n\n---\n\nFor those of you who are misunderstanding this post:\n- ~~Don't share your work~~ Don't share easy to fork notebooks in the last days of the competition. \n- People should understand that this is not an easy competition. A lot of people have been working hard for months and it's not fair for them to loose to someone who didn't work hard but ended up submitting someone else's work in the last days. \n- Kaggle is as much of competition site as that of a learning site for us. Please try to appreciate it. Learning can wait few days if people's hardwork is at stake. \n\n---\n\nSo please I request all of you who are considering making high scoring public notebooks, please wait till the end of the competition. Thanks 😃",
      "votes": 63
    },
    {
      "id": 1128786,
      "postDate": "2020-12-27T18:31:54.020Z",
      "content": "<p>Might be an unpopular opinion (hot take?), but this line of reasoning confuses me.</p>\n<p>In the past, there was no requirements whatsoever. This makes sense from Kaggle's perspective because their goal is to have the best result for their clients, and the hosts too want the best result for the amount of time (money) they have spent retaining Kaggle.</p>\n<p>Then eventually through community discussion, Kaggle added requirement to disclose non-competition specific details such as external datasets within 7 days before competition end. The driving force here was also in-line with Kaggle's overarching objective, of increasing the overall quality of submissions at the end of competition for their client.</p>\n<p>Again through more maturing communal discourse, Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients (not us competing data scientists) so no wonder it was a \"soft\" suggestion and no one gets banned for breaking it.</p>\n<p>We aren't at t-7 days to competition close yet. If someone shares something amazing, either on the forums or by means of a kernel, what is the problem? In MISH people were complaining 20 days before competition end and now here people are complaining 10 days before competition end. Kaggle already has the warning for t-7 days before competition end, do we need more regulations than that? If so, discuss on the global kaggle suggestions forums rather than competition specific forums, since it's the same conversation that keeps coming up month after month. I too was crushed by the public MISH kernels and God knows how many hours I've pushed into riiid (and spousal fights as a result of it). But if your goal is truly to learn, then learn from wherever the knowledge comes from. And if that knowledge comes t-7 days before the competition end date, as was agreed upon by the community guidelines, then be grateful that someone shared something useful and see how best you can integrate it into your solution. If the ground breaking happens after that then sure that sucks. But trying to restrict the flow of ideas -10 or -20 days before… that makes no sense to me. Why not just restrict sharing entirely \\s</p>",
      "rawMarkdown": "Might be an unpopular opinion (hot take?), but this line of reasoning confuses me.\n\nIn the past, there was no requirements whatsoever. This makes sense from Kaggle's perspective because their goal is to have the best result for their clients, and the hosts too want the best result for the amount of time (money) they have spent retaining Kaggle.\n\nThen eventually through community discussion, Kaggle added requirement to disclose non-competition specific details such as external datasets within 7 days before competition end. The driving force here was also in-line with Kaggle's overarching objective, of increasing the overall quality of submissions at the end of competition for their client.\n\nAgain through more maturing communal discourse, Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients (not us competing data scientists) so no wonder it was a \"soft\" suggestion and no one gets banned for breaking it.\n\nWe aren't at t-7 days to competition close yet. If someone shares something amazing, either on the forums or by means of a kernel, what is the problem? In MISH people were complaining 20 days before competition end and now here people are complaining 10 days before competition end. Kaggle already has the warning for t-7 days before competition end, do we need more regulations than that? If so, discuss on the global kaggle suggestions forums rather than competition specific forums, since it's the same conversation that keeps coming up month after month. I too was crushed by the public MISH kernels and God knows how many hours I've pushed into riiid (and spousal fights as a result of it). But if your goal is truly to learn, then learn from wherever the knowledge comes from. And if that knowledge comes t-7 days before the competition end date, as was agreed upon by the community guidelines, then be grateful that someone shared something useful and see how best you can integrate it into your solution. If the ground breaking happens after that then sure that sucks. But trying to restrict the flow of ideas -10 or -20 days before... that makes no sense to me. Why not just restrict sharing entirely \\s",
      "votes": 7,
      "replies": [
        {
          "id": 1128790,
          "postDate": "2020-12-27T18:35:22.927Z",
          "content": "<p>Btw the people who have discovered something crazy and share for discussion / notebook reputation points, aren't these the same people who would likely private share anyway? If they're gonna do that, better at least to have it done public rather than just benefit a small in-circle group of people. Sorry just my 2cents.</p>",
          "rawMarkdown": "Btw the people who have discovered something crazy and share for discussion / notebook reputation points, aren't these the same people who would likely private share anyway? If they're gonna do that, better at least to have it done public rather than just benefit a small in-circle group of people. Sorry just my 2cents."
        },
        {
          "id": 1128800,
          "postDate": "2020-12-27T18:51:34.183Z",
          "content": "<p>Great points! And yes, if someone s here for learning, then this competition is a great opportunity! It just can't get better than this. Cool LB/CV sync, lot of cool notebooks/ideas/threads/papers/self ideas etc to try and see how it works and then enjoy being on the top-k% of the LB. Till here it's good but the trouble comes when people who just fork not with the intention to learn anything about what the author did etc but more with the intention to later share and be proud of their fake medal, they harm everyone, including the one's who have actually put into efforts, tremendous efforts rather. The analogy is similar to merits/demerits of using Internet let's say, how you use it, tells a lot about you as a person.</p>\n<p>And in case someone has shared something cool, look into it with a microscope and see how fast you can integrate and test the same pretty much! If it helps, keep it, if it doesn't, then you already have what you need!</p>\n<p>It's very important for people to communicate their ideas, you never know how that idea can be bended and take a new form! And avoid private sharing, there's no point. So if restricted, then Idk what will happen and that will give rise to more private sharing for sure.</p>\n<p>Speaking for myself, have learnt a ton of things from this comp, writing cool code, tests, quick and easily manageable pipelines, managing such a vol of data and many other things!</p>\n<p>Also, don't focus too much on LB scores, focus on your CV as public notebooks might bring you down in private LB.</p>",
          "rawMarkdown": "Great points! And yes, if someone s here for learning, then this competition is a great opportunity! It just can't get better than this. Cool LB/CV sync, lot of cool notebooks/ideas/threads/papers/self ideas etc to try and see how it works and then enjoy being on the top-k% of the LB. Till here it's good but the trouble comes when people who just fork not with the intention to learn anything about what the author did etc but more with the intention to later share and be proud of their fake medal, they harm everyone, including the one's who have actually put into efforts, tremendous efforts rather. The analogy is similar to merits/demerits of using Internet let's say, how you use it, tells a lot about you as a person.\n\n\nAnd in case someone has shared something cool, look into it with a microscope and see how fast you can integrate and test the same pretty much! If it helps, keep it, if it doesn't, then you already have what you need!\n\nIt's very important for people to communicate their ideas, you never know how that idea can be bended and take a new form! And avoid private sharing, there's no point. So if restricted, then Idk what will happen and that will give rise to more private sharing for sure.\n\n\nSpeaking for myself, have learnt a ton of things from this comp, writing cool code, tests, quick and easily manageable pipelines, managing such a vol of data and many other things!\n\n\n\nAlso, don't focus too much on LB scores, focus on your CV as public notebooks might bring you down in private LB.",
          "votes": 3
        },
        {
          "id": 1128857,
          "postDate": "2020-12-27T20:03:10.610Z",
          "content": "<p>You argument doesn't make sense. Learning is important, but the motivation to win the competition is also an important factor of Kaggle.</p>\n<p>Following your argument, why not ask Kaggle to cancel the competition, but only running as an learning platform? Do you think there will be the same amount of great people coming here to contribute?</p>\n<p>And about learning, where is the learning when people just fork and submit? Or even fork someone's kernel and publish it as he/she is the original author?</p>\n<p>This is not a 0 or 1 situation. We want people to win, but also want to respect and evaluate people's hard work! Some kind of rules should be maintained anyway.</p>",
          "rawMarkdown": "You argument doesn't make sense. Learning is important, but the motivation to win the competition is also an important factor of Kaggle.\n\nFollowing your argument, why not ask Kaggle to cancel the competition, but only running as an learning platform? Do you think there will be the same amount of great people coming here to contribute?\n\nAnd about learning, where is the learning when people just fork and submit? Or even fork someone's kernel and publish it as he/she is the original author?\n\nThis is not a 0 or 1 situation. We want people to win, but also want to respect and evaluate people's hard work! Some kind of rules should be maintained anyway.",
          "votes": 1
        },
        {
          "id": 1128865,
          "postDate": "2020-12-27T20:13:06.197Z",
          "content": "<p>I am not disagree with you. All I am saying is that the very gray and arbitrary \"too late to share\" cutoff which is unofficially renegotiated at every competition by participants should be done at the global kaggle community level (as been done in the past) if people really want to have hard deadlines. There already exist a cutoff period that the community and Kaggle have agreed to, which is t-7; so why should people be upset if someone shares something good at t-8? or t-10? The people who have worked hard like you and others, can take whatever is shared and bolster your own solutions further propelling you forward. The people who just fork and haven't put in any work, there will be a limit to how far they can go—both on this platform as well as in life..</p>\n<p>And just statistically speaking, the number of people in the top ranks for any given competition is extremely limited. I know well the desire to win (anyone who's competed for years is right there with you, we're all addicted). But if the goal in competing at Kaggle is <em>just</em> to win, or even <em>primarily</em> to win, then I daresay 99.9% of the people on this platform myself included are failing horribly.</p>",
          "rawMarkdown": "I am not disagree with you. All I am saying is that the very gray and arbitrary \"too late to share\" cutoff which is unofficially renegotiated at every competition by participants should be done at the global kaggle community level (as been done in the past) if people really want to have hard deadlines. There already exist a cutoff period that the community and Kaggle have agreed to, which is t-7; so why should people be upset if someone shares something good at t-8? or t-10? The people who have worked hard like you and others, can take whatever is shared and bolster your own solutions further propelling you forward. The people who just fork and haven't put in any work, there will be a limit to how far they can go—both on this platform as well as in life..\n\nAnd just statistically speaking, the number of people in the top ranks for any given competition is extremely limited. I know well the desire to win (anyone who's competed for years is right there with you, we're all addicted). But if the goal in competing at Kaggle is *just* to win, or even *primarily* to win, then I daresay 99.9% of the people on this platform myself included are failing horribly."
        },
        {
          "id": 1128871,
          "postDate": "2020-12-27T20:26:04.987Z",
          "content": "<p>Well, for myself,  even if I don't win in a competition, I still hope the results reflect better the reality. For example, say I am near the bronze zone (say rank 300). And someone publish a high score kernel in a silver zone? And his pipeline is totally different from mine (or even, the framework I never used before)?<br>\nIn this case, either I fork blindly, or I play honest and get a low rank, say 1000?</p>",
          "rawMarkdown": "Well, for myself,  even if I don't win in a competition, I still hope the results reflect better the reality. For example, say I am near the bronze zone (say rank 300). And someone publish a high score kernel in a silver zone? And his pipeline is totally different from mine (or even, the framework I never used before)?\nIn this case, either I fork blindly, or I play honest and get a low rank, say 1000?",
          "votes": 1
        },
        {
          "id": 1128920,
          "postDate": "2020-12-27T22:10:55.613Z",
          "content": "<blockquote>\n  <p>Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients</p>\n</blockquote>\n<p>I strongly disagree with this assumption. </p>\n<p>What makes the value out of the competitions (for clients, users, everyone) it's to have thousands of minds working on the same problem each one from a different angle, contributing to original, creative, innovative approaches.    <br>\nThe spreading of high scoring kernels only attracted users who try to find a shortcut and hope they can win a medal by slightly tweaking some parameters of an existing script. This is <strong>not</strong> of any value for anyone, I am sorry. </p>",
          "rawMarkdown": "> Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients\n\nI strongly disagree with this assumption. \n\nWhat makes the value out of the competitions (for clients, users, everyone) it's to have thousands of minds working on the same problem each one from a different angle, contributing to original, creative, innovative approaches.    \nThe spreading of high scoring kernels only attracted users who try to find a shortcut and hope they can win a medal by slightly tweaking some parameters of an existing script. This is **not** of any value for anyone, I am sorry. ",
          "votes": 3
        },
        {
          "id": 1128961,
          "postDate": "2020-12-27T23:57:31.227Z",
          "content": "<p>That's nice and great to hold as conjecture, but the point you raise has never been realized as policy, neither soft or otherwise. Whereas all the points I raised were <strong>historical</strong>.</p>\n<p>So either the people in charge do in fact feel there is benefit to doing it the way they've been doing it; or perhaps, as I've repeated now for the third time, that the proper place to bring up policy changes across Kaggle as a whole would be the suggestions / feedback forums, which have been purposely built for exactly this type of discourse.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc70667212f0f75e6b17b7e4385a02327%2FScreen%20Shot%202020-12-27%20at%207.16.50%20PM.png?generation=1609118245701021&amp;alt=media\" alt=\"\"></p>\n<p>Last thing I'll mention about this. While we did have a \"high scoring\" - if we can call it that - kernel published 2 days ago, the reality of the matter is another kernel with the exact same score was available 10 days ago. And we can see sorting by best public kernel that there really hasn't been any action in the last +14 day. In other words, almost no real sharing for quite a while. If people would be more open, we'd see a lot more interesting solutions imo, especially since there def is some good FE going on as evidenced by the public LB.</p>",
          "rawMarkdown": "That's nice and great to hold as conjecture, but the point you raise has never been realized as policy, neither soft or otherwise. Whereas all the points I raised were **historical**.\n\nSo either the people in charge do in fact feel there is benefit to doing it the way they've been doing it; or perhaps, as I've repeated now for the third time, that the proper place to bring up policy changes across Kaggle as a whole would be the suggestions / feedback forums, which have been purposely built for exactly this type of discourse.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc70667212f0f75e6b17b7e4385a02327%2FScreen%20Shot%202020-12-27%20at%207.16.50%20PM.png?generation=1609118245701021&alt=media)\n\nLast thing I'll mention about this. While we did have a \"high scoring\" - if we can call it that - kernel published 2 days ago, the reality of the matter is another kernel with the exact same score was available 10 days ago. And we can see sorting by best public kernel that there really hasn't been any action in the last +14 day. In other words, almost no real sharing for quite a while. If people would be more open, we'd see a lot more interesting solutions imo, especially since there def is some good FE going on as evidenced by the public LB.",
          "votes": 1
        },
        {
          "id": 1129021,
          "postDate": "2020-12-28T02:14:10.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> you have raised great points. I think what we are discussing is in the grey area. If a kernel is published now on something that is genuinely good work and is helpful to everyone, then everyone would be happy and take it in a positive way. All I am concerned about is about the fork of fork of forks. If you look at the the number of forks vs upvotes, it's too disproportionate. I still think that it depends on the intention of the author. They have to weight the positives vs negatives before publishing one. All I want from this post is to remind them of that. If it's something that crushes leaderboard in the last few days, it can wait till the end of the competition. No one complains after that. </p>",
          "rawMarkdown": "@authman you have raised great points. I think what we are discussing is in the grey area. If a kernel is published now on something that is genuinely good work and is helpful to everyone, then everyone would be happy and take it in a positive way. All I am concerned about is about the fork of fork of forks. If you look at the the number of forks vs upvotes, it's too disproportionate. I still think that it depends on the intention of the author. They have to weight the positives vs negatives before publishing one. All I want from this post is to remind them of that. If it's something that crushes leaderboard in the last few days, it can wait till the end of the competition. No one complains after that. ",
          "votes": 1
        },
        {
          "id": 1129679,
          "postDate": "2020-12-28T14:03:02.800Z",
          "content": "<blockquote>\n  <p>that the proper place to bring up policy changes across Kaggle as a whole would be the suggestions / feedback forums</p>\n</blockquote>\n<p>And so what? Why can't we have this discussion here too? If good points are raised, are they illegitimate just because they are not posted in the optimal channel?</p>\n<blockquote>\n  <p>That's nice and great to hold as conjecture, but the point you raise has never been realized as policy, neither soft or otherwise</p>\n</blockquote>\n<p>3 people replied to your comment, who and what are you referring to? </p>\n<blockquote>\n  <p>And we can see sorting by best public kernel that there really hasn't been any action in the last +14 day. In other words, almost no real sharing for quite a while</p>\n</blockquote>\n<p>If you read the post carefully, there is an invitation to stop publishing high scoring kernels from now on. Which, incidentally, is upvoted by 44 people already (as of writing).</p>",
          "rawMarkdown": "> that the proper place to bring up policy changes across Kaggle as a whole would be the suggestions / feedback forums\n\nAnd so what? Why can't we have this discussion here too? If good points are raised, are they illegitimate just because they are not posted in the optimal channel?\n\n> That's nice and great to hold as conjecture, but the point you raise has never been realized as policy, neither soft or otherwise\n\n3 people replied to your comment, who and what are you referring to? \n\n> And we can see sorting by best public kernel that there really hasn't been any action in the last +14 day. In other words, almost no real sharing for quite a while\n\nIf you read the post carefully, there is an invitation to stop publishing high scoring kernels from now on. Which, incidentally, is upvoted by 44 people already (as of writing)."
        }
      ]
    },
    {
      "id": 1129734,
      "postDate": "2020-12-28T14:15:08.690Z",
      "content": "<p>I think the main issue is not the public kernels with high scores pubished last minute but more the people that blindly fork and resubmit. I always like to pick up ideas from top kernels, even if it is at the last minute if someone completly outscore my own submission.</p>\n<p>What I don't like is seeing 500 persons climbing the ladder with absolutly no work/no merite and just hunting for the medal ☹️</p>",
      "rawMarkdown": "I think the main issue is not the public kernels with high scores pubished last minute but more the people that blindly fork and resubmit. I always like to pick up ideas from top kernels, even if it is at the last minute if someone completly outscore my own submission.\n\nWhat I don't like is seeing 500 persons climbing the ladder with absolutly no work/no merite and just hunting for the medal ☹️",
      "votes": 6,
      "replies": [
        {
          "id": 1129756,
          "postDate": "2020-12-28T14:29:56.343Z",
          "content": "<p>full agree ,what can learn from the blindly fork notebook ? turing hyperparameter?</p>",
          "rawMarkdown": "full agree ,what can learn from the blindly fork notebook ? turing hyperparameter?"
        },
        {
          "id": 1129774,
          "postDate": "2020-12-28T14:40:24.903Z",
          "content": "<p>Yeah… <br>\nOne way of possibly going would be to disable fork of notebooks from a certain moment in the competition and having no possibility to c/c cells… </p>\n<p>It could be still possible to see solutions of others, but it would requiert to actually recode it to include it in your own solution. </p>\n<p>Still possible to recode a full kernel, but I'm sure it would discourage most of the serial resubmiters 😄</p>",
          "rawMarkdown": "Yeah... \nOne way of possibly going would be to disable fork of notebooks from a certain moment in the competition and having no possibility to c/c cells... \n\nIt could be still possible to see solutions of others, but it would requiert to actually recode it to include it in your own solution. \n\nStill possible to recode a full kernel, but I'm sure it would discourage most of the serial resubmiters 😄"
        }
      ]
    },
    {
      "id": 1128725,
      "postDate": "2020-12-27T17:18:26.443Z",
      "content": "<p>Totally agree. I was thinking of writing a similar post. If we look at public notebooks sorted by hotness only few of them are quite interesting but the others are fork of forks.. I think there is a still an important gap between 0.78x and 0.79x but I'm pretty afraid of waking-up a morning and seeing a public kernel scoring 0.79x. <br>\nHowever I remain optimistic that people who worked hard will be rewarded </p>",
      "rawMarkdown": "Totally agree. I was thinking of writing a similar post. If we look at public notebooks sorted by hotness only few of them are quite interesting but the others are fork of forks.. I think there is a still an important gap between 0.78x and 0.79x but I'm pretty afraid of waking-up a morning and seeing a public kernel scoring 0.79x. \nHowever I remain optimistic that people who worked hard will be rewarded ",
      "votes": 3,
      "replies": [
        {
          "id": 1128732,
          "postDate": "2020-12-27T17:21:28.940Z",
          "content": "<p>Yeah, this is what I am most worried about. 0.79x is not achievable without hard work. I am trying my best but still in the 0.77x. I haven't finished my LGBM and SAINT+. I want to hope that no one releases their saint model now. </p>",
          "rawMarkdown": "Yeah, this is what I am most worried about. 0.79x is not achievable without hard work. I am trying my best but still in the 0.77x. I haven't finished my LGBM and SAINT+. I want to hope that no one releases their saint model now. ",
          "votes": 2
        },
        {
          "id": 1128742,
          "postDate": "2020-12-27T17:28:00.840Z",
          "content": "<p>Releasing a high scoring SAINT in the last 10 days of competition would probably destroy the LB</p>",
          "rawMarkdown": "Releasing a high scoring SAINT in the last 10 days of competition would probably destroy the LB"
        }
      ]
    },
    {
      "id": 1128746,
      "postDate": "2020-12-27T17:34:29.273Z",
      "content": "<p>What is most surprising to me it's the little Kaggle has done to prevent this from happening… Given it's an outstanding issue since years…</p>",
      "rawMarkdown": "What is most surprising to me it's the little Kaggle has done to prevent this from happening... Given it's an outstanding issue since years...",
      "votes": 4
    },
    {
      "id": 1128749,
      "postDate": "2020-12-27T17:37:31.763Z",
      "content": "<p>It was lucky to participate in this competition for me ,learned a lot.I'm trying my SANIT model though time is not enough.<br>\nMy best time so far is 0.782, and it looks like I'll be dropping crazy places again tomorrow。</p>",
      "rawMarkdown": "It was lucky to participate in this competition for me ,learned a lot.I'm trying my SANIT model though time is not enough.\nMy best time so far is 0.782, and it looks like I'll be dropping crazy places again tomorrow。",
      "votes": 2,
      "replies": [
        {
          "id": 1128760,
          "postDate": "2020-12-27T17:52:35.267Z",
          "content": "<p>Don't give up. I think single SAINT could beat LightGBM and SAKT ensemble hands down. </p>",
          "rawMarkdown": "Don't give up. I think single SAINT could beat LightGBM and SAKT ensemble hands down. "
        },
        {
          "id": 1128869,
          "postDate": "2020-12-27T20:25:10.497Z",
          "content": "<p>of course, I think I can do better with single lgbm ,0.782 not end . and try SAINT then ensemble .</p>",
          "rawMarkdown": "of course, I think I can do better with single lgbm ,0.782 not end . and try SAINT then ensemble .",
          "votes": 1
        },
        {
          "id": 1129023,
          "postDate": "2020-12-28T02:15:50.390Z",
          "content": "<p>Yeah. It seems atleast  <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801#1116607\" target=\"_blank\">0.786</a> is achievable using LightGBM with proper feature engineering. </p>",
          "rawMarkdown": "Yeah. It seems atleast ~~0.887~~ [0.786](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801#1116607) is achievable using LightGBM with proper feature engineering. "
        },
        {
          "id": 1129030,
          "postDate": "2020-12-28T02:37:48.697Z",
          "content": "<p>lol I think you mean 0.787 not 0.887 <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a> . But even 0.787 seems hard to achieve, for my case using LGBM, but a bit easier using SAINT.</p>",
          "rawMarkdown": "lol I think you mean 0.787 not 0.887 @manikanthr5 . But even 0.787 seems hard to achieve, for my case using LGBM, but a bit easier using SAINT.",
          "votes": 1
        },
        {
          "id": 1129036,
          "postDate": "2020-12-28T02:56:33.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdessalemboukil\" target=\"_blank\">@abdessalemboukil</a> My single LGBM model achieves 0.806 CV score, submission is pending (it might go below 0.8)</p>",
          "rawMarkdown": "@abdessalemboukil My single LGBM model achieves 0.806 CV score, submission is pending (it might go below 0.8)"
        },
        {
          "id": 1129037,
          "postDate": "2020-12-28T02:57:23.900Z",
          "content": "<p>Yes, my bad. correcting it. Thanks for pointing out. </p>",
          "rawMarkdown": "Yes, my bad. correcting it. Thanks for pointing out. "
        },
        {
          "id": 1129921,
          "postDate": "2020-12-28T16:38:23.367Z",
          "content": "<p>In discussions was lb score 0.8+ with single lgbm</p>",
          "rawMarkdown": "In discussions was lb score 0.8+ with single lgbm"
        }
      ]
    },
    {
      "id": 1137494,
      "postDate": "2021-01-04T01:57:45.330Z",
      "content": "<p>I learned a lot from shared notebook. But I don't have interest to copy and resubmit it after slight change. I like any ideas shared and try to integrate it into my notebook. However, I feel many things are frustrating in this competition. The scoring error drove my crazy, and nobody can actually help!</p>",
      "rawMarkdown": "I learned a lot from shared notebook. But I don't have interest to copy and resubmit it after slight change. I like any ideas shared and try to integrate it into my notebook. However, I feel many things are frustrating in this competition. The scoring error drove my crazy, and nobody can actually help!",
      "votes": 1,
      "replies": [
        {
          "id": 1137522,
          "postDate": "2021-01-04T02:44:06.623Z",
          "content": "<p>There are some baseline notebooks with very low scores, but their submission abilities are totally ok. You may want to change the code (adding your code for features etc.), leaving only the last submission part untouched and figure out when your submission actually stops working. Most likely you are encountering timeout error or simply send the wrong submission file in the end (for example, you dont add the \"previous right answers\" element to the table or whatever it was called and etc.).</p>",
          "rawMarkdown": "There are some baseline notebooks with very low scores, but their submission abilities are totally ok. You may want to change the code (adding your code for features etc.), leaving only the last submission part untouched and figure out when your submission actually stops working. Most likely you are encountering timeout error or simply send the wrong submission file in the end (for example, you dont add the \"previous right answers\" element to the table or whatever it was called and etc.)."
        },
        {
          "id": 1137543,
          "postDate": "2021-01-04T03:04:59.467Z",
          "content": "<p>We're you able to solve your issue? If it is working fine before inference then most likely issue is with handling lectures. If it is then the error will be caused if you have removed lectures and are trying to add them to your label and user answer columns. I used <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">tito's api simulator notebook</a> to debug my issues. </p>\n<pre><code>previous_test_df = None\nfor (current_test, _) in iter_test:\n    if previous_test_df is not None:\n        actual_answer = np.array(eval(current_test[\"prior_group_answers_correct\"].iloc[0]))\n        previous_test_df['answered_correctly'] = actual_answer[actual_answer != -1]\n        user_answer = np.array(eval(current_test[\"prior_group_responses\"].iloc[0]))\n        previous_test_df['user_answer'] = user_answer[user_answer != -1]\n</code></pre>",
          "rawMarkdown": "We're you able to solve your issue? If it is working fine before inference then most likely issue is with handling lectures. If it is then the error will be caused if you have removed lectures and are trying to add them to your label and user answer columns. I used [tito's api simulator notebook](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) to debug my issues. \n\n```\nprevious_test_df = None\nfor (current_test, _) in iter_test:\n    if previous_test_df is not None:\n        actual_answer = np.array(eval(current_test[\"prior_group_answers_correct\"].iloc[0]))\n        previous_test_df['answered_correctly'] = actual_answer[actual_answer != -1]\n        user_answer = np.array(eval(current_test[\"prior_group_responses\"].iloc[0]))\n        previous_test_df['user_answer'] = user_answer[user_answer != -1]\n```"
        },
        {
          "id": 1137634,
          "postDate": "2021-01-04T05:40:25.743Z",
          "content": "<p>I finally find out the reason. After I make the submission environment ( i.e., env = riiideducation.make_env(), iter_test = env.iter_test()), I cannot use my pipeline function. It simply creates submission error even thought the pipeline function should be fine. I just copied the contents in the pipeline function, then I can submit it without errors. But I wasted almost one day trying to figure out the reason. Very sad, because I didn't learn anything new. </p>",
          "rawMarkdown": "I finally find out the reason. After I make the submission environment ( i.e., env = riiideducation.make_env(), iter_test = env.iter_test()), I cannot use my pipeline function. It simply creates submission error even thought the pipeline function should be fine. I just copied the contents in the pipeline function, then I can submit it without errors. But I wasted almost one day trying to figure out the reason. Very sad, because I didn't learn anything new. "
        }
      ]
    },
    {
      "id": 1130013,
      "postDate": "2020-12-28T17:43:42.357Z",
      "content": "<p>I am way off in this competition but I do agree with this sentiment. the ranks from 300 to 800 are all on the same score and I feel that it's mostly whoever can find a high-scoring notebook and fork it to submit it. I thought there was some sort of system there to prevent serial submitters from actually winning the medals but apparently, there isn't one.  </p>",
      "rawMarkdown": "I am way off in this competition but I do agree with this sentiment. the ranks from 300 to 800 are all on the same score and I feel that it's mostly whoever can find a high-scoring notebook and fork it to submit it. I thought there was some sort of system there to prevent serial submitters from actually winning the medals but apparently, there isn't one.  ",
      "votes": 1
    },
    {
      "id": 1128826,
      "postDate": "2020-12-27T19:12:45.163Z",
      "content": "<p>Same here, three weeks ago I took a break of five days. I was the 126th in ranking, and after the break I lost 500 positions due to public kernels shared. I had to work hard these last weeks to recover from that.</p>",
      "rawMarkdown": "Same here, three weeks ago I took a break of five days. I was the 126th in ranking, and after the break I lost 500 positions due to public kernels shared. I had to work hard these last weeks to recover from that.",
      "votes": 1
    },
    {
      "id": 1128707,
      "postDate": "2020-12-27T17:05:45.040Z",
      "content": "<p>I also agree, and the same happened to me last week when the 0.782 notebook was released which was very frustrating.</p>\n<p>If we look at the global leaderboard, we can infer that the global distribution (without the public kernels) shall be centered around 0.752, as the two picks we see later (around 0.773 and 0.781-2-3) correspond to the public kernels.</p>\n<p>Figure updated today:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcc97a7eae7d48a3fbb36347528d9f659%2FSans%20titre.png?generation=1609088616735233&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I also agree, and the same happened to me last week when the 0.782 notebook was released which was very frustrating.\n\nIf we look at the global leaderboard, we can infer that the global distribution (without the public kernels) shall be centered around 0.752, as the two picks we see later (around 0.773 and 0.781-2-3) correspond to the public kernels.\n\nFigure updated today:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcc97a7eae7d48a3fbb36347528d9f659%2FSans%20titre.png?generation=1609088616735233&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1128716,
          "postDate": "2020-12-27T17:12:27.413Z",
          "content": "<p>Yeah, 0.75+ is achievable with simple lgbm or sakt. That should be a baseline for beginners who are working on their own models. </p>\n<p>Now 0.781 is at peak but tomorrow it's gonna shift to 0.783. I hope no one publishes notebook with more than 0.785+. </p>",
          "rawMarkdown": "Yeah, 0.75+ is achievable with simple lgbm or sakt. That should be a baseline for beginners who are working on their own models. \n\nNow 0.781 is at peak but tomorrow it's gonna shift to 0.783. I hope no one publishes notebook with more than 0.785+. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1137349,
      "postDate": "2021-01-03T21:19:06.987Z",
      "content": "<p>Guys, plz stop doing these complaints for each competition. No one knows which notebook will score better in the private LB. It is a share and learn platform.  If u find sth amazing to show off (even if it is a bit better model than the original), then just go for it and share with the community. U should not relate to what others do since in this specific topic everyone has a different opinion. Thus, let everyone act as the Kaggle regulates. Cheers.</p>",
      "rawMarkdown": "Guys, plz stop doing these complaints for each competition. No one knows which notebook will score better in the private LB. It is a share and learn platform.  If u find sth amazing to show off (even if it is a bit better model than the original), then just go for it and share with the community. U should not relate to what others do since in this specific topic everyone has a different opinion. Thus, let everyone act as the Kaggle regulates. Cheers.",
      "votes": -1,
      "replies": [
        {
          "id": 1137550,
          "postDate": "2021-01-04T03:16:29.080Z",
          "content": "<p>I understand what you are saying. I have learnt a lot in the last few months by using public notebooks. All my public notebooks are slight modifications of popular public notebooks. I am one of those who would love to see the community grow together. But… You should also try to understand the situation. Try to look at the number of forks (each 940+ and 1000+) of the top scoring notebooks and also the people with scores of 781/0.783 with &lt;= 2 submissions. This is just utter madness. These are good contributions with modifications of popular notebooks (0.783 one uses my SAKT notebook output). I am happy that people are contributing but what I don't want is how people are taking it at the time. They are just submitting someone else's work without any additions of their own. Even if we don't know that they may not score as good in private but the opposite might also be possible (look for the top MOA public notebook). During the end of the competitions this is detrimental to the community as a whole. Those who loose medals due to the public forking will loose the spirit to share their ideas with the community. I myself am reconsidering whether should I share my SAINT+ training kernel without submission part. </p>",
          "rawMarkdown": "I understand what you are saying. I have learnt a lot in the last few months by using public notebooks. All my public notebooks are slight modifications of popular public notebooks. I am one of those who would love to see the community grow together. But... You should also try to understand the situation. Try to look at the number of forks (each 940+ and 1000+) of the top scoring notebooks and also the people with scores of 781/0.783 with <= 2 submissions. This is just utter madness. These are good contributions with modifications of popular notebooks (0.783 one uses my SAKT notebook output). I am happy that people are contributing but what I don't want is how people are taking it at the time. They are just submitting someone else's work without any additions of their own. Even if we don't know that they may not score as good in private but the opposite might also be possible (look for the top MOA public notebook). During the end of the competitions this is detrimental to the community as a whole. Those who loose medals due to the public forking will loose the spirit to share their ideas with the community. I myself am reconsidering whether should I share my SAINT+ training kernel without submission part. "
        },
        {
          "id": 1138236,
          "postDate": "2021-01-04T14:14:39.607Z",
          "content": "<p>I do understand ur statements, however, there is nothing that should be done to prevent that (this is my opinion). Coz, for almost 99 percent of all competitions, there is no one in the top 200 or 300 who has his notebook made public. What I mean is, as a result of these public notebooks, competitiveness grows even more as people see new ways to improve the LB score. Anyways, for me, Kaggle and their notebooks are already regulated to the positive of the community and growth of the competitiveness among competitors. Cheers.</p>",
          "rawMarkdown": "I do understand ur statements, however, there is nothing that should be done to prevent that (this is my opinion). Coz, for almost 99 percent of all competitions, there is no one in the top 200 or 300 who has his notebook made public. What I mean is, as a result of these public notebooks, competitiveness grows even more as people see new ways to improve the LB score. Anyways, for me, Kaggle and their notebooks are already regulated to the positive of the community and growth of the competitiveness among competitors. Cheers.",
          "votes": -1
        },
        {
          "id": 1138638,
          "postDate": "2021-01-04T20:30:07.150Z",
          "content": "<p><a href=\"https://www.kaggle.com/elvinagammed\" target=\"_blank\">@elvinagammed</a> </p>\n<blockquote>\n  <p>It is a share and learn platform</p>\n</blockquote>\n<p>not only, it's also a competition platform. And there is plenty of time to share after the competition is over. But apparently many people are less interested in learning after the competition is over…</p>\n<blockquote>\n  <p>as a result of these public notebooks, competitiveness grows even more as people see new ways to improve the LB score</p>\n</blockquote>\n<p>I very much disagree with this statement. High scoring kernels distract people from trying their own code and this reduces the diversity and the originality of models. It won't distract the top kagglers but it does for many competitors who fight for bronze/silver medals.</p>\n<blockquote>\n  <p>plz stop doing these complaints for each competition</p>\n</blockquote>\n<p>Kaggle is a community and in a community opinions are discussed and rules are changed based on discussions. In the past many rules have been added or adjusted based on discussions in the forum. Inviting people to \"stop\" discussing about topics is against the spirit of this community.</p>",
          "rawMarkdown": "@elvinagammed \n\n> It is a share and learn platform\n\nnot only, it's also a competition platform. And there is plenty of time to share after the competition is over. But apparently many people are less interested in learning after the competition is over...\n\n > as a result of these public notebooks, competitiveness grows even more as people see new ways to improve the LB score\n\nI very much disagree with this statement. High scoring kernels distract people from trying their own code and this reduces the diversity and the originality of models. It won't distract the top kagglers but it does for many competitors who fight for bronze/silver medals.\n\n\n> plz stop doing these complaints for each competition\n\nKaggle is a community and in a community opinions are discussed and rules are changed based on discussions. In the past many rules have been added or adjusted based on discussions in the forum. Inviting people to \"stop\" discussing about topics is against the spirit of this community.",
          "votes": 3
        },
        {
          "id": 1138651,
          "postDate": "2021-01-04T20:57:41.443Z",
          "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> thanks for ur valuable opinions. <br>\nAbout the first part, I see it is a competition platform but have u seen anyone who has a public notebook being in the prize ranking with a fork of sb else? No, it is not possible as competitiveness is much higher than the existence of those forks. (My opinions should not sound like I support those fork, submit notebooks…)<br>\nand lastly, there is a difference between a discussion and a complaint. These topics have been discussed several times already, a conclusion is made to preserve the state since probably Kaggle team thinks this contributes to the community more as it supports newbies, and doesn't affect grandmasters badly at the same time. </p>\n<p>In any case, ur opinions are respected. Cheers)</p>",
          "rawMarkdown": "@stecasasso thanks for ur valuable opinions. \nAbout the first part, I see it is a competition platform but have u seen anyone who has a public notebook being in the prize ranking with a fork of sb else? No, it is not possible as competitiveness is much higher than the existence of those forks. (My opinions should not sound like I support those fork, submit notebooks...)\nand lastly, there is a difference between a discussion and a complaint. These topics have been discussed several times already, a conclusion is made to preserve the state since probably Kaggle team thinks this contributes to the community more as it supports newbies, and doesn't affect grandmasters badly at the same time. \n\nIn any case, ur opinions are respected. Cheers)",
          "votes": -1
        }
      ]
    },
    {
      "id": 1128684,
      "postDate": "2020-12-27T16:57:16.850Z",
      "content": "<p>Very agree with you. Publishing high scoring notebooks frustrate people who make a lot of efforts in the competition.<br>\nHope we can have good competition till the end.</p>",
      "rawMarkdown": "Very agree with you. Publishing high scoring notebooks frustrate people who make a lot of efforts in the competition.\nHope we can have good competition till the end.\n",
      "votes": 2,
      "replies": [
        {
          "id": 1128700,
          "postDate": "2020-12-27T17:02:21.683Z",
          "content": "<p>I do hope that this becomes a good competition for everyone who put effort into it. But it's scary to look at the 0.781 and 0.783 scores in the leaderboard, especially the increase in the number of people with the scores who have only 1/2 submissions , in the last 2 days 😣. </p>",
          "rawMarkdown": "I do hope that this becomes a good competition for everyone who put effort into it. But it's scary to look at the 0.781 and 0.783 scores in the leaderboard, especially the increase in the number of people with the scores who have only 1/2 submissions , in the last 2 days 😣. "
        }
      ]
    },
    {
      "id": 1136278,
      "postDate": "2021-01-02T23:05:36.633Z",
      "content": "<p>The worst is when one forks a high-scoring notebook and adds no effort of his own and claim the work is his. I do hope that the Judges of the competition will check for plagiarism - the first published/submitted work is the original copy. </p>",
      "rawMarkdown": "The worst is when one forks a high-scoring notebook and adds no effort of his own and claim the work is his. I do hope that the Judges of the competition will check for plagiarism - the first published/submitted work is the original copy. "
    },
    {
      "id": 1131337,
      "postDate": "2020-12-29T17:03:42.110Z",
      "content": "<blockquote>\n  <p><em>This has become rat race to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else</em></p>\n</blockquote>\n<p>Implying that current 0.770+ tier isnt a rat race for getting 0.001 more from tuning your hyperparameters, adding 10 more features or adding SAKT/ whatever to get the mentioned above 0.001 increase.</p>",
      "rawMarkdown": "> *This has become rat race to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else*\n\nImplying that current 0.770+ tier isnt a rat race for getting 0.001 more from tuning your hyperparameters, adding 10 more features or adding SAKT/ whatever to get the mentioned above 0.001 increase.",
      "replies": [
        {
          "id": 1131350,
          "postDate": "2020-12-29T17:10:49.920Z",
          "content": "<p>It depends on your perspective. For me, submitting someone elses notebook without even adding your contribution is an endless, self-defeating, or pointless pursuit. Tuning hyper-parameters by self or adding features that add value is one of the skills we have to develop as data scientists. </p>",
          "rawMarkdown": "It depends on your perspective. For me, submitting someone elses notebook without even adding your contribution is an endless, self-defeating, or pointless pursuit. Tuning hyper-parameters by self or adding features that add value is one of the skills we have to develop as data scientists. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1129721,
      "postDate": "2020-12-28T14:09:45.257Z",
      "content": "<p>Totally agree, I will be mad if I saw anyone public a notebook with a score &gt;= 0.79</p>",
      "rawMarkdown": "Totally agree, I will be mad if I saw anyone public a notebook with a score >= 0.79"
    },
    {
      "id": 1128945,
      "postDate": "2020-12-27T22:58:35.533Z",
      "content": "<p>I have not seen any Notebook which had a medal-winning score. I forked and submitted a Notebook with 0.773 LB score as well, just to figure out the submission system.</p>",
      "rawMarkdown": "I have not seen any Notebook which had a medal-winning score. I forked and submitted a Notebook with 0.773 LB score as well, just to figure out the submission system.",
      "replies": [
        {
          "id": 1129531,
          "postDate": "2020-12-28T12:17:06.010Z",
          "content": "<p>It changes over time but 0.783 is in bronze range now. When I made 0.773 SAKT model public, it was not in bronze range. </p>",
          "rawMarkdown": "It changes over time but 0.783 is in bronze range now. When I made 0.773 SAKT model public, it was not in bronze range. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1128740,
      "postDate": "2020-12-27T17:27:01.320Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1128786,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-12-27T18:31:54.020000",
      "content": "<p>Might be an unpopular opinion (hot take?), but this line of reasoning confuses me.</p>\n<p>In the past, there was no requirements whatsoever. This makes sense from Kaggle's perspective because their goal is to have the best result for their clients, and the hosts too want the best result for the amount of time (money) they have spent retaining Kaggle.</p>\n<p>Then eventually through community discussion, Kaggle added requirement to disclose non-competition specific details such as external datasets within 7 days before competition end. The driving force here was also in-line with Kaggle's overarching objective, of increasing the overall quality of submissions at the end of competition for their client.</p>\n<p>Again through more maturing communal discourse, Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients (not us competing data scientists) so no wonder it was a \"soft\" suggestion and no one gets banned for breaking it.</p>\n<p>We aren't at t-7 days to competition close yet. If someone shares something amazing, either on the forums or by means of a kernel, what is the problem? In MISH people were complaining 20 days before competition end and now here people are complaining 10 days before competition end. Kaggle already has the warning for t-7 days before competition end, do we need more regulations than that? If so, discuss on the global kaggle suggestions forums rather than competition specific forums, since it's the same conversation that keeps coming up month after month. I too was crushed by the public MISH kernels and God knows how many hours I've pushed into riiid (and spousal fights as a result of it). But if your goal is truly to learn, then learn from wherever the knowledge comes from. And if that knowledge comes t-7 days before the competition end date, as was agreed upon by the community guidelines, then be grateful that someone shared something useful and see how best you can integrate it into your solution. If the ground breaking happens after that then sure that sucks. But trying to restrict the flow of ideas -10 or -20 days before… that makes no sense to me. Why not just restrict sharing entirely \\s</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1128790,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-27T18:35:22.927000",
          "content": "<p>Btw the people who have discovered something crazy and share for discussion / notebook reputation points, aren't these the same people who would likely private share anyway? If they're gonna do that, better at least to have it done public rather than just benefit a small in-circle group of people. Sorry just my 2cents.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128800,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-27T18:51:34.183000",
          "content": "<p>Great points! And yes, if someone s here for learning, then this competition is a great opportunity! It just can't get better than this. Cool LB/CV sync, lot of cool notebooks/ideas/threads/papers/self ideas etc to try and see how it works and then enjoy being on the top-k% of the LB. Till here it's good but the trouble comes when people who just fork not with the intention to learn anything about what the author did etc but more with the intention to later share and be proud of their fake medal, they harm everyone, including the one's who have actually put into efforts, tremendous efforts rather. The analogy is similar to merits/demerits of using Internet let's say, how you use it, tells a lot about you as a person.</p>\n<p>And in case someone has shared something cool, look into it with a microscope and see how fast you can integrate and test the same pretty much! If it helps, keep it, if it doesn't, then you already have what you need!</p>\n<p>It's very important for people to communicate their ideas, you never know how that idea can be bended and take a new form! And avoid private sharing, there's no point. So if restricted, then Idk what will happen and that will give rise to more private sharing for sure.</p>\n<p>Speaking for myself, have learnt a ton of things from this comp, writing cool code, tests, quick and easily manageable pipelines, managing such a vol of data and many other things!</p>\n<p>Also, don't focus too much on LB scores, focus on your CV as public notebooks might bring you down in private LB.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1128857,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-27T20:03:10.610000",
          "content": "<p>You argument doesn't make sense. Learning is important, but the motivation to win the competition is also an important factor of Kaggle.</p>\n<p>Following your argument, why not ask Kaggle to cancel the competition, but only running as an learning platform? Do you think there will be the same amount of great people coming here to contribute?</p>\n<p>And about learning, where is the learning when people just fork and submit? Or even fork someone's kernel and publish it as he/she is the original author?</p>\n<p>This is not a 0 or 1 situation. We want people to win, but also want to respect and evaluate people's hard work! Some kind of rules should be maintained anyway.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1128865,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-27T20:13:06.197000",
          "content": "<p>I am not disagree with you. All I am saying is that the very gray and arbitrary \"too late to share\" cutoff which is unofficially renegotiated at every competition by participants should be done at the global kaggle community level (as been done in the past) if people really want to have hard deadlines. There already exist a cutoff period that the community and Kaggle have agreed to, which is t-7; so why should people be upset if someone shares something good at t-8? or t-10? The people who have worked hard like you and others, can take whatever is shared and bolster your own solutions further propelling you forward. The people who just fork and haven't put in any work, there will be a limit to how far they can go—both on this platform as well as in life..</p>\n<p>And just statistically speaking, the number of people in the top ranks for any given competition is extremely limited. I know well the desire to win (anyone who's competed for years is right there with you, we're all addicted). But if the goal in competing at Kaggle is <em>just</em> to win, or even <em>primarily</em> to win, then I daresay 99.9% of the people on this platform myself included are failing horribly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128871,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-27T20:26:04.987000",
          "content": "<p>Well, for myself,  even if I don't win in a competition, I still hope the results reflect better the reality. For example, say I am near the bronze zone (say rank 300). And someone publish a high score kernel in a silver zone? And his pipeline is totally different from mine (or even, the framework I never used before)?<br>\nIn this case, either I fork blindly, or I play honest and get a low rank, say 1000?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1128920,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2020-12-27T22:10:55.613000",
          "content": "<blockquote>\n  <p>Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients</p>\n</blockquote>\n<p>I strongly disagree with this assumption. </p>\n<p>What makes the value out of the competitions (for clients, users, everyone) it's to have thousands of minds working on the same problem each one from a different angle, contributing to original, creative, innovative approaches.    <br>\nThe spreading of high scoring kernels only attracted users who try to find a shortcut and hope they can win a medal by slightly tweaking some parameters of an existing script. This is <strong>not</strong> of any value for anyone, I am sorry. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1128961,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-27T23:57:31.227000",
          "content": "<p>That's nice and great to hold as conjecture, but the point you raise has never been realized as policy, neither soft or otherwise. Whereas all the points I raised were <strong>historical</strong>.</p>\n<p>So either the people in charge do in fact feel there is benefit to doing it the way they've been doing it; or perhaps, as I've repeated now for the third time, that the proper place to bring up policy changes across Kaggle as a whole would be the suggestions / feedback forums, which have been purposely built for exactly this type of discourse.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fc70667212f0f75e6b17b7e4385a02327%2FScreen%20Shot%202020-12-27%20at%207.16.50%20PM.png?generation=1609118245701021&amp;alt=media\" alt=\"\"></p>\n<p>Last thing I'll mention about this. While we did have a \"high scoring\" - if we can call it that - kernel published 2 days ago, the reality of the matter is another kernel with the exact same score was available 10 days ago. And we can see sorting by best public kernel that there really hasn't been any action in the last +14 day. In other words, almost no real sharing for quite a while. If people would be more open, we'd see a lot more interesting solutions imo, especially since there def is some good FE going on as evidenced by the public LB.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1129021,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-28T02:14:10.963000",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> you have raised great points. I think what we are discussing is in the grey area. If a kernel is published now on something that is genuinely good work and is helpful to everyone, then everyone would be happy and take it in a positive way. All I am concerned about is about the fork of fork of forks. If you look at the the number of forks vs upvotes, it's too disproportionate. I still think that it depends on the intention of the author. They have to weight the positives vs negatives before publishing one. All I want from this post is to remind them of that. If it's something that crushes leaderboard in the last few days, it can wait till the end of the competition. No one complains after that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1129679,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2020-12-28T14:03:02.800000",
          "content": "<blockquote>\n  <p>that the proper place to bring up policy changes across Kaggle as a whole would be the suggestions / feedback forums</p>\n</blockquote>\n<p>And so what? Why can't we have this discussion here too? If good points are raised, are they illegitimate just because they are not posted in the optimal channel?</p>\n<blockquote>\n  <p>That's nice and great to hold as conjecture, but the point you raise has never been realized as policy, neither soft or otherwise</p>\n</blockquote>\n<p>3 people replied to your comment, who and what are you referring to? </p>\n<blockquote>\n  <p>And we can see sorting by best public kernel that there really hasn't been any action in the last +14 day. In other words, almost no real sharing for quite a while</p>\n</blockquote>\n<p>If you read the post carefully, there is an invitation to stop publishing high scoring kernels from now on. Which, incidentally, is upvoted by 44 people already (as of writing).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1129734,
      "author_name": "Jacky",
      "author_url": "",
      "post_date": "2020-12-28T14:15:08.690000",
      "content": "<p>I think the main issue is not the public kernels with high scores pubished last minute but more the people that blindly fork and resubmit. I always like to pick up ideas from top kernels, even if it is at the last minute if someone completly outscore my own submission.</p>\n<p>What I don't like is seeing 500 persons climbing the ladder with absolutly no work/no merite and just hunting for the medal ☹️</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1129756,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-28T14:29:56.343000",
          "content": "<p>full agree ,what can learn from the blindly fork notebook ? turing hyperparameter?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129774,
          "author_name": "Jacky",
          "author_url": "",
          "post_date": "2020-12-28T14:40:24.903000",
          "content": "<p>Yeah… <br>\nOne way of possibly going would be to disable fork of notebooks from a certain moment in the competition and having no possibility to c/c cells… </p>\n<p>It could be still possible to see solutions of others, but it would requiert to actually recode it to include it in your own solution. </p>\n<p>Still possible to recode a full kernel, but I'm sure it would discourage most of the serial resubmiters 😄</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1128725,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-12-27T17:18:26.443000",
      "content": "<p>Totally agree. I was thinking of writing a similar post. If we look at public notebooks sorted by hotness only few of them are quite interesting but the others are fork of forks.. I think there is a still an important gap between 0.78x and 0.79x but I'm pretty afraid of waking-up a morning and seeing a public kernel scoring 0.79x. <br>\nHowever I remain optimistic that people who worked hard will be rewarded </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1128732,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-27T17:21:28.940000",
          "content": "<p>Yeah, this is what I am most worried about. 0.79x is not achievable without hard work. I am trying my best but still in the 0.77x. I haven't finished my LGBM and SAINT+. I want to hope that no one releases their saint model now. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1128742,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-12-27T17:28:00.840000",
          "content": "<p>Releasing a high scoring SAINT in the last 10 days of competition would probably destroy the LB</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1128746,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2020-12-27T17:34:29.273000",
      "content": "<p>What is most surprising to me it's the little Kaggle has done to prevent this from happening… Given it's an outstanding issue since years…</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1128749,
      "author_name": "qiaqia",
      "author_url": "",
      "post_date": "2020-12-27T17:37:31.763000",
      "content": "<p>It was lucky to participate in this competition for me ,learned a lot.I'm trying my SANIT model though time is not enough.<br>\nMy best time so far is 0.782, and it looks like I'll be dropping crazy places again tomorrow。</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1128760,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-27T17:52:35.267000",
          "content": "<p>Don't give up. I think single SAINT could beat LightGBM and SAKT ensemble hands down. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128869,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-27T20:25:10.497000",
          "content": "<p>of course, I think I can do better with single lgbm ,0.782 not end . and try SAINT then ensemble .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1129023,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-28T02:15:50.390000",
          "content": "<p>Yeah. It seems atleast  <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801#1116607\" target=\"_blank\">0.786</a> is achievable using LightGBM with proper feature engineering. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129030,
          "author_name": "Abdessalem Boukil",
          "author_url": "",
          "post_date": "2020-12-28T02:37:48.697000",
          "content": "<p>lol I think you mean 0.787 not 0.887 <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a> . But even 0.787 seems hard to achieve, for my case using LGBM, but a bit easier using SAINT.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1129036,
          "author_name": "Bekir",
          "author_url": "",
          "post_date": "2020-12-28T02:56:33.137000",
          "content": "<p><a href=\"https://www.kaggle.com/abdessalemboukil\" target=\"_blank\">@abdessalemboukil</a> My single LGBM model achieves 0.806 CV score, submission is pending (it might go below 0.8)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129037,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-28T02:57:23.900000",
          "content": "<p>Yes, my bad. correcting it. Thanks for pointing out. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129921,
          "author_name": "Alyona Pasevieva",
          "author_url": "",
          "post_date": "2020-12-28T16:38:23.367000",
          "content": "<p>In discussions was lb score 0.8+ with single lgbm</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1137494,
      "author_name": "Noah Xi",
      "author_url": "",
      "post_date": "2021-01-04T01:57:45.330000",
      "content": "<p>I learned a lot from shared notebook. But I don't have interest to copy and resubmit it after slight change. I like any ideas shared and try to integrate it into my notebook. However, I feel many things are frustrating in this competition. The scoring error drove my crazy, and nobody can actually help!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1137522,
          "author_name": "majoraregalia",
          "author_url": "",
          "post_date": "2021-01-04T02:44:06.623000",
          "content": "<p>There are some baseline notebooks with very low scores, but their submission abilities are totally ok. You may want to change the code (adding your code for features etc.), leaving only the last submission part untouched and figure out when your submission actually stops working. Most likely you are encountering timeout error or simply send the wrong submission file in the end (for example, you dont add the \"previous right answers\" element to the table or whatever it was called and etc.).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1137543,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-04T03:04:59.467000",
          "content": "<p>We're you able to solve your issue? If it is working fine before inference then most likely issue is with handling lectures. If it is then the error will be caused if you have removed lectures and are trying to add them to your label and user answer columns. I used <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">tito's api simulator notebook</a> to debug my issues. </p>\n<pre><code>previous_test_df = None\nfor (current_test, _) in iter_test:\n    if previous_test_df is not None:\n        actual_answer = np.array(eval(current_test[\"prior_group_answers_correct\"].iloc[0]))\n        previous_test_df['answered_correctly'] = actual_answer[actual_answer != -1]\n        user_answer = np.array(eval(current_test[\"prior_group_responses\"].iloc[0]))\n        previous_test_df['user_answer'] = user_answer[user_answer != -1]\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1137634,
          "author_name": "Noah Xi",
          "author_url": "",
          "post_date": "2021-01-04T05:40:25.743000",
          "content": "<p>I finally find out the reason. After I make the submission environment ( i.e., env = riiideducation.make_env(), iter_test = env.iter_test()), I cannot use my pipeline function. It simply creates submission error even thought the pipeline function should be fine. I just copied the contents in the pipeline function, then I can submit it without errors. But I wasted almost one day trying to figure out the reason. Very sad, because I didn't learn anything new. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1130013,
      "author_name": "Sumit Singh",
      "author_url": "",
      "post_date": "2020-12-28T17:43:42.357000",
      "content": "<p>I am way off in this competition but I do agree with this sentiment. the ranks from 300 to 800 are all on the same score and I feel that it's mostly whoever can find a high-scoring notebook and fork it to submit it. I thought there was some sort of system there to prevent serial submitters from actually winning the medals but apparently, there isn't one.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1128826,
      "author_name": "Abdessalem Boukil",
      "author_url": "",
      "post_date": "2020-12-27T19:12:45.163000",
      "content": "<p>Same here, three weeks ago I took a break of five days. I was the 126th in ranking, and after the break I lost 500 positions due to public kernels shared. I had to work hard these last weeks to recover from that.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1128707,
      "author_name": "Jacky",
      "author_url": "",
      "post_date": "2020-12-27T17:05:45.040000",
      "content": "<p>I also agree, and the same happened to me last week when the 0.782 notebook was released which was very frustrating.</p>\n<p>If we look at the global leaderboard, we can infer that the global distribution (without the public kernels) shall be centered around 0.752, as the two picks we see later (around 0.773 and 0.781-2-3) correspond to the public kernels.</p>\n<p>Figure updated today:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcc97a7eae7d48a3fbb36347528d9f659%2FSans%20titre.png?generation=1609088616735233&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1128716,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-27T17:12:27.413000",
          "content": "<p>Yeah, 0.75+ is achievable with simple lgbm or sakt. That should be a baseline for beginners who are working on their own models. </p>\n<p>Now 0.781 is at peak but tomorrow it's gonna shift to 0.783. I hope no one publishes notebook with more than 0.785+. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1137349,
      "author_name": "Elvin Aghammadzada",
      "author_url": "",
      "post_date": "2021-01-03T21:19:06.987000",
      "content": "<p>Guys, plz stop doing these complaints for each competition. No one knows which notebook will score better in the private LB. It is a share and learn platform.  If u find sth amazing to show off (even if it is a bit better model than the original), then just go for it and share with the community. U should not relate to what others do since in this specific topic everyone has a different opinion. Thus, let everyone act as the Kaggle regulates. Cheers.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 1137550,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-04T03:16:29.080000",
          "content": "<p>I understand what you are saying. I have learnt a lot in the last few months by using public notebooks. All my public notebooks are slight modifications of popular public notebooks. I am one of those who would love to see the community grow together. But… You should also try to understand the situation. Try to look at the number of forks (each 940+ and 1000+) of the top scoring notebooks and also the people with scores of 781/0.783 with &lt;= 2 submissions. This is just utter madness. These are good contributions with modifications of popular notebooks (0.783 one uses my SAKT notebook output). I am happy that people are contributing but what I don't want is how people are taking it at the time. They are just submitting someone else's work without any additions of their own. Even if we don't know that they may not score as good in private but the opposite might also be possible (look for the top MOA public notebook). During the end of the competitions this is detrimental to the community as a whole. Those who loose medals due to the public forking will loose the spirit to share their ideas with the community. I myself am reconsidering whether should I share my SAINT+ training kernel without submission part. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1138236,
          "author_name": "Elvin Aghammadzada",
          "author_url": "",
          "post_date": "2021-01-04T14:14:39.607000",
          "content": "<p>I do understand ur statements, however, there is nothing that should be done to prevent that (this is my opinion). Coz, for almost 99 percent of all competitions, there is no one in the top 200 or 300 who has his notebook made public. What I mean is, as a result of these public notebooks, competitiveness grows even more as people see new ways to improve the LB score. Anyways, for me, Kaggle and their notebooks are already regulated to the positive of the community and growth of the competitiveness among competitors. Cheers.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1138638,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2021-01-04T20:30:07.150000",
          "content": "<p><a href=\"https://www.kaggle.com/elvinagammed\" target=\"_blank\">@elvinagammed</a> </p>\n<blockquote>\n  <p>It is a share and learn platform</p>\n</blockquote>\n<p>not only, it's also a competition platform. And there is plenty of time to share after the competition is over. But apparently many people are less interested in learning after the competition is over…</p>\n<blockquote>\n  <p>as a result of these public notebooks, competitiveness grows even more as people see new ways to improve the LB score</p>\n</blockquote>\n<p>I very much disagree with this statement. High scoring kernels distract people from trying their own code and this reduces the diversity and the originality of models. It won't distract the top kagglers but it does for many competitors who fight for bronze/silver medals.</p>\n<blockquote>\n  <p>plz stop doing these complaints for each competition</p>\n</blockquote>\n<p>Kaggle is a community and in a community opinions are discussed and rules are changed based on discussions. In the past many rules have been added or adjusted based on discussions in the forum. Inviting people to \"stop\" discussing about topics is against the spirit of this community.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1138651,
          "author_name": "Elvin Aghammadzada",
          "author_url": "",
          "post_date": "2021-01-04T20:57:41.443000",
          "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> thanks for ur valuable opinions. <br>\nAbout the first part, I see it is a competition platform but have u seen anyone who has a public notebook being in the prize ranking with a fork of sb else? No, it is not possible as competitiveness is much higher than the existence of those forks. (My opinions should not sound like I support those fork, submit notebooks…)<br>\nand lastly, there is a difference between a discussion and a complaint. These topics have been discussed several times already, a conclusion is made to preserve the state since probably Kaggle team thinks this contributes to the community more as it supports newbies, and doesn't affect grandmasters badly at the same time. </p>\n<p>In any case, ur opinions are respected. Cheers)</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1128684,
      "author_name": "Dean",
      "author_url": "",
      "post_date": "2020-12-27T16:57:16.850000",
      "content": "<p>Very agree with you. Publishing high scoring notebooks frustrate people who make a lot of efforts in the competition.<br>\nHope we can have good competition till the end.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1128700,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-27T17:02:21.683000",
          "content": "<p>I do hope that this becomes a good competition for everyone who put effort into it. But it's scary to look at the 0.781 and 0.783 scores in the leaderboard, especially the increase in the number of people with the scores who have only 1/2 submissions , in the last 2 days 😣. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1136278,
      "author_name": "Nawas Naziru",
      "author_url": "",
      "post_date": "2021-01-02T23:05:36.633000",
      "content": "<p>The worst is when one forks a high-scoring notebook and adds no effort of his own and claim the work is his. I do hope that the Judges of the competition will check for plagiarism - the first published/submitted work is the original copy. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1131337,
      "author_name": "majoraregalia",
      "author_url": "",
      "post_date": "2020-12-29T17:03:42.110000",
      "content": "<blockquote>\n  <p><em>This has become rat race to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else</em></p>\n</blockquote>\n<p>Implying that current 0.770+ tier isnt a rat race for getting 0.001 more from tuning your hyperparameters, adding 10 more features or adding SAKT/ whatever to get the mentioned above 0.001 increase.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1131350,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-29T17:10:49.920000",
          "content": "<p>It depends on your perspective. For me, submitting someone elses notebook without even adding your contribution is an endless, self-defeating, or pointless pursuit. Tuning hyper-parameters by self or adding features that add value is one of the skills we have to develop as data scientists. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1129721,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2020-12-28T14:09:45.257000",
      "content": "<p>Totally agree, I will be mad if I saw anyone public a notebook with a score &gt;= 0.79</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1128945,
      "author_name": "Bekir",
      "author_url": "",
      "post_date": "2020-12-27T22:58:35.533000",
      "content": "<p>I have not seen any Notebook which had a medal-winning score. I forked and submitted a Notebook with 0.773 LB score as well, just to figure out the submission system.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1129531,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2020-12-28T12:17:06.010000",
          "content": "<p>It changes over time but 0.783 is in bronze range now. When I made 0.773 SAKT model public, it was not in bronze range. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1128740,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-27T17:27:01.320000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1128648": "Hi all, I think it's about time we stop posting high scoring notebooks (one's that are easy to fork and submit without adding any contribution). We have about only 10 days left and I don't think it's fair anymore to publish any high scoring notebooks. \n- In LISH MoA competition, 1 high scoring public notebook released in the last few days ended up contributing to most of the silver and bronze medals. People who have worked hard got frustrated (check [this](https://www.kaggle.com/c/lish-moa/discussion/200586) and [this](https://www.kaggle.com/c/lish-moa/discussion/200535)). This competition is going through the same phase and might even become worse. (I myself lost 500+ positions last week since I was **on vacation**. I want to work on this one till the end but I do understand that others may not be able to work till the end. There are a lot of people who have worked on this from the very beginning and might be on vacation due to the year end).\n- All the latest high scoring notebooks are just **forks of forks of forks** (Just look at the notebook names) with funda of **\"Do hyper-parameter tuning and make the notebook public\"**. I think this is **not adding any value**. \n- **1200 forks** from just 2 notebooks (900, 300), this is scary stuff. **This has become rat race (endless, self-defeating, or pointless pursuit) to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else**. Those who are just forking and submitting public notebooks without actually working on this competition, you guys have my pity (I don't enjoy playing my video games from someone else's *final checkpoint*, I want to beat the final boss by myself). This competition is one of the few gems for learning feature engineering, understanding how lgbm is too different from xgboost and catboost in its tree building, how transformers work and the most important - building efficient pipelines with limited RAM and TIME. If you haven't learnt these....\n\n---\n\nFor those of you who are misunderstanding this post:\n- ~~Don't share your work~~ Don't share easy to fork notebooks in the last days of the competition. \n- People should understand that this is not an easy competition. A lot of people have been working hard for months and it's not fair for them to loose to someone who didn't work hard but ended up submitting someone else's work in the last days. \n- Kaggle is as much of competition site as that of a learning site for us. Please try to appreciate it. Learning can wait few days if people's hardwork is at stake. \n\n---\n\nSo please I request all of you who are considering making high scoring public notebooks, please wait till the end of the competition. Thanks 😃",
    "1128786": "Might be an unpopular opinion (hot take?), but this line of reasoning confuses me.\n\nIn the past, there was no requirements whatsoever. This makes sense from Kaggle's perspective because their goal is to have the best result for their clients, and the hosts too want the best result for the amount of time (money) they have spent retaining Kaggle.\n\nThen eventually through community discussion, Kaggle added requirement to disclose non-competition specific details such as external datasets within 7 days before competition end. The driving force here was also in-line with Kaggle's overarching objective, of increasing the overall quality of submissions at the end of competition for their client.\n\nAgain through more maturing communal discourse, Kaggle added the soft warning on Kernels page not to publish high scoring kernels within t-7 days before competition closure. This time, this acts against the interest of their paying clients (not us competing data scientists) so no wonder it was a \"soft\" suggestion and no one gets banned for breaking it.\n\nWe aren't at t-7 days to competition close yet. If someone shares something amazing, either on the forums or by means of a kernel, what is the problem? In MISH people were complaining 20 days before competition end and now here people are complaining 10 days before competition end. Kaggle already has the warning for t-7 days before competition end, do we need more regulations than that? If so, discuss on the global kaggle suggestions forums rather than competition specific forums, since it's the same conversation that keeps coming up month after month. I too was crushed by the public MISH kernels and God knows how many hours I've pushed into riiid (and spousal fights as a result of it). But if your goal is truly to learn, then learn from wherever the knowledge comes from. And if that knowledge comes t-7 days before the competition end date, as was agreed upon by the community guidelines, then be grateful that someone shared something useful and see how best you can integrate it into your solution. If the ground breaking happens after that then sure that sucks. But trying to restrict the flow of ideas -10 or -20 days before... that makes no sense to me. Why not just restrict sharing entirely \\s",
    "1129734": "I think the main issue is not the public kernels with high scores pubished last minute but more the people that blindly fork and resubmit. I always like to pick up ideas from top kernels, even if it is at the last minute if someone completly outscore my own submission.\n\nWhat I don't like is seeing 500 persons climbing the ladder with absolutly no work/no merite and just hunting for the medal ☹️",
    "1128725": "Totally agree. I was thinking of writing a similar post. If we look at public notebooks sorted by hotness only few of them are quite interesting but the others are fork of forks.. I think there is a still an important gap between 0.78x and 0.79x but I'm pretty afraid of waking-up a morning and seeing a public kernel scoring 0.79x. \nHowever I remain optimistic that people who worked hard will be rewarded ",
    "1128746": "What is most surprising to me it's the little Kaggle has done to prevent this from happening... Given it's an outstanding issue since years...",
    "1128749": "It was lucky to participate in this competition for me ,learned a lot.I'm trying my SANIT model though time is not enough.\nMy best time so far is 0.782, and it looks like I'll be dropping crazy places again tomorrow。",
    "1137494": "I learned a lot from shared notebook. But I don't have interest to copy and resubmit it after slight change. I like any ideas shared and try to integrate it into my notebook. However, I feel many things are frustrating in this competition. The scoring error drove my crazy, and nobody can actually help!",
    "1130013": "I am way off in this competition but I do agree with this sentiment. the ranks from 300 to 800 are all on the same score and I feel that it's mostly whoever can find a high-scoring notebook and fork it to submit it. I thought there was some sort of system there to prevent serial submitters from actually winning the medals but apparently, there isn't one.  ",
    "1128826": "Same here, three weeks ago I took a break of five days. I was the 126th in ranking, and after the break I lost 500 positions due to public kernels shared. I had to work hard these last weeks to recover from that.",
    "1128707": "I also agree, and the same happened to me last week when the 0.782 notebook was released which was very frustrating.\n\nIf we look at the global leaderboard, we can infer that the global distribution (without the public kernels) shall be centered around 0.752, as the two picks we see later (around 0.773 and 0.781-2-3) correspond to the public kernels.\n\nFigure updated today:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcc97a7eae7d48a3fbb36347528d9f659%2FSans%20titre.png?generation=1609088616735233&alt=media)",
    "1137349": "Guys, plz stop doing these complaints for each competition. No one knows which notebook will score better in the private LB. It is a share and learn platform.  If u find sth amazing to show off (even if it is a bit better model than the original), then just go for it and share with the community. U should not relate to what others do since in this specific topic everyone has a different opinion. Thus, let everyone act as the Kaggle regulates. Cheers.",
    "1128684": "Very agree with you. Publishing high scoring notebooks frustrate people who make a lot of efforts in the competition.\nHope we can have good competition till the end.\n",
    "1136278": "The worst is when one forks a high-scoring notebook and adds no effort of his own and claim the work is his. I do hope that the Judges of the competition will check for plagiarism - the first published/submitted work is the original copy. ",
    "1131337": "> *This has become rat race to continuously check the notebook section for high scoring notebooks and to fork and submit it before anyone else*\n\nImplying that current 0.770+ tier isnt a rat race for getting 0.001 more from tuning your hyperparameters, adding 10 more features or adding SAKT/ whatever to get the mentioned above 0.001 increase.",
    "1129721": "Totally agree, I will be mad if I saw anyone public a notebook with a score >= 0.79",
    "1128945": "I have not seen any Notebook which had a medal-winning score. I forked and submitted a Notebook with 0.773 LB score as well, just to figure out the submission system.",
    "1128740": ""
  }
}