{
  "id": 123980,
  "title": "High Scoring Public Kernels",
  "url": "/competitions/bengaliai-cv19/discussion/123980",
  "author_name": "",
  "post_date": "2020-01-01T03:40:21.746828300Z",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Can we stop making high-scoring \"inference\" kernels public on the leaderboard? \nIt just feels unfair that some people submit other peoples model weights and place higher than people who are trying their best. It's quite depressing. \nCan we work around this anyhow? </p>\n\n<p>No offense to anyone.\nedited : I meant inference kernels, not starters. </p>",
  "messages": [
    {
      "id": "707554",
      "postDate": "01/01/2020 03:40:21",
      "content": "<p>Can we stop making high-scoring \"inference\" kernels public on the leaderboard? \nIt just feels unfair that some people submit other peoples model weights and place higher than people who are trying their best. It's quite depressing. \nCan we work around this anyhow? </p>\n\n<p>No offense to anyone.\nedited : I meant inference kernels, not starters. </p>",
      "rawMarkdown": "Can we stop making high-scoring \"inference\" kernels public on the leaderboard? \nIt just feels unfair that some people submit other peoples model weights and place higher than people who are trying their best. It's quite depressing. \nCan we work around this anyhow? \n\nNo offense to anyone.\nedited : I meant inference kernels, not starters.",
      "votes": null
    },
    {
      "id": "707659",
      "postDate": "01/01/2020 08:56:32",
      "content": "<p>Hi Rafid,\nYou should not be worrying about high scoring kernels being made public, since they can be overfitting on the public test set. If you look at other competitions, teams can have upto an 80 rank jump from public to private leaderboard. Whatever hypothesis you deem correct is something you should pursue and if it is better than the public kernels, it'll just need one small jump to reach the top! :D\nCheers!</p>",
      "rawMarkdown": "Hi Rafid,\nYou should not be worrying about high scoring kernels being made public, since they can be overfitting on the public test set. If you look at other competitions, teams can have upto an 80 rank jump from public to private leaderboard. Whatever hypothesis you deem correct is something you should pursue and if it is better than the public kernels, it'll just need one small jump to reach the top! :D\nCheers!",
      "votes": null
    },
    {
      "id": "707666",
      "postDate": "01/01/2020 09:13:47",
      "content": "<p>hi Rafid,\nIn a recent competition I have seen triple(close to 1000) digit jumps because of overfitting. My submission jumped 400 ranks in final private test. Always trust your validation techniques. Secondly, I think its good that high ranking starter kernels are released as these help us understand stuff that we may be missing in our pipeline. You can pick parts of the kernel and test on your pipeline which is typically what we do in our jobs also, take a new feature engineering technique, a new augmentation, new loss and experiment for results.</p>",
      "rawMarkdown": "hi Rafid,\nIn a recent competition I have seen triple(close to 1000) digit jumps because of overfitting. My submission jumped 400 ranks in final private test. Always trust your validation techniques. Secondly, I think its good that high ranking starter kernels are released as these help us understand stuff that we may be missing in our pipeline. You can pick parts of the kernel and test on your pipeline which is typically what we do in our jobs also, take a new feature engineering technique, a new augmentation, new loss and experiment for results.",
      "votes": null
    },
    {
      "id": "707698",
      "postDate": "01/01/2020 10:17:55",
      "content": "<p>Hi, I wrongly addressed them as starter kernels. I meant to indicate inference kernels. I mean starter kernels are great for learning stuff,  but inference kernels do not fall in the criteria. \nLook at the leaderboard right now, you'll see what I mean.</p>",
      "rawMarkdown": "Hi, I wrongly addressed them as starter kernels. I meant to indicate inference kernels. I mean starter kernels are great for learning stuff,  but inference kernels do not fall in the criteria. \nLook at the leaderboard right now, you'll see what I mean.",
      "votes": null
    },
    {
      "id": "707700",
      "postDate": "01/01/2020 10:20:19",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "707756",
      "postDate": "01/01/2020 12:32:04",
      "content": "<p>Leader-board shake is one of the most interesting things on Kaggle.\nFor example: <a href=\"https://www.kaggle.com/robikscube/ashrae-leaderboard-and-shake\">https://www.kaggle.com/robikscube/ashrae-leaderboard-and-shake</a></p>",
      "rawMarkdown": "Leader-board shake is one of the most interesting things on Kaggle.\nFor example: https://www.kaggle.com/robikscube/ashrae-leaderboard-and-shake",
      "votes": null
    },
    {
      "id": "707768",
      "postDate": "01/01/2020 13:04:56",
      "content": "<p>As we know, some time ago Kaggle introduced dataset ranking. These kernels is one of the ways to get upvotes to datasets (even though these datasets contain only model weights).\nSo if something is to be blamed, then it is gamification.</p>",
      "rawMarkdown": "As we know, some time ago Kaggle introduced dataset ranking. These kernels is one of the ways to get upvotes to datasets (even though these datasets contain only model weights).\nSo if something is to be blamed, then it is gamification.",
      "votes": null
    },
    {
      "id": "707777",
      "postDate": "01/01/2020 13:20:20",
      "content": "<p>Oh, then it makes sense. Inference doesn't help if there is no preprocessing, data augmentation or validation information.</p>",
      "rawMarkdown": "Oh, then it makes sense. Inference doesn't help if there is no preprocessing, data augmentation or validation information.",
      "votes": null
    },
    {
      "id": "707785",
      "postDate": "01/01/2020 13:35:12",
      "content": "<p>Yeah, I couldn't express what I meant in the origin post. My bad.</p>",
      "rawMarkdown": "Yeah, I couldn't express what I meant in the origin post. My bad.",
      "votes": null
    },
    {
      "id": "707790",
      "postDate": "01/01/2020 13:46:33",
      "content": "<p>I actually don’t think this has much to do with the datasets. IMHO it’s mostly a function of two things - kernel-only competitions, and extremely limited GPU resources.</p>",
      "rawMarkdown": "I actually don’t think this has much to do with the datasets. IMHO it’s mostly a function of two things - kernel-only competitions, and extremely limited GPU resources.",
      "votes": null
    },
    {
      "id": "707950",
      "postDate": "01/01/2020 18:46:52",
      "content": "<p>Also, for what it's worth, I am not a big fan of kernel competitions, especially these with \"instant\" private scoring. I usually end up spending more time troubleshooting and guessing how to make a proper submission kernel. The positive effect that these high-scoring kernels have had is to help me in that regard - made my own submission kernels <em>much</em> more efficient, error-proof, and streamlined. </p>",
      "rawMarkdown": "Also, for what it's worth, I am not a big fan of kernel competitions, especially these with \"instant\" private scoring. I usually end up spending more time troubleshooting and guessing how to make a proper submission kernel. The positive effect that these high-scoring kernels have had is to help me in that regard - made my own submission kernels *much* more efficient, error-proof, and streamlined.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 707659,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "01/01/2020 08:56:32",
      "content": "<p>Hi Rafid,\nYou should not be worrying about high scoring kernels being made public, since they can be overfitting on the public test set. If you look at other competitions, teams can have upto an 80 rank jump from public to private leaderboard. Whatever hypothesis you deem correct is something you should pursue and if it is better than the public kernels, it'll just need one small jump to reach the top! :D\nCheers!</p>",
      "votes": null,
      "replies": [
        {
          "id": 707700,
          "author_name": "abyaadrafid",
          "author_url": "",
          "post_date": "01/01/2020 10:20:19",
          "content": "<p>Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 707666,
      "author_name": "iamkhader",
      "author_url": "",
      "post_date": "01/01/2020 09:13:47",
      "content": "<p>hi Rafid,\nIn a recent competition I have seen triple(close to 1000) digit jumps because of overfitting. My submission jumped 400 ranks in final private test. Always trust your validation techniques. Secondly, I think its good that high ranking starter kernels are released as these help us understand stuff that we may be missing in our pipeline. You can pick parts of the kernel and test on your pipeline which is typically what we do in our jobs also, take a new feature engineering technique, a new augmentation, new loss and experiment for results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 707698,
          "author_name": "abyaadrafid",
          "author_url": "",
          "post_date": "01/01/2020 10:17:55",
          "content": "<p>Hi, I wrongly addressed them as starter kernels. I meant to indicate inference kernels. I mean starter kernels are great for learning stuff,  but inference kernels do not fall in the criteria. \nLook at the leaderboard right now, you'll see what I mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707777,
          "author_name": "iamkhader",
          "author_url": "",
          "post_date": "01/01/2020 13:20:20",
          "content": "<p>Oh, then it makes sense. Inference doesn't help if there is no preprocessing, data augmentation or validation information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707785,
          "author_name": "abyaadrafid",
          "author_url": "",
          "post_date": "01/01/2020 13:35:12",
          "content": "<p>Yeah, I couldn't express what I meant in the origin post. My bad.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 707756,
      "author_name": "shayekh",
      "author_url": "",
      "post_date": "01/01/2020 12:32:04",
      "content": "<p>Leader-board shake is one of the most interesting things on Kaggle.\nFor example: <a href=\"https://www.kaggle.com/robikscube/ashrae-leaderboard-and-shake\">https://www.kaggle.com/robikscube/ashrae-leaderboard-and-shake</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 707768,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "01/01/2020 13:04:56",
      "content": "<p>As we know, some time ago Kaggle introduced dataset ranking. These kernels is one of the ways to get upvotes to datasets (even though these datasets contain only model weights).\nSo if something is to be blamed, then it is gamification.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 707790,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "01/01/2020 13:46:33",
      "content": "<p>I actually don’t think this has much to do with the datasets. IMHO it’s mostly a function of two things - kernel-only competitions, and extremely limited GPU resources.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 707950,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "01/01/2020 18:46:52",
      "content": "<p>Also, for what it's worth, I am not a big fan of kernel competitions, especially these with \"instant\" private scoring. I usually end up spending more time troubleshooting and guessing how to make a proper submission kernel. The positive effect that these high-scoring kernels have had is to help me in that regard - made my own submission kernels <em>much</em> more efficient, error-proof, and streamlined. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "707554": "Can we stop making high-scoring \"inference\" kernels public on the leaderboard? \nIt just feels unfair that some people submit other peoples model weights and place higher than people who are trying their best. It's quite depressing. \nCan we work around this anyhow? \n\nNo offense to anyone.\nedited : I meant inference kernels, not starters.",
    "707659": "Hi Rafid,\nYou should not be worrying about high scoring kernels being made public, since they can be overfitting on the public test set. If you look at other competitions, teams can have upto an 80 rank jump from public to private leaderboard. Whatever hypothesis you deem correct is something you should pursue and if it is better than the public kernels, it'll just need one small jump to reach the top! :D\nCheers!",
    "707666": "hi Rafid,\nIn a recent competition I have seen triple(close to 1000) digit jumps because of overfitting. My submission jumped 400 ranks in final private test. Always trust your validation techniques. Secondly, I think its good that high ranking starter kernels are released as these help us understand stuff that we may be missing in our pipeline. You can pick parts of the kernel and test on your pipeline which is typically what we do in our jobs also, take a new feature engineering technique, a new augmentation, new loss and experiment for results.",
    "707698": "Hi, I wrongly addressed them as starter kernels. I meant to indicate inference kernels. I mean starter kernels are great for learning stuff,  but inference kernels do not fall in the criteria. \nLook at the leaderboard right now, you'll see what I mean.",
    "707700": "Thank you.",
    "707756": "Leader-board shake is one of the most interesting things on Kaggle.\nFor example: https://www.kaggle.com/robikscube/ashrae-leaderboard-and-shake",
    "707768": "As we know, some time ago Kaggle introduced dataset ranking. These kernels is one of the ways to get upvotes to datasets (even though these datasets contain only model weights).\nSo if something is to be blamed, then it is gamification.",
    "707777": "Oh, then it makes sense. Inference doesn't help if there is no preprocessing, data augmentation or validation information.",
    "707785": "Yeah, I couldn't express what I meant in the origin post. My bad.",
    "707790": "I actually don’t think this has much to do with the datasets. IMHO it’s mostly a function of two things - kernel-only competitions, and extremely limited GPU resources.",
    "707950": "Also, for what it's worth, I am not a big fan of kernel competitions, especially these with \"instant\" private scoring. I usually end up spending more time troubleshooting and guessing how to make a proper submission kernel. The positive effect that these high-scoring kernels have had is to help me in that regard - made my own submission kernels *much* more efficient, error-proof, and streamlined."
  },
  "source": "meta"
}