{
  "id": 307231,
  "title": "Submitting results from public notebooks",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307231",
  "author_name": "",
  "post_date": "2022-02-13T11:00:14.936913400Z",
  "votes": 17,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>How does everybody feel about submissions made from unchanged public kernels?<br>\nI have put in some effort in creating my own solution which, while not entirely hopeless, does not compare against the winning solutions at all.</p>\n<p>Then I cloned and ran the one that is currently holding the top positions and now I'm in a ninth place without any effort at all.<br>\nWhile I greatly appreciate the effort other kagglers have gone through to create the kernel and make it public I would feel better if I used it to learn from rather than just executing it and submitting the result. Personally this feels like cheating to me, what is the public standpoint on this behavior?</p>",
  "messages": [
    {
      "id": "1688061",
      "postDate": "02/13/2022 11:00:14",
      "content": "<p>Hi all,</p>\n<p>How does everybody feel about submissions made from unchanged public kernels?<br>\nI have put in some effort in creating my own solution which, while not entirely hopeless, does not compare against the winning solutions at all.</p>\n<p>Then I cloned and ran the one that is currently holding the top positions and now I'm in a ninth place without any effort at all.<br>\nWhile I greatly appreciate the effort other kagglers have gone through to create the kernel and make it public I would feel better if I used it to learn from rather than just executing it and submitting the result. Personally this feels like cheating to me, what is the public standpoint on this behavior?</p>",
      "rawMarkdown": "Hi all,\n\nHow does everybody feel about submissions made from unchanged public kernels?\nI have put in some effort in creating my own solution which, while not entirely hopeless, does not compare against the winning solutions at all.\n\nThen I cloned and ran the one that is currently holding the top positions and now I'm in a ninth place without any effort at all.\nWhile I greatly appreciate the effort other kagglers have gone through to create the kernel and make it public I would feel better if I used it to learn from rather than just executing it and submitting the result. Personally this feels like cheating to me, what is the public standpoint on this behavior?",
      "votes": null
    },
    {
      "id": "1688091",
      "postDate": "02/13/2022 11:29:52",
      "content": "<p>I think its about what you are optimising for: If you just want to climb to a higher position on the <strong>Public</strong> Leaderboard, its a great way to copy pasta, ensemble 3-4 high performing kernels and make it to the top of 😎</p>\n<p>Its also not wrong to start with doing that. After a point, we all want to learn and hopefully apply this knowledge to problems that we all care about-for that we need to understand how to build solutions. </p>\n<p>I think Kaggle Grandmaster <a href=\"http://kaggle.com/cpmpml/\" target=\"_blank\">CPMP</a> mentioned that he never uses code from outside without understanding it. </p>\n<p>TL;DR its okay to start by copy pasta and tweaking things but eventually to improve its a necessity to also understand what you're forking and how to build on your own ideas and find the right ideas. </p>",
      "rawMarkdown": "I think its about what you are optimising for: If you just want to climb to a higher position on the **Public** Leaderboard, its a great way to copy pasta, ensemble 3-4 high performing kernels and make it to the top of 😎\n\nIts also not wrong to start with doing that. After a point, we all want to learn and hopefully apply this knowledge to problems that we all care about-for that we need to understand how to build solutions. \n\nI think Kaggle Grandmaster [CPMP](http://kaggle.com/cpmpml/) mentioned that he never uses code from outside without understanding it. \n\nTL;DR its okay to start by copy pasta and tweaking things but eventually to improve its a necessity to also understand what you're forking and how to build on your own ideas and find the right ideas.",
      "votes": null
    },
    {
      "id": "1688105",
      "postDate": "02/13/2022 11:47:00",
      "content": "<p>That certainly makes sense - thanks for the reply. 😊<br>\nI get the point that public LB score is not the same as final score but still since this model scores 0.655 right out of the box and my best try so far with my own solution has scored &lt; 0.15 it still feels like an unfair advantage to me.<br>\nI certainly also agree that studying other people's solutions is the way to improve but still, tweaking a solution that was handed to you doesn't seem to me like the same as building your own, regardless of wether I understand it or not.</p>",
      "rawMarkdown": "That certainly makes sense - thanks for the reply. 😊\nI get the point that public LB score is not the same as final score but still since this model scores 0.655 right out of the box and my best try so far with my own solution has scored < 0.15 it still feels like an unfair advantage to me.\nI certainly also agree that studying other people's solutions is the way to improve but still, tweaking a solution that was handed to you doesn't seem to me like the same as building your own, regardless of wether I understand it or not.",
      "votes": null
    },
    {
      "id": "1688133",
      "postDate": "02/13/2022 12:04:36",
      "content": "<p>I share the same frustration as you, we all are here to learn, even the Grandmasters. We're all at different learning stages however. </p>\n<p>I think remembering that the people putting out the kernels that get you high scores also started a while ago at the same position as us is helpful. At the end of any competition if we optimize for learning, hopefully our starter solutions will improve in accuracy a lot too :) </p>",
      "rawMarkdown": "I share the same frustration as you, we all are here to learn, even the Grandmasters. We're all at different learning stages however. \n\nI think remembering that the people putting out the kernels that get you high scores also started a while ago at the same position as us is helpful. At the end of any competition if we optimize for learning, hopefully our starter solutions will improve in accuracy a lot too :)",
      "votes": null
    },
    {
      "id": "1688601",
      "postDate": "02/13/2022 18:10:21",
      "content": "<p>For the record, publicly shared code is open, and you are free to use it as you like. It’s up to you what you think is best.</p>\n<p>In my opinion it’s all good because public lb is not the final lb, so it’s no big deal. I suggest you try to figure out why a public solution is performing better. Try breakdown ALL the differences between that solution and yours. Slowly add the pieces you think will work. Test it with cross validation test sets. After all that work, understanding, and reimplementing I say you deserve (although you don’t need any permission) to use those ideas. Then try to think of any way to slightly improve the new solution you have. </p>\n<p>With all the testing you might find that you can’t get the same score unless you copy the solution EXACTLY the same with all parameters and seeds. This would tell me the model is overfit to the public lb, which is the most awesome thing, because all your work will allow to jump those who just submitted the solution as is.</p>",
      "rawMarkdown": "For the record, publicly shared code is open, and you are free to use it as you like. It’s up to you what you think is best.\n\nIn my opinion it’s all good because public lb is not the final lb, so it’s no big deal. I suggest you try to figure out why a public solution is performing better. Try breakdown ALL the differences between that solution and yours. Slowly add the pieces you think will work. Test it with cross validation test sets. After all that work, understanding, and reimplementing I say you deserve (although you don’t need any permission) to use those ideas. Then try to think of any way to slightly improve the new solution you have. \n\nWith all the testing you might find that you can’t get the same score unless you copy the solution EXACTLY the same with all parameters and seeds. This would tell me the model is overfit to the public lb, which is the most awesome thing, because all your work will allow to jump those who just submitted the solution as is.",
      "votes": null
    },
    {
      "id": "1689383",
      "postDate": "02/14/2022 07:35:12",
      "content": "<p>Thank you very much for your advice Chris. Much appreciated. 😊<br>\nWith regards to comparing the individual parts of the solution I think the approach I have chosen is too different for that to work. I guess there is an important learning in that as well; The solution I have chosen is simply not optimal for this problem so I should move on and try something else instead of tweaking what I have (and the smart thing would be to use what I have been given in the public solution for the base of that).</p>",
      "rawMarkdown": "Thank you very much for your advice Chris. Much appreciated. 😊\nWith regards to comparing the individual parts of the solution I think the approach I have chosen is too different for that to work. I guess there is an important learning in that as well; The solution I have chosen is simply not optimal for this problem so I should move on and try something else instead of tweaking what I have (and the smart thing would be to use what I have been given in the public solution for the base of that).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1688091,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/13/2022 11:29:52",
      "content": "<p>I think its about what you are optimising for: If you just want to climb to a higher position on the <strong>Public</strong> Leaderboard, its a great way to copy pasta, ensemble 3-4 high performing kernels and make it to the top of 😎</p>\n<p>Its also not wrong to start with doing that. After a point, we all want to learn and hopefully apply this knowledge to problems that we all care about-for that we need to understand how to build solutions. </p>\n<p>I think Kaggle Grandmaster <a href=\"http://kaggle.com/cpmpml/\" target=\"_blank\">CPMP</a> mentioned that he never uses code from outside without understanding it. </p>\n<p>TL;DR its okay to start by copy pasta and tweaking things but eventually to improve its a necessity to also understand what you're forking and how to build on your own ideas and find the right ideas. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1688105,
          "author_name": "larsmadsen",
          "author_url": "",
          "post_date": "02/13/2022 11:47:00",
          "content": "<p>That certainly makes sense - thanks for the reply. 😊<br>\nI get the point that public LB score is not the same as final score but still since this model scores 0.655 right out of the box and my best try so far with my own solution has scored &lt; 0.15 it still feels like an unfair advantage to me.<br>\nI certainly also agree that studying other people's solutions is the way to improve but still, tweaking a solution that was handed to you doesn't seem to me like the same as building your own, regardless of wether I understand it or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688133,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/13/2022 12:04:36",
          "content": "<p>I share the same frustration as you, we all are here to learn, even the Grandmasters. We're all at different learning stages however. </p>\n<p>I think remembering that the people putting out the kernels that get you high scores also started a while ago at the same position as us is helpful. At the end of any competition if we optimize for learning, hopefully our starter solutions will improve in accuracy a lot too :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1688601,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "02/13/2022 18:10:21",
      "content": "<p>For the record, publicly shared code is open, and you are free to use it as you like. It’s up to you what you think is best.</p>\n<p>In my opinion it’s all good because public lb is not the final lb, so it’s no big deal. I suggest you try to figure out why a public solution is performing better. Try breakdown ALL the differences between that solution and yours. Slowly add the pieces you think will work. Test it with cross validation test sets. After all that work, understanding, and reimplementing I say you deserve (although you don’t need any permission) to use those ideas. Then try to think of any way to slightly improve the new solution you have. </p>\n<p>With all the testing you might find that you can’t get the same score unless you copy the solution EXACTLY the same with all parameters and seeds. This would tell me the model is overfit to the public lb, which is the most awesome thing, because all your work will allow to jump those who just submitted the solution as is.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689383,
          "author_name": "larsmadsen",
          "author_url": "",
          "post_date": "02/14/2022 07:35:12",
          "content": "<p>Thank you very much for your advice Chris. Much appreciated. 😊<br>\nWith regards to comparing the individual parts of the solution I think the approach I have chosen is too different for that to work. I guess there is an important learning in that as well; The solution I have chosen is simply not optimal for this problem so I should move on and try something else instead of tweaking what I have (and the smart thing would be to use what I have been given in the public solution for the base of that).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1688061": "Hi all,\n\nHow does everybody feel about submissions made from unchanged public kernels?\nI have put in some effort in creating my own solution which, while not entirely hopeless, does not compare against the winning solutions at all.\n\nThen I cloned and ran the one that is currently holding the top positions and now I'm in a ninth place without any effort at all.\nWhile I greatly appreciate the effort other kagglers have gone through to create the kernel and make it public I would feel better if I used it to learn from rather than just executing it and submitting the result. Personally this feels like cheating to me, what is the public standpoint on this behavior?",
    "1688091": "I think its about what you are optimising for: If you just want to climb to a higher position on the **Public** Leaderboard, its a great way to copy pasta, ensemble 3-4 high performing kernels and make it to the top of 😎\n\nIts also not wrong to start with doing that. After a point, we all want to learn and hopefully apply this knowledge to problems that we all care about-for that we need to understand how to build solutions. \n\nI think Kaggle Grandmaster [CPMP](http://kaggle.com/cpmpml/) mentioned that he never uses code from outside without understanding it. \n\nTL;DR its okay to start by copy pasta and tweaking things but eventually to improve its a necessity to also understand what you're forking and how to build on your own ideas and find the right ideas.",
    "1688105": "That certainly makes sense - thanks for the reply. 😊\nI get the point that public LB score is not the same as final score but still since this model scores 0.655 right out of the box and my best try so far with my own solution has scored < 0.15 it still feels like an unfair advantage to me.\nI certainly also agree that studying other people's solutions is the way to improve but still, tweaking a solution that was handed to you doesn't seem to me like the same as building your own, regardless of wether I understand it or not.",
    "1688133": "I share the same frustration as you, we all are here to learn, even the Grandmasters. We're all at different learning stages however. \n\nI think remembering that the people putting out the kernels that get you high scores also started a while ago at the same position as us is helpful. At the end of any competition if we optimize for learning, hopefully our starter solutions will improve in accuracy a lot too :)",
    "1688601": "For the record, publicly shared code is open, and you are free to use it as you like. It’s up to you what you think is best.\n\nIn my opinion it’s all good because public lb is not the final lb, so it’s no big deal. I suggest you try to figure out why a public solution is performing better. Try breakdown ALL the differences between that solution and yours. Slowly add the pieces you think will work. Test it with cross validation test sets. After all that work, understanding, and reimplementing I say you deserve (although you don’t need any permission) to use those ideas. Then try to think of any way to slightly improve the new solution you have. \n\nWith all the testing you might find that you can’t get the same score unless you copy the solution EXACTLY the same with all parameters and seeds. This would tell me the model is overfit to the public lb, which is the most awesome thing, because all your work will allow to jump those who just submitted the solution as is.",
    "1689383": "Thank you very much for your advice Chris. Much appreciated. 😊\nWith regards to comparing the individual parts of the solution I think the approach I have chosen is too different for that to work. I guess there is an important learning in that as well; The solution I have chosen is simply not optimal for this problem so I should move on and try something else instead of tweaking what I have (and the smart thing would be to use what I have been given in the public solution for the base of that)."
  },
  "source": "meta"
}