{
  "id": 78239,
  "title": "My Takeaways from the NFL Punt Competition",
  "url": "/competitions/NFL-Punt-Analytics-Competition/discussion/78239",
  "author_name": "",
  "post_date": "2019-01-21T15:14:17.389912600Z",
  "votes": 27,
  "comment_count": 28,
  "views": 0,
  "content": "<p>I’m very thankful for Kaggle and the NFL for hosting this competition, it was a lot of fun and I learned a lot. The winners clearly did an amazing job and should be applauded for their hard work. As with most competitions, there are more losers than winners- and I'm sure many of us wish we were chosen. Regardless I do feel like it’s important to look back and grow from the experience. It’s been a few days since the winners were announced and I’ve had a chance to collect my thoughts and put them on paper. This may be more of a therapeutic way for me to finally close the book on this competition than anything :)</p>\n\n<p>If I could do it all over again:</p>\n\n<ul>\n<li>I would have focused more on the slides and presentation. Most of my work was spent digging into every angle of the data. Trying to understand and visualize the NGS data, and generally gaining an understanding the results and formations of punting plays. I wanted to make sure I investigated any possible area that could reduce the possibility of concussions. I assumed that the slides would come naturally as an extension of this analysis. Now that it’s over my gut feeling is that the judges used the slides as a starting point to evaluate submissions and dove into the kernels to confirm that it supported their conclusions. If I had known this going in I would’ve worked backwards from a polished power-point and then supported it with my kernel.</li>\n<li>I would have made my presentation much, much shorter and to the point. I had assumed that some of non-technical judges might only look at the slides slides only- so I tried to include a lot of my analysis in my slides to show all the word I had done. The unintended consequence of this was that it was way too much to digest in a 5-10 minute presentation. At the end of the day this is for a PR event where the NFL wanted to hear quick pitches about proposals backed by data- not digest an entire research report.</li>\n<li>Similar to my last point I would’ve made my analysis less exhaustive, more focused, and concise. I added a lot of analysis to my kernel that was more exploratory and didn’t impact my final rule proposals directly- I don’t think adding this extra analysis helped to keep things clear. I noticed that ‘concise’ was specifically mentioned when describing the winning submissions. Again, this makes sense with the whole 5-10 minute presentation thing.</li>\n<li>I would've focused more on telling a cohesive story. The winning submissions did an excellent job of this in their kernels and I'm sure also in their slides. This is always the hardest part.</li>\n<li>I would have tried to think more outside the box with my rule suggestions to differentiate it from others. It does appear that the winners each had their own angle on the rule proposal and the judges selected them in a way that there wasn’t a lot of overlap. As frustrating as it is to very similar proposals being selected, but the judges obviously thought they did a better job reaching their conclusions in a more unique, and clearer way than I did. I would have never thought of suggesting something like putting sensors in helmets as a rule proposal- it was very clever and was rewarded for being so.</li>\n<li>I would have worked smarter, not harder. Especially when it came to the NGS data. Plotting and reviewing player routes was a big part of my approach, but in the end wasn’t digestible for a 5 minute presentation- so it didn’t matter. The winning submissions focused less on the NGS data and instead on supporting their rule changes in a way that anyone could understand. My approach to quantify each play’s risk level was not the type of thing I was going to be able to pitch to a general audience. I will say, after going back and reviewing @awainger submission I was blown away by the clean and clear way he used the NGS data. Really outstanding work and I learned a lot.</li>\n<li>I would've made my preprocessing all run within a kernel. By the time I found out this was a requirement I had already gone too far into preprocessing offline. It probably wasn't a deciding factor in the end, but it's something I regret not doing it and can't help but notice all of the winning submissions ran everything in kernels.</li>\n<li>Would've shared my EDA kernel publicly early on. I took the approach that many others did of not wanting my analysis to get out, but I think I could've benefited from receiving feedback earlier. Share more, share early, everyone benefits.</li>\n<li>Probably the biggest for me: I would’ve not believed the hype. No matter how many people tell you you’re a ‘shoe-in’ for something, it doesn’t mean anything until the results are in. I appreciated all the kind comments I got about my kernel, but I also knew that it was setting myself for disappointment- and made missing the cut hurt even more. I guess that’s just a general life lesson.</li>\n</ul>\n\n<p>Suggestions and feedback to Kaggle for future Analytics competitions:</p>\n\n<ul>\n<li>This competition had a fairly quick turnaround (which I really liked). I appreciate the responsiveness and feedback we did receive, but since things were moving relatively fast, I think the message board needed quick responses to questions from participants. This is especially true with the post detailing what content should be in the slides/kernel which that came out about a week before the deadline. If we had known this information up front I think many of us would’ve approached things differently.</li>\n<li>Be clear about how the submissions are being judged and who will be reviewing them. Knowing the audience turned out to be really important. The evaluation section of the competition said the submissions would be judged on Solution Efficacy and Game Integrity- but it didn’t explain who was deciding what is or isn’t an effective solution. The audience was loosely clarified in a discussion post but should have been stated up front. Ideally this type of competition would say “We have a panel of X judges with positions A, B, and C who will each vote on their top 4 and the submissions”. This would also help with transparency and make it feel like submissions weren’t cherry picked by a company to support their preconceived beliefs (I’m not saying that’s the case here, but transparency couldn’t hurt). It also might help the participants who didn't win at least see if they were close to making the cut by seeing how many votes they received.</li>\n<li>Take into account peer reviews. I think Kaggle should seriously consider this in future competitions. If a large portion of the participants seem to appreciate a submission, it should bear some weight in the review process. If it doesn't then either we are being told: (1) many of the participants didn't actually understand what the objective of the competition actually was or (2) The judges know better than those who had spent time working on the competition. I could see there being a valid argument for #2 - but isn't that the whole point of crowdsourcing for ideas? I know Kaggle comments/upvotes can be easily doctored or manipulated- so that can be a concern. Still, there must be some be clever way to include peer reviews (even if only as a superlative) in future competitions while minimizing the possibility for manipulation.</li>\n</ul>\n\n<p>....and by the length of this post it's clear I haven't learned how to be brief and concise yet :D</p>\n\n<p>So in the end, knowing what I do now- Do I regret having spending all the time I did working on this competition? Yes.  If the NFL and Kaggle announced another competition today on a different topic would I participate? Absolutely yes- but I’d take an entirely different approach.</p>\n\n<p>Best of luck to all the winners in Atlanta!</p>",
  "messages": [
    {
      "id": "459312",
      "postDate": "01/21/2019 15:14:17",
      "content": "<p>I’m very thankful for Kaggle and the NFL for hosting this competition, it was a lot of fun and I learned a lot. The winners clearly did an amazing job and should be applauded for their hard work. As with most competitions, there are more losers than winners- and I'm sure many of us wish we were chosen. Regardless I do feel like it’s important to look back and grow from the experience. It’s been a few days since the winners were announced and I’ve had a chance to collect my thoughts and put them on paper. This may be more of a therapeutic way for me to finally close the book on this competition than anything :)</p>\n\n<p>If I could do it all over again:</p>\n\n<ul>\n<li>I would have focused more on the slides and presentation. Most of my work was spent digging into every angle of the data. Trying to understand and visualize the NGS data, and generally gaining an understanding the results and formations of punting plays. I wanted to make sure I investigated any possible area that could reduce the possibility of concussions. I assumed that the slides would come naturally as an extension of this analysis. Now that it’s over my gut feeling is that the judges used the slides as a starting point to evaluate submissions and dove into the kernels to confirm that it supported their conclusions. If I had known this going in I would’ve worked backwards from a polished power-point and then supported it with my kernel.</li>\n<li>I would have made my presentation much, much shorter and to the point. I had assumed that some of non-technical judges might only look at the slides slides only- so I tried to include a lot of my analysis in my slides to show all the word I had done. The unintended consequence of this was that it was way too much to digest in a 5-10 minute presentation. At the end of the day this is for a PR event where the NFL wanted to hear quick pitches about proposals backed by data- not digest an entire research report.</li>\n<li>Similar to my last point I would’ve made my analysis less exhaustive, more focused, and concise. I added a lot of analysis to my kernel that was more exploratory and didn’t impact my final rule proposals directly- I don’t think adding this extra analysis helped to keep things clear. I noticed that ‘concise’ was specifically mentioned when describing the winning submissions. Again, this makes sense with the whole 5-10 minute presentation thing.</li>\n<li>I would've focused more on telling a cohesive story. The winning submissions did an excellent job of this in their kernels and I'm sure also in their slides. This is always the hardest part.</li>\n<li>I would have tried to think more outside the box with my rule suggestions to differentiate it from others. It does appear that the winners each had their own angle on the rule proposal and the judges selected them in a way that there wasn’t a lot of overlap. As frustrating as it is to very similar proposals being selected, but the judges obviously thought they did a better job reaching their conclusions in a more unique, and clearer way than I did. I would have never thought of suggesting something like putting sensors in helmets as a rule proposal- it was very clever and was rewarded for being so.</li>\n<li>I would have worked smarter, not harder. Especially when it came to the NGS data. Plotting and reviewing player routes was a big part of my approach, but in the end wasn’t digestible for a 5 minute presentation- so it didn’t matter. The winning submissions focused less on the NGS data and instead on supporting their rule changes in a way that anyone could understand. My approach to quantify each play’s risk level was not the type of thing I was going to be able to pitch to a general audience. I will say, after going back and reviewing @awainger submission I was blown away by the clean and clear way he used the NGS data. Really outstanding work and I learned a lot.</li>\n<li>I would've made my preprocessing all run within a kernel. By the time I found out this was a requirement I had already gone too far into preprocessing offline. It probably wasn't a deciding factor in the end, but it's something I regret not doing it and can't help but notice all of the winning submissions ran everything in kernels.</li>\n<li>Would've shared my EDA kernel publicly early on. I took the approach that many others did of not wanting my analysis to get out, but I think I could've benefited from receiving feedback earlier. Share more, share early, everyone benefits.</li>\n<li>Probably the biggest for me: I would’ve not believed the hype. No matter how many people tell you you’re a ‘shoe-in’ for something, it doesn’t mean anything until the results are in. I appreciated all the kind comments I got about my kernel, but I also knew that it was setting myself for disappointment- and made missing the cut hurt even more. I guess that’s just a general life lesson.</li>\n</ul>\n\n<p>Suggestions and feedback to Kaggle for future Analytics competitions:</p>\n\n<ul>\n<li>This competition had a fairly quick turnaround (which I really liked). I appreciate the responsiveness and feedback we did receive, but since things were moving relatively fast, I think the message board needed quick responses to questions from participants. This is especially true with the post detailing what content should be in the slides/kernel which that came out about a week before the deadline. If we had known this information up front I think many of us would’ve approached things differently.</li>\n<li>Be clear about how the submissions are being judged and who will be reviewing them. Knowing the audience turned out to be really important. The evaluation section of the competition said the submissions would be judged on Solution Efficacy and Game Integrity- but it didn’t explain who was deciding what is or isn’t an effective solution. The audience was loosely clarified in a discussion post but should have been stated up front. Ideally this type of competition would say “We have a panel of X judges with positions A, B, and C who will each vote on their top 4 and the submissions”. This would also help with transparency and make it feel like submissions weren’t cherry picked by a company to support their preconceived beliefs (I’m not saying that’s the case here, but transparency couldn’t hurt). It also might help the participants who didn't win at least see if they were close to making the cut by seeing how many votes they received.</li>\n<li>Take into account peer reviews. I think Kaggle should seriously consider this in future competitions. If a large portion of the participants seem to appreciate a submission, it should bear some weight in the review process. If it doesn't then either we are being told: (1) many of the participants didn't actually understand what the objective of the competition actually was or (2) The judges know better than those who had spent time working on the competition. I could see there being a valid argument for #2 - but isn't that the whole point of crowdsourcing for ideas? I know Kaggle comments/upvotes can be easily doctored or manipulated- so that can be a concern. Still, there must be some be clever way to include peer reviews (even if only as a superlative) in future competitions while minimizing the possibility for manipulation.</li>\n</ul>\n\n<p>....and by the length of this post it's clear I haven't learned how to be brief and concise yet :D</p>\n\n<p>So in the end, knowing what I do now- Do I regret having spending all the time I did working on this competition? Yes.  If the NFL and Kaggle announced another competition today on a different topic would I participate? Absolutely yes- but I’d take an entirely different approach.</p>\n\n<p>Best of luck to all the winners in Atlanta!</p>",
      "rawMarkdown": "I’m very thankful for Kaggle and the NFL for hosting this competition, it was a lot of fun and I learned a lot. The winners clearly did an amazing job and should be applauded for their hard work. As with most competitions, there are more losers than winners- and I'm sure many of us wish we were chosen. Regardless I do feel like it’s important to look back and grow from the experience. It’s been a few days since the winners were announced and I’ve had a chance to collect my thoughts and put them on paper. This may be more of a therapeutic way for me to finally close the book on this competition than anything :)\n\nIf I could do it all over again:\n\n- I would have focused more on the slides and presentation. Most of my work was spent digging into every angle of the data. Trying to understand and visualize the NGS data, and generally gaining an understanding the results and formations of punting plays. I wanted to make sure I investigated any possible area that could reduce the possibility of concussions. I assumed that the slides would come naturally as an extension of this analysis. Now that it’s over my gut feeling is that the judges used the slides as a starting point to evaluate submissions and dove into the kernels to confirm that it supported their conclusions. If I had known this going in I would’ve worked backwards from a polished power-point and then supported it with my kernel.\n- I would have made my presentation much, much shorter and to the point. I had assumed that some of non-technical judges might only look at the slides slides only- so I tried to include a lot of my analysis in my slides to show all the word I had done. The unintended consequence of this was that it was way too much to digest in a 5-10 minute presentation. At the end of the day this is for a PR event where the NFL wanted to hear quick pitches about proposals backed by data- not digest an entire research report.\n- Similar to my last point I would’ve made my analysis less exhaustive, more focused, and concise. I added a lot of analysis to my kernel that was more exploratory and didn’t impact my final rule proposals directly- I don’t think adding this extra analysis helped to keep things clear. I noticed that ‘concise’ was specifically mentioned when describing the winning submissions. Again, this makes sense with the whole 5-10 minute presentation thing.\n- I would've focused more on telling a cohesive story. The winning submissions did an excellent job of this in their kernels and I'm sure also in their slides. This is always the hardest part.\n- I would have tried to think more outside the box with my rule suggestions to differentiate it from others. It does appear that the winners each had their own angle on the rule proposal and the judges selected them in a way that there wasn’t a lot of overlap. As frustrating as it is to very similar proposals being selected, but the judges obviously thought they did a better job reaching their conclusions in a more unique, and clearer way than I did. I would have never thought of suggesting something like putting sensors in helmets as a rule proposal- it was very clever and was rewarded for being so.\n- I would have worked smarter, not harder. Especially when it came to the NGS data. Plotting and reviewing player routes was a big part of my approach, but in the end wasn’t digestible for a 5 minute presentation- so it didn’t matter. The winning submissions focused less on the NGS data and instead on supporting their rule changes in a way that anyone could understand. My approach to quantify each play’s risk level was not the type of thing I was going to be able to pitch to a general audience. I will say, after going back and reviewing @awainger submission I was blown away by the clean and clear way he used the NGS data. Really outstanding work and I learned a lot.\n- I would've made my preprocessing all run within a kernel. By the time I found out this was a requirement I had already gone too far into preprocessing offline. It probably wasn't a deciding factor in the end, but it's something I regret not doing it and can't help but notice all of the winning submissions ran everything in kernels.\n- Would've shared my EDA kernel publicly early on. I took the approach that many others did of not wanting my analysis to get out, but I think I could've benefited from receiving feedback earlier. Share more, share early, everyone benefits.\n- Probably the biggest for me: I would’ve not believed the hype. No matter how many people tell you you’re a ‘shoe-in’ for something, it doesn’t mean anything until the results are in. I appreciated all the kind comments I got about my kernel, but I also knew that it was setting myself for disappointment- and made missing the cut hurt even more. I guess that’s just a general life lesson.\n\nSuggestions and feedback to Kaggle for future Analytics competitions:\n\n- This competition had a fairly quick turnaround (which I really liked). I appreciate the responsiveness and feedback we did receive, but since things were moving relatively fast, I think the message board needed quick responses to questions from participants. This is especially true with the post detailing what content should be in the slides/kernel which that came out about a week before the deadline. If we had known this information up front I think many of us would’ve approached things differently.\n- Be clear about how the submissions are being judged and who will be reviewing them. Knowing the audience turned out to be really important. The evaluation section of the competition said the submissions would be judged on Solution Efficacy and Game Integrity- but it didn’t explain who was deciding what is or isn’t an effective solution. The audience was loosely clarified in a discussion post but should have been stated up front. Ideally this type of competition would say “We have a panel of X judges with positions A, B, and C who will each vote on their top 4 and the submissions”. This would also help with transparency and make it feel like submissions weren’t cherry picked by a company to support their preconceived beliefs (I’m not saying that’s the case here, but transparency couldn’t hurt). It also might help the participants who didn't win at least see if they were close to making the cut by seeing how many votes they received.\n- Take into account peer reviews. I think Kaggle should seriously consider this in future competitions. If a large portion of the participants seem to appreciate a submission, it should bear some weight in the review process. If it doesn't then either we are being told: (1) many of the participants didn't actually understand what the objective of the competition actually was or (2) The judges know better than those who had spent time working on the competition. I could see there being a valid argument for #2 - but isn't that the whole point of crowdsourcing for ideas? I know Kaggle comments/upvotes can be easily doctored or manipulated- so that can be a concern. Still, there must be some be clever way to include peer reviews (even if only as a superlative) in future competitions while minimizing the possibility for manipulation.\n\n....and by the length of this post it's clear I haven't learned how to be brief and concise yet :D\n\nSo in the end, knowing what I do now- Do I regret having spending all the time I did working on this competition? Yes.  If the NFL and Kaggle announced another competition today on a different topic would I participate? Absolutely yes- but I’d take an entirely different approach.\n\nBest of luck to all the winners in Atlanta!",
      "votes": null
    },
    {
      "id": "459415",
      "postDate": "01/21/2019 17:22:55",
      "content": "<h3>Our main goals</h3>\n\n<p>Given that this competition forum is not that active I do not want to start a new thread.\nI agree with most points of Rob. I will simply flood this topic with my additional thoughts.\nI try to add several comments each have the best intent to improve similar further competitions.\nObviously I am not unbiased, as Rob (and sure others as well) we were hoping/expected to go to Atlanta. </p>\n\n<p>We read the <a href=\"https://www.kaggle.com/c/NFL-Punt-Analytics-Competition#evaluation\">Evaluation</a> page very carefully and always kept in mind each and every point.</p>\n\n<p>Our main goals were</p>\n\n<ul>\n<li>Create an excellent quality presentation as many of the judges might not care about the source code</li>\n<li>Explore all the provided data do not leave rocks unturned</li>\n<li>Collect more data to have more solid, statistically significant results</li>\n<li>Expect broad audience and do not go into too much technical details in the slides and summary</li>\n<li>Try to suggest several different rule modifications but highlight the one with the biggest impact</li>\n</ul>",
      "rawMarkdown": "### Our main goals\nGiven that this competition forum is not that active I do not want to start a new thread.\nI agree with most points of Rob. I will simply flood this topic with my additional thoughts.\nI try to add several comments each have the best intent to improve similar further competitions.\nObviously I am not unbiased, as Rob (and sure others as well) we were hoping/expected to go to Atlanta. \n \nWe read the [Evaluation](https://www.kaggle.com/c/NFL-Punt-Analytics-Competition#evaluation) page very carefully and always kept in mind each and every point.\n\nOur main goals were\n\n* Create an excellent quality presentation as many of the judges might not care about the source code\n* Explore all the provided data do not leave rocks unturned\n* Collect more data to have more solid, statistically significant results\n* Expect broad audience and do not go into too much technical details in the slides and summary\n* Try to suggest several different rule modifications but highlight the one with the biggest impact",
      "votes": null
    },
    {
      "id": "459418",
      "postDate": "01/21/2019 17:32:39",
      "content": "<h3>Anonimity</h3>\n\n<p>There was a strange question in the beginning of the competition about bias and discrimination.</p>\n\n<p>&gt; <strong>Chris Crawford wrote</strong>\n&gt; \n&gt; &gt; The answer is yes and no. We do a little to prevent bias. For example, we don't show your username or email address to the hosts when we give them the final list of submissions, but they'll obviously see who you are when they see your kernel. \n&gt; \n&gt; If you're worried about it, I give everyone permission to use my avatar so we all look the same :)</p>\n\n<p>I believe we had a compelling story (everyone likes to cheer to the small team who does not really has a chance, right? :))\nWe decided to put our work into focus and only shared who we are after the results were finalized. \nGiven that we are talking about PR event that was probably a mistake.  </p>",
      "rawMarkdown": "### Anonimity \n\nThere was a strange question in the beginning of the competition about bias and discrimination.\n\n&gt; **Chris Crawford wrote**\n&gt; \n&gt; &gt; The answer is yes and no. We do a little to prevent bias. For example, we don't show your username or email address to the hosts when we give them the final list of submissions, but they'll obviously see who you are when they see your kernel. \n&gt; \n&gt; If you're worried about it, I give everyone permission to use my avatar so we all look the same :)\n\nI believe we had a compelling story (everyone likes to cheer to the small team who does not really has a chance, right? :))\nWe decided to put our work into focus and only shared who we are after the results were finalized. \nGiven that we are talking about PR event that was probably a mistake.",
      "votes": null
    },
    {
      "id": "459423",
      "postDate": "01/21/2019 17:48:13",
      "content": "<h3>USA &amp; Domain Knowledge</h3>\n\n<p>We live in Hungary where soccer is way more popular than any other sports (imho way more than it should be). \nEven though my friend is a huge NFL fan and knows the game quite well we had to work hard to understand the fine details of the rules.</p>\n\n<p>Actually we were more afraid of professional sport analysts/ PhD students with relevant research area. </p>\n\n<p>I am not saying this to claim some kind of consolation prize.\nThis is a competition and let the best team win!\nWe certainly had more chance here than would had against Prof. Keld Helsgaun and William Cook in the other \n<a href=\"https://www.kaggle.com/c/traveling-santa-2018-prime-paths/discussion/77134\">Traveling Santa Competition</a> :)</p>",
      "rawMarkdown": "### USA &amp; Domain Knowledge\n\nWe live in Hungary where soccer is way more popular than any other sports (imho way more than it should be). \nEven though my friend is a huge NFL fan and knows the game quite well we had to work hard to understand the fine details of the rules.\n\nActually we were more afraid of professional sport analysts/ PhD students with relevant research area. \n\nI am not saying this to claim some kind of consolation prize.\nThis is a competition and let the best team win!\nWe certainly had more chance here than would had against Prof. Keld Helsgaun and William Cook in the other \n[Traveling Santa Competition](https://www.kaggle.com/c/traveling-santa-2018-prime-paths/discussion/77134) :)",
      "votes": null
    },
    {
      "id": "459436",
      "postDate": "01/21/2019 18:21:07",
      "content": "<h3>Transparency of the decision</h3>\n\n<p>In the other thread many others raised good points about how to make the selection process more transparent. I won't repeat them just wanted to emphasize the importance of it. We learned a lot during the competition but we don't have a clue what we missed or what to improve next. </p>\n\n<p>We talked about single and double coverage, it feels that we had an additional <a href=\"https://en.wikipedia.org/wiki/Maximum_coverage_problem\"><strong>maximum coverage problem</strong></a>.\n<img src=\"https://s3-eu-west-1.amazonaws.com/nfl-punt-analytics/MaxCov.png\" alt=\"\"></p>\n\n<p>It is rational from the host. We tried to mitigate the risk by suggesting more rules and having unique methods and suggestions. Unfortunately the best submissions had quite a few overlapping rule sets. I have to agree with Rob from the other thread:</p>\n\n<blockquote>\n  <p><strong>Rob Mulla wrote</strong>\n  I had an idea for my rule change proposals fairly early on and was pretty devisdtated when I saw many others with the same ideas. But in the end, just like you said it’s pretty cool to see people coming to the same conclusions from different angles. </p>\n</blockquote>",
      "rawMarkdown": "### Transparency of the decision\nIn the other thread many others raised good points about how to make the selection process more transparent. I won't repeat them just wanted to emphasize the importance of it. We learned a lot during the competition but we don't have a clue what we missed or what to improve next. \n\nWe talked about single and double coverage, it feels that we had an additional [**maximum coverage problem**](https://en.wikipedia.org/wiki/Maximum_coverage_problem).\n![](https://s3-eu-west-1.amazonaws.com/nfl-punt-analytics/MaxCov.png)\n\nIt is rational from the host. We tried to mitigate the risk by suggesting more rules and having unique methods and suggestions. Unfortunately the best submissions had quite a few overlapping rule sets. I have to agree with Rob from the other thread:\n\n&gt; **Rob Mulla wrote**\n&gt; I had an idea for my rule change proposals fairly early on and was pretty devisdtated when I saw many others with the same ideas. But in the end, just like you said it’s pretty cool to see people coming to the same conclusions from different angles.",
      "votes": null
    },
    {
      "id": "459445",
      "postDate": "01/21/2019 18:42:31",
      "content": "<h3>Ambiguity of requirements</h3>\n\n<p>Even the last minute <a href=\"https://www.kaggle.com/c/NFL-Punt-Analytics-Competition/discussion/76622\">clarification</a> was not clear about what is required from the presentation. 5-10 minutes presentation with 1-50 slides is a bit vague. We aimed to be somewhere in the middle and thanks to my teammate we had an excellent deck with ~24 slides. Maybe it was to long, but we knew we could present the main part easily to any audience who already read our summary report.</p>\n\n<p>Not knowing the audience (e.g. former players, coaches, journalists, doctors, analysts, data scientists etc.) made it very difficult to balance between being too shallow and going too deep.</p>",
      "rawMarkdown": "### Ambiguity of requirements\n\nEven the last minute [clarification](https://www.kaggle.com/c/NFL-Punt-Analytics-Competition/discussion/76622) was not clear about what is required from the presentation. 5-10 minutes presentation with 1-50 slides is a bit vague. We aimed to be somewhere in the middle and thanks to my teammate we had an excellent deck with ~24 slides. Maybe it was to long, but we knew we could present the main part easily to any audience who already read our summary report.\n\nNot knowing the audience (e.g. former players, coaches, journalists, doctors, analysts, data scientists etc.) made it very difficult to balance between being too shallow and going too deep.",
      "votes": null
    },
    {
      "id": "459451",
      "postDate": "01/21/2019 19:02:33",
      "content": "<p>@RobMulla, let me remind you that you  did a great job. Your risk normalization was a brilliant idea, data augmentation when it was really needed.  Plus you were one of the few who kept alive the almost desert forums.</p>\n\n<p>About votes in this competition, I think there was some distortion because of the competition format. There are many kernels that possibly under \"normal\" circumstances would have had some votes that are almost unnoticed (Funnily, I had not even read the elegant solution of <a href=\"/awainger\">@awainger</a> before winners were announced. 135 kernels are a lot...)</p>\n\n<p>About actual optimal format of presentation... who knows. I agree that it would be good to have more clear guidelines, or even a template,so that efforts can be better directed. But in almost all Kaggle competition I've taken part in there's noise and some degree of surprises...\nI find a huge problem the fact that kernels can not be shared before and can not be digested after... Maybe too \npesimistic but I'm sure early sharing would lead to other kind of problems.</p>\n\n<p>I personally enjoyed, I had zero domain knowledge before beginning, a \"martian\" seeing people crash following a ball. This, I think, was a huge dissadvantage in this competition. But disciplined process led to some signal in the noise what was a big satisfaction. Also to experience this format was interesting for me.</p>\n\n<p>About the format, I think it probably achieves expected result for the organizers but it is much harder for kagglers  that take part on it;  ironically there's a high chance your hard worked kernel will get a fraction of the votes a quality kernel would get in a \"normal\" competition. And no ranking feedback, or points...</p>",
      "rawMarkdown": "RobMulla, let me remind you that you  did a great job. Your risk normalization was a brilliant idea, data augmentation when it was really needed.  Plus you were one of the few who kept alive the almost desert forums.\n\nAbout votes in this competition, I think there was some distortion because of the competition format. There are many kernels that possibly under \"normal\" circumstances would have had some votes that are almost unnoticed (Funnily, I had not even read the elegant solution of @awainger before winners were announced. 135 kernels are a lot...)\n\nAbout actual optimal format of presentation... who knows. I agree that it would be good to have more clear guidelines, or even a template,so that efforts can be better directed. But in almost all Kaggle competition I've taken part in there's noise and some degree of surprises...\nI find a huge problem the fact that kernels can not be shared before and can not be digested after... Maybe too \npesimistic but I'm sure early sharing would lead to other kind of problems.\n\nI personally enjoyed, I had zero domain knowledge before beginning, a \"martian\" seeing people crash following a ball. This, I think, was a huge dissadvantage in this competition. But disciplined process led to some signal in the noise what was a big satisfaction. Also to experience this format was interesting for me.\n\nAbout the format, I think it probably achieves expected result for the organizers but it is much harder for kagglers  that take part on it;  ironically there's a high chance your hard worked kernel will get a fraction of the votes a quality kernel would get in a \"normal\" competition. And no ranking feedback, or points...",
      "votes": null
    },
    {
      "id": "459473",
      "postDate": "01/21/2019 20:12:04",
      "content": "<p>I guess this is the post-competition support group, so I'll share some of my thoughts here as well. </p>\n\n<p>For my part,  I focused on coming up with a unique solution that did not function by reducing punts or punt returns, as such rules seemed too obvious and I feared they might be considered to conflict with the integrity of the game. I figured that even if they did select a couple projects with \"obvious\" rule recommendations, they'd maybe choose one with a more unique solution.</p>\n\n<p>I had the same hunch as Rob when I saw the winners announced: that the judges might have looked at the slides/presentations first and used those to come up with a short list of top candidates. When you think about it, it actually makes a lot of sense to do it that way, because reading through entire kernels takes a lot of time and the slides were the one part of the competition that couldn't be copied by someone else and that will be seen by a wider audience. Still, it feels disappointing to pour so much time into a kernel that may never have even been looked at by the judges. It would be nice if they made the winning presentations public now that the competition is over.</p>\n\n<p>I agree with the others posts here that for future analytics competitions, the details of judging should be made more clear. I'm not sure that peer reviews should be taken into account in judging though, because people would definitely find ways to manipulate the system and plenty of good kernels don't get much attention from other users.</p>\n\n<p>All things considered, I don't regret the time I spent because I learned a lot and tried my best and I don't think trying your best is something to regret even if things don't turn out how you want.</p>",
      "rawMarkdown": "I guess this is the post-competition support group, so I'll share some of my thoughts here as well. \n\nFor my part,  I focused on coming up with a unique solution that did not function by reducing punts or punt returns, as such rules seemed too obvious and I feared they might be considered to conflict with the integrity of the game. I figured that even if they did select a couple projects with \"obvious\" rule recommendations, they'd maybe choose one with a more unique solution.\n\nI had the same hunch as Rob when I saw the winners announced: that the judges might have looked at the slides/presentations first and used those to come up with a short list of top candidates. When you think about it, it actually makes a lot of sense to do it that way, because reading through entire kernels takes a lot of time and the slides were the one part of the competition that couldn't be copied by someone else and that will be seen by a wider audience. Still, it feels disappointing to pour so much time into a kernel that may never have even been looked at by the judges. It would be nice if they made the winning presentations public now that the competition is over.\n\nI agree with the others posts here that for future analytics competitions, the details of judging should be made more clear. I'm not sure that peer reviews should be taken into account in judging though, because people would definitely find ways to manipulate the system and plenty of good kernels don't get much attention from other users.\n\nAll things considered, I don't regret the time I spent because I learned a lot and tried my best and I don't think trying your best is something to regret even if things don't turn out how you want.",
      "votes": null
    },
    {
      "id": "459476",
      "postDate": "01/21/2019 20:35:04",
      "content": "<h3>Reward structure &amp; Lack of collaboration</h3>\n\n<p>I missed the open discussions and sharing that usually happens in other competitions.\nThe total prize pool was more than enough. Maybe too much...</p>\n\n<p>This was not even the first kaggle competition where judgdes decided the winners.\nJust two recent examples: </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/kaggle/kaggle-survey-2018/home\">2018 Kaggle ML &amp; DS Survey Challenge ($28k in prizes)</a></li>\n<li><a href=\"https://www.kaggle.com/kiva/data-science-for-good-kiva-crowdfunding/home\">Data Science for Good: Kiva Crowdfunding ($30k in prizes)</a></li>\n</ul>\n\n<p>Both competitions had the majority of the prize pool awarded by judges but they also had smaller (1000-2000) prizes to increase participation. Weekly awards, overall popularity awards, awards for the most used external datasets helped active collaboration and I think they improved the final results.\nThose competitions had ~1600 comments on the kernels and ~3500 total kernel upvotes each.</p>\n\n<p>Beside the financial rewards there are other ways to motivate people. \nI really liked the honorable mention idea in the other threads.</p>",
      "rawMarkdown": "### Reward structure &amp; Lack of collaboration\n\nI missed the open discussions and sharing that usually happens in other competitions.\nThe total prize pool was more than enough. Maybe too much...\n\nThis was not even the first kaggle competition where judgdes decided the winners.\nJust two recent examples: \n\n* [2018 Kaggle ML &amp; DS Survey Challenge ($28k in prizes)](https://www.kaggle.com/kaggle/kaggle-survey-2018/home)\n* [Data Science for Good: Kiva Crowdfunding ($30k in prizes)](https://www.kaggle.com/kiva/data-science-for-good-kiva-crowdfunding/home)\n\nBoth competitions had the majority of the prize pool awarded by judges but they also had smaller (1000-2000) prizes to increase participation. Weekly awards, overall popularity awards, awards for the most used external datasets helped active collaboration and I think they improved the final results.\nThose competitions had ~1600 comments on the kernels and ~3500 total kernel upvotes each.\n\nBeside the financial rewards there are other ways to motivate people. \nI really liked the honorable mention idea in the other threads.",
      "votes": null
    },
    {
      "id": "459517",
      "postDate": "01/21/2019 22:36:06",
      "content": "<p>I really like the idea of an analytics competition, but I wish it had been more \"kaggle\" like and emphasized analytics more.  Maybe have more focused, smaller kernels that people can vote and discuss.  There were some great visualizations and interactive graphics that I would have liked to have seen rewarded.  Best use of NGS data could have been another area for recognition.  Separating the kernels from the presentation would have been good.  Pick the 4 best kernels and then those 4 people make presentations.  </p>\n\n<p>This competition emphasized story telling.   It was geared towards the management consultant more than a data scientist.  Chris joked at the beginning about using a TI-83, but honestly, a management consultant with excel would crush a data scientist.  Story telling is a very important part of being a data scientist.  Many organizations force their data scientists to report to non-technical business managers, so it's important to be able  to tell persuasive stories.  My personal problem with that org structure is that it rewards well presented ideas over well founded ideas.  Left unchecked, data scientists end up hand waving the data science and become powerpoint engineers.</p>",
      "rawMarkdown": "I really like the idea of an analytics competition, but I wish it had been more \"kaggle\" like and emphasized analytics more.  Maybe have more focused, smaller kernels that people can vote and discuss.  There were some great visualizations and interactive graphics that I would have liked to have seen rewarded.  Best use of NGS data could have been another area for recognition.  Separating the kernels from the presentation would have been good.  Pick the 4 best kernels and then those 4 people make presentations.  \n\nThis competition emphasized story telling.   It was geared towards the management consultant more than a data scientist.  Chris joked at the beginning about using a TI-83, but honestly, a management consultant with excel would crush a data scientist.  Story telling is a very important part of being a data scientist.  Many organizations force their data scientists to report to non-technical business managers, so it's important to be able  to tell persuasive stories.  My personal problem with that org structure is that it rewards well presented ideas over well founded ideas.  Left unchecked, data scientists end up hand waving the data science and become powerpoint engineers.",
      "votes": null
    },
    {
      "id": "459519",
      "postDate": "01/21/2019 22:41:59",
      "content": "<p>2 last thoughts.\n1. Is everyone else waiting for the 2018 concussion data to come out? <br>\n2. Given a mulligan, my rule proposal would be no new punt specific rules needed.  Continue working on making the overall game safer.</p>",
      "rawMarkdown": "2 last thoughts.\n1. Is everyone else waiting for the 2018 concussion data to come out?  \n2. Given a mulligan, my rule proposal would be no new punt specific rules needed.  Continue working on making the overall game safer.",
      "votes": null
    },
    {
      "id": "459583",
      "postDate": "01/22/2019 03:25:56",
      "content": "<p>Did wonder in the rules or evaluation if some preference against or exclusion applied to outside USA teams. Certainly in the Big Data Bowl the requirement to go to Indianapolis to present was stated in a way that if a team could not commit to that then they would not be eligible. Also it seemed to have some recruitment aspects.  It was not stated anywhere I could see, but from the point of view of travel time, expense, etc. maybe they leant toward those entrants that they could tell were already in the States and short flights away.   To some extent this goes to your point on anonymity as well.  Being a first time event for them, maybe they'd rather select a management consultant they can see from their profile that works at , so they can feel like the selections be professionals and have experience doing presentations, etc.   Not like sour grapes, but you could look at it like given a tie in the 4 places, they might have picked US teams over ex-pats, known company profiles over unknowns, gold medallists over novices, etc.     </p>\n\n<p>With respect to Travelling Santa and the team of Helsgaun and Cook - much was made of their participation by a few, but at the end of the day, all their sources were available to everyone and they agreed to let everyone use Concorde, LKH from the start of the competition, though normally only allowed for academic use.  There was a lot of sharing of ideas and kernels, some high placed teams ideas even surprised them, like the penalty schedule.   This goes to your point on the lack of discussion, collaboration here. But also with no leaderboard, no idea who the teams were in this competition, there was no way to gauge the silent participants. Of course even in Travelling Santa a last minute top ten snuck in!       </p>",
      "rawMarkdown": "Did wonder in the rules or evaluation if some preference against or exclusion applied to outside USA teams. Certainly in the Big Data Bowl the requirement to go to Indianapolis to present was stated in a way that if a team could not commit to that then they would not be eligible. Also it seemed to have some recruitment aspects.  It was not stated anywhere I could see, but from the point of view of travel time, expense, etc. maybe they leant toward those entrants that they could tell were already in the States and short flights away.   To some extent this goes to your point on anonymity as well.  Being a first time event for them, maybe they'd rather select a management consultant they can see from their profile that works at",
      "votes": null
    },
    {
      "id": "459590",
      "postDate": "01/22/2019 03:43:04",
      "content": "<p>Rob - if there were an MVP or Walter Peyton award for the competition that would surely be yours. You made great contributions to the discussions, kernels, and keeping up the interest levels in an otherwise pretty quiet competition.  You should feel proud of your achievements and hopefully learning some things in the experience has some rewards now or in the future.</p>\n\n<p>On the flip side - if punt rules change and the public hate them, no trolling will come your way!    </p>",
      "rawMarkdown": "Rob - if there were an MVP or Walter Peyton award for the competition that would surely be yours. You made great contributions to the discussions, kernels, and keeping up the interest levels in an otherwise pretty quiet competition.  You should feel proud of your achievements and hopefully learning some things in the experience has some rewards now or in the future.\n\nOn the flip side - if punt rules change and the public hate them, no trolling will come your way!",
      "votes": null
    },
    {
      "id": "459598",
      "postDate": "01/22/2019 03:57:19",
      "content": "<p>Looking at concussion in another code of football - their points for primary prevention were - \nMinimise head contact - rule changes and enforcement of rules that limit contact\nMinimise impact energy - challenging\nMinimise impact forces - head linear and angular acceleration with helmets, head guards, etc.</p>\n\n<p>Quite  a lot of focus was on skills training, and to a certain extent, paying more attention to punts, tackles  or near collisions, you could observe the experienced, agile almost acrobatic players techniques that seemed to help avoid head contact.  So rule changes are just one aspect and maybe not the only or best.</p>\n\n<p>Thanks for your contributions to the competition Eric! </p>",
      "rawMarkdown": "Looking at concussion in another code of football - their points for primary prevention were - \nMinimise head contact - rule changes and enforcement of rules that limit contact\nMinimise impact energy - challenging\nMinimise impact forces - head linear and angular acceleration with helmets, head guards, etc.\n\nQuite  a lot of focus was on skills training, and to a certain extent, paying more attention to punts, tackles  or near collisions, you could observe the experienced, agile almost acrobatic players techniques that seemed to help avoid head contact.  So rule changes are just one aspect and maybe not the only or best.\n\nThanks for your contributions to the competition Eric!",
      "votes": null
    },
    {
      "id": "460032",
      "postDate": "01/22/2019 21:20:44",
      "content": "<p>Good to see that it's not just me upset with the outcome of the competition.  It's always tough when you pour so much time and effort into something, only to have it subjectively rejected.  Yet, I don't regret the time spent on the competition nor do I \"blame\" the judges for their decision.  They must've had an incredibly hard time judging very similar submissions.  I just wish there was a bit more transparency into how they came to their conclusions.  I spent a lot of time on my slides and was confident in them, so I would love to be able to compare my slides with the winning submissions.  </p>",
      "rawMarkdown": "Good to see that it's not just me upset with the outcome of the competition.  It's always tough when you pour so much time and effort into something, only to have it subjectively rejected.  Yet, I don't regret the time spent on the competition nor do I \"blame\" the judges for their decision.  They must've had an incredibly hard time judging very similar submissions.  I just wish there was a bit more transparency into how they came to their conclusions.  I spent a lot of time on my slides and was confident in them, so I would love to be able to compare my slides with the winning submissions.",
      "votes": null
    },
    {
      "id": "460517",
      "postDate": "01/23/2019 21:45:08",
      "content": "<p>Rob and everyone, thank you for the feedback. I know you've all been really active during this competition and that made it really exciting. I'm certainly taking notes for the future and right now it's <em>blaringly</em> obvious that we need to get some more feedback for everyone.  I'd be happy to respond to feedback if anyone wants (just tag me), otherwise I'll sit back and keep taking notes.  </p>",
      "rawMarkdown": "Rob and everyone, thank you for the feedback. I know you've all been really active during this competition and that made it really exciting. I'm certainly taking notes for the future and right now it's *blaringly* obvious that we need to get some more feedback for everyone.  I'd be happy to respond to feedback if anyone wants (just tag me), otherwise I'll sit back and keep taking notes.",
      "votes": null
    },
    {
      "id": "460593",
      "postDate": "01/24/2019 03:45:09",
      "content": "<p>Very nice post. I did not participate in the competition due to my lack of domain knowledge.</p>\n\n<p>Rob, I have lost many competitions where I have <code>a lot of hard work</code> and as you say <code>not smart work</code>, but learned a lot which helped me in future competitions and otherwise.</p>\n\n<p>You have done yourself good because writing things out is extremely therapeutic. You have also done a major work for the community by writing the <strong>Lessons Learnt</strong>.</p>\n\n<p>Thank you for your great post!</p>",
      "rawMarkdown": "Very nice post. I did not participate in the competition due to my lack of domain knowledge.\n\nRob, I have lost many competitions where I have ` a lot of hard work` and as you say `not smart work`, but learned a lot which helped me in future competitions and otherwise.\n\nYou have done yourself good because writing things out is extremely therapeutic. You have also done a major work for the community by writing the **Lessons Learnt**.\n\nThank you for your great post!",
      "votes": null
    },
    {
      "id": "460851",
      "postDate": "01/24/2019 14:46:27",
      "content": "<p>One problem was that there was no public leaderboard.  Obviously, we couldn't ask someone to manually score every kernel every time that is submitted.  That would be insane.  Two way we could simulate a leaderboard:\n1. Self-Scoring -- post a scoring rubric at the beginning of the competition.  Then we can self-grade as we go.  It wouldn't be perfect, but it would give us an idea if we are hitting all the right points.\n2. Kaggle competition to score analytics competition -- Honestly, not quite sure how this would work.  Like a real data science project, collecting and labeling the data would be very hard.  I just think it would be very cool to use Kaggle to improve Kaggle.  </p>",
      "rawMarkdown": "One problem was that there was no public leaderboard.  Obviously, we couldn't ask someone to manually score every kernel every time that is submitted.  That would be insane.  Two way we could simulate a leaderboard:\n1. Self-Scoring -- post a scoring rubric at the beginning of the competition.  Then we can self-grade as we go.  It wouldn't be perfect, but it would give us an idea if we are hitting all the right points.\n2. Kaggle competition to score analytics competition -- Honestly, not quite sure how this would work.  Like a real data science project, collecting and labeling the data would be very hard.  I just think it would be very cool to use Kaggle to improve Kaggle.",
      "votes": null
    },
    {
      "id": "460859",
      "postDate": "01/24/2019 15:00:14",
      "content": "<p>I also had a thought about the nature of this competition.  In hindsight, this might have been a very good test case for a co-operative solution instead of a competitive solution.  There was a limited solution space, which we saw by the heavy overlap in proposed solutions.  It might have been better to work on the proposal as a crowd.  IIRC, the first kernel was just a wall of text that said \"58% of concussions happen on returns of less than 5 yards.  Therefore, give 5 yards for a fair catch.\"  Hours of analysis later using NGS data, making graphs, doing statistics, I came back to that very same conclusion.  </p>\n\n<p>I think we could all have benefited from early feedback from others.  \"Have you thought of this?\", \"That graph doesn't make sense\",  etc.  Instead of us all parsing the play information, one person could do it and we could all use it.  One person makes a graph, the next person makes it better.   </p>\n\n<p>I don't know how to score that kind of thing, and maybe there's no special scoring, just regular kernel and discussion points.  I've forked a number of the kernels that won or that I liked, and I learned a lot by working through their code.  </p>\n\n<p>Which brings me back to the above post.  Creating and scoring enough kernels to have data to develop a ML scorer would be a good candidate for a co-operative project.  We might even (if we are ambitious) develop a project to create and score kernels.  I've been wanting to play with GAN's and that could be a good use case.</p>",
      "rawMarkdown": "I also had a thought about the nature of this competition.  In hindsight, this might have been a very good test case for a co-operative solution instead of a competitive solution.  There was a limited solution space, which we saw by the heavy overlap in proposed solutions.  It might have been better to work on the proposal as a crowd.  IIRC, the first kernel was just a wall of text that said \"58% of concussions happen on returns of less than 5 yards.  Therefore, give 5 yards for a fair catch.\"  Hours of analysis later using NGS data, making graphs, doing statistics, I came back to that very same conclusion.  \n\nI think we could all have benefited from early feedback from others.  \"Have you thought of this?\", \"That graph doesn't make sense\",  etc.  Instead of us all parsing the play information, one person could do it and we could all use it.  One person makes a graph, the next person makes it better.   \n\nI don't know how to score that kind of thing, and maybe there's no special scoring, just regular kernel and discussion points.  I've forked a number of the kernels that won or that I liked, and I learned a lot by working through their code.  \n\nWhich brings me back to the above post.  Creating and scoring enough kernels to have data to develop a ML scorer would be a good candidate for a co-operative project.  We might even (if we are ambitious) develop a project to create and score kernels.  I've been wanting to play with GAN's and that could be a good use case.",
      "votes": null
    },
    {
      "id": "460943",
      "postDate": "01/24/2019 19:35:08",
      "content": "<p>I made a post <a href=\"https://hamelg.blogspot.com/2019/01/thoughts-on-kaggle-analytics.html\">on my blog</a> with some extended thoughts on the competition if you care to read it and can stand some criticism. It is a little long so I won't post the whole thing here.</p>",
      "rawMarkdown": "I made a post [on my blog](https://hamelg.blogspot.com/2019/01/thoughts-on-kaggle-analytics.html) with some extended thoughts on the competition if you care to read it and can stand some criticism. It is a little long so I won't post the whole thing here.",
      "votes": null
    },
    {
      "id": "460979",
      "postDate": "01/24/2019 23:15:58",
      "content": "<p><a href=\"/ericfreeman\">@ericfreeman</a> You might be interested in the Data Science for Good competitions. We aren't running one at the moment but they're the same kind of format and are <em>maybe, possibly, kind of</em> a little more cooperative.  </p>\n\n<p><a href=\"/hamelg\">@hamelg</a> Nice blog post. Thanks for the feedback and criticism. </p>",
      "rawMarkdown": "ericfreeman You might be interested in the Data Science for Good competitions. We aren't running one at the moment but they're the same kind of format and are *maybe, possibly, kind of* a little more cooperative.  \n\n\n@hamelg Nice blog post. Thanks for the feedback and criticism.",
      "votes": null
    },
    {
      "id": "460984",
      "postDate": "01/24/2019 23:40:22",
      "content": "<p>Hi Chris, can you give us a rough timeframe for when the next Data science for Good competition will start, just out of curiosity?</p>",
      "rawMarkdown": "Hi Chris, can you give us a rough timeframe for when the next Data science for Good competition will start, just out of curiosity?",
      "votes": null
    },
    {
      "id": "460993",
      "postDate": "01/25/2019 00:28:20",
      "content": "<p><a href=\"/garlsham\">@garlsham</a> We don't have a date yet but probably in the next couple of weeks. </p>",
      "rawMarkdown": "garlsham We don't have a date yet but probably in the next couple of weeks.",
      "votes": null
    },
    {
      "id": "460994",
      "postDate": "01/25/2019 00:31:40",
      "content": "<p>Ok thanks.</p>",
      "rawMarkdown": "Ok thanks.",
      "votes": null
    },
    {
      "id": "461024",
      "postDate": "01/25/2019 02:47:55",
      "content": "<p>\" I'd be happy to respond to feedback if anyone wants (just tag me)\"</p>\n\n<p><a href=\"/crawford\">@crawford</a> I would love to hear some feedback, please.</p>",
      "rawMarkdown": "\" I'd be happy to respond to feedback if anyone wants (just tag me)\"\n\n@crawford I would love to hear some feedback, please.",
      "votes": null
    },
    {
      "id": "461399",
      "postDate": "01/25/2019 22:52:13",
      "content": "<p><a href=\"/sp007c\">@sp007c</a> I meant that I would be happy to respond to any comments or complaints about the competition</p>",
      "rawMarkdown": "sp007c I meant that I would be happy to respond to any comments or complaints about the competition",
      "votes": null
    },
    {
      "id": "461431",
      "postDate": "01/26/2019 02:35:54",
      "content": "<p>With respect to public LB, scoring, etc.  There could be a peer review for a given scoring rubric submission file format - e.g. if teams needed to review 2-5 kernels or whatever is reasonable and of course not their own!  Ideally the same points to be used for evaluation.  Then at least it would give some feedback and create a pseudo LB.  </p>\n\n<p>It would probably be good if competition kernels could be made public only to those in the competition until it is over, maybe a generic competition collaborator could be added when ready for reviews rather than made public.   Along with that, perhaps having a timeline that made the kernels finish a week or two prior to presentation pdf/ppt submissions could be good.  Would give some time for reviews, feedback, comments, etc. to refine presentations, and to focus on the different requirements. </p>",
      "rawMarkdown": "With respect to public LB, scoring, etc.  There could be a peer review for a given scoring rubric submission file format - e.g. if teams needed to review 2-5 kernels or whatever is reasonable and of course not their own!  Ideally the same points to be used for evaluation.  Then at least it would give some feedback and create a pseudo LB.  \n\nIt would probably be good if competition kernels could be made public only to those in the competition until it is over, maybe a generic competition collaborator could be added when ready for reviews rather than made public.   Along with that, perhaps having a timeline that made the kernels finish a week or two prior to presentation pdf/ppt submissions could be good.  Would give some time for reviews, feedback, comments, etc. to refine presentations, and to focus on the different requirements.",
      "votes": null
    },
    {
      "id": "462139",
      "postDate": "01/27/2019 19:04:20",
      "content": "<p>2018 concussion data will be highly beneficial, keeping in mind the rule changes. We will be able to realise the impact of the \"lowering head' rule change. Most of the concussion in this data were caused when player used their helmet to tackle someone. When we have the 2018 data , we can see if the rule managed to lower the tackles with helmet.</p>",
      "rawMarkdown": "2018 concussion data will be highly beneficial, keeping in mind the rule changes. We will be able to realise the impact of the \"lowering head' rule change. Most of the concussion in this data were caused when player used their helmet to tackle someone. When we have the 2018 data , we can see if the rule managed to lower the tackles with helmet.",
      "votes": null
    },
    {
      "id": "466722",
      "postDate": "02/05/2019 21:31:48",
      "content": "<p>Dear Rob,\n                   I can understand your emotions. Although many people deserve more than me who actually happened to submit their work unfortunately I missed my submission due to lack of track of where to start and end at the last minute and I missed my Hail Mary pass. Nevertheless I would not consider that I regret my time and efforts as I got an exclusive opportunity to review those NFL datasets, understand internal details of how games and teams function and learn pythons/R platform that gave boost to my career for sure. I would love to mention this in my resume and present like a feather in a cap. I do agree that you have spend enormous time and efforts looking at your kernel overall but I did have a nervous feeling due to loaded details with visualization and lines and curves. The way I looked at it i had feeling that a non-techie would have it like an overwhelming series of bouts. \nT.M.I my friend. T.M.I\nNext time I would see myself in judges shoes and review with someone and very important do some time management for myself which i missed from time to time.\nKaggle wins too !!!!</p>",
      "rawMarkdown": "Dear Rob,\n                   I can understand your emotions. Although many people deserve more than me who actually happened to submit their work unfortunately I missed my submission due to lack of track of where to start and end at the last minute and I missed my Hail Mary pass. Nevertheless I would not consider that I regret my time and efforts as I got an exclusive opportunity to review those NFL datasets, understand internal details of how games and teams function and learn pythons/R platform that gave boost to my career for sure. I would love to mention this in my resume and present like a feather in a cap. I do agree that you have spend enormous time and efforts looking at your kernel overall but I did have a nervous feeling due to loaded details with visualization and lines and curves. The way I looked at it i had feeling that a non-techie would have it like an overwhelming series of bouts. \nT.M.I my friend. T.M.I\nNext time I would see myself in judges shoes and review with someone and very important do some time management for myself which i missed from time to time.\nKaggle wins too !!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 459415,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/21/2019 17:22:55",
      "content": "<h3>Our main goals</h3>\n\n<p>Given that this competition forum is not that active I do not want to start a new thread.\nI agree with most points of Rob. I will simply flood this topic with my additional thoughts.\nI try to add several comments each have the best intent to improve similar further competitions.\nObviously I am not unbiased, as Rob (and sure others as well) we were hoping/expected to go to Atlanta. </p>\n\n<p>We read the <a href=\"https://www.kaggle.com/c/NFL-Punt-Analytics-Competition#evaluation\">Evaluation</a> page very carefully and always kept in mind each and every point.</p>\n\n<p>Our main goals were</p>\n\n<ul>\n<li>Create an excellent quality presentation as many of the judges might not care about the source code</li>\n<li>Explore all the provided data do not leave rocks unturned</li>\n<li>Collect more data to have more solid, statistically significant results</li>\n<li>Expect broad audience and do not go into too much technical details in the slides and summary</li>\n<li>Try to suggest several different rule modifications but highlight the one with the biggest impact</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459418,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/21/2019 17:32:39",
      "content": "<h3>Anonimity</h3>\n\n<p>There was a strange question in the beginning of the competition about bias and discrimination.</p>\n\n<p>&gt; <strong>Chris Crawford wrote</strong>\n&gt; \n&gt; &gt; The answer is yes and no. We do a little to prevent bias. For example, we don't show your username or email address to the hosts when we give them the final list of submissions, but they'll obviously see who you are when they see your kernel. \n&gt; \n&gt; If you're worried about it, I give everyone permission to use my avatar so we all look the same :)</p>\n\n<p>I believe we had a compelling story (everyone likes to cheer to the small team who does not really has a chance, right? :))\nWe decided to put our work into focus and only shared who we are after the results were finalized. \nGiven that we are talking about PR event that was probably a mistake.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459423,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/21/2019 17:48:13",
      "content": "<h3>USA &amp; Domain Knowledge</h3>\n\n<p>We live in Hungary where soccer is way more popular than any other sports (imho way more than it should be). \nEven though my friend is a huge NFL fan and knows the game quite well we had to work hard to understand the fine details of the rules.</p>\n\n<p>Actually we were more afraid of professional sport analysts/ PhD students with relevant research area. </p>\n\n<p>I am not saying this to claim some kind of consolation prize.\nThis is a competition and let the best team win!\nWe certainly had more chance here than would had against Prof. Keld Helsgaun and William Cook in the other \n<a href=\"https://www.kaggle.com/c/traveling-santa-2018-prime-paths/discussion/77134\">Traveling Santa Competition</a> :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 459583,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "01/22/2019 03:25:56",
          "content": "<p>Did wonder in the rules or evaluation if some preference against or exclusion applied to outside USA teams. Certainly in the Big Data Bowl the requirement to go to Indianapolis to present was stated in a way that if a team could not commit to that then they would not be eligible. Also it seemed to have some recruitment aspects.  It was not stated anywhere I could see, but from the point of view of travel time, expense, etc. maybe they leant toward those entrants that they could tell were already in the States and short flights away.   To some extent this goes to your point on anonymity as well.  Being a first time event for them, maybe they'd rather select a management consultant they can see from their profile that works at , so they can feel like the selections be professionals and have experience doing presentations, etc.   Not like sour grapes, but you could look at it like given a tie in the 4 places, they might have picked US teams over ex-pats, known company profiles over unknowns, gold medallists over novices, etc.     </p>\n\n<p>With respect to Travelling Santa and the team of Helsgaun and Cook - much was made of their participation by a few, but at the end of the day, all their sources were available to everyone and they agreed to let everyone use Concorde, LKH from the start of the competition, though normally only allowed for academic use.  There was a lot of sharing of ideas and kernels, some high placed teams ideas even surprised them, like the penalty schedule.   This goes to your point on the lack of discussion, collaboration here. But also with no leaderboard, no idea who the teams were in this competition, there was no way to gauge the silent participants. Of course even in Travelling Santa a last minute top ten snuck in!       </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 459436,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/21/2019 18:21:07",
      "content": "<h3>Transparency of the decision</h3>\n\n<p>In the other thread many others raised good points about how to make the selection process more transparent. I won't repeat them just wanted to emphasize the importance of it. We learned a lot during the competition but we don't have a clue what we missed or what to improve next. </p>\n\n<p>We talked about single and double coverage, it feels that we had an additional <a href=\"https://en.wikipedia.org/wiki/Maximum_coverage_problem\"><strong>maximum coverage problem</strong></a>.\n<img src=\"https://s3-eu-west-1.amazonaws.com/nfl-punt-analytics/MaxCov.png\" alt=\"\"></p>\n\n<p>It is rational from the host. We tried to mitigate the risk by suggesting more rules and having unique methods and suggestions. Unfortunately the best submissions had quite a few overlapping rule sets. I have to agree with Rob from the other thread:</p>\n\n<blockquote>\n  <p><strong>Rob Mulla wrote</strong>\n  I had an idea for my rule change proposals fairly early on and was pretty devisdtated when I saw many others with the same ideas. But in the end, just like you said it’s pretty cool to see people coming to the same conclusions from different angles. </p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459445,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/21/2019 18:42:31",
      "content": "<h3>Ambiguity of requirements</h3>\n\n<p>Even the last minute <a href=\"https://www.kaggle.com/c/NFL-Punt-Analytics-Competition/discussion/76622\">clarification</a> was not clear about what is required from the presentation. 5-10 minutes presentation with 1-50 slides is a bit vague. We aimed to be somewhere in the middle and thanks to my teammate we had an excellent deck with ~24 slides. Maybe it was to long, but we knew we could present the main part easily to any audience who already read our summary report.</p>\n\n<p>Not knowing the audience (e.g. former players, coaches, journalists, doctors, analysts, data scientists etc.) made it very difficult to balance between being too shallow and going too deep.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459451,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "01/21/2019 19:02:33",
      "content": "<p>@RobMulla, let me remind you that you  did a great job. Your risk normalization was a brilliant idea, data augmentation when it was really needed.  Plus you were one of the few who kept alive the almost desert forums.</p>\n\n<p>About votes in this competition, I think there was some distortion because of the competition format. There are many kernels that possibly under \"normal\" circumstances would have had some votes that are almost unnoticed (Funnily, I had not even read the elegant solution of <a href=\"/awainger\">@awainger</a> before winners were announced. 135 kernels are a lot...)</p>\n\n<p>About actual optimal format of presentation... who knows. I agree that it would be good to have more clear guidelines, or even a template,so that efforts can be better directed. But in almost all Kaggle competition I've taken part in there's noise and some degree of surprises...\nI find a huge problem the fact that kernels can not be shared before and can not be digested after... Maybe too \npesimistic but I'm sure early sharing would lead to other kind of problems.</p>\n\n<p>I personally enjoyed, I had zero domain knowledge before beginning, a \"martian\" seeing people crash following a ball. This, I think, was a huge dissadvantage in this competition. But disciplined process led to some signal in the noise what was a big satisfaction. Also to experience this format was interesting for me.</p>\n\n<p>About the format, I think it probably achieves expected result for the organizers but it is much harder for kagglers  that take part on it;  ironically there's a high chance your hard worked kernel will get a fraction of the votes a quality kernel would get in a \"normal\" competition. And no ranking feedback, or points...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459473,
      "author_name": "hamelg",
      "author_url": "",
      "post_date": "01/21/2019 20:12:04",
      "content": "<p>I guess this is the post-competition support group, so I'll share some of my thoughts here as well. </p>\n\n<p>For my part,  I focused on coming up with a unique solution that did not function by reducing punts or punt returns, as such rules seemed too obvious and I feared they might be considered to conflict with the integrity of the game. I figured that even if they did select a couple projects with \"obvious\" rule recommendations, they'd maybe choose one with a more unique solution.</p>\n\n<p>I had the same hunch as Rob when I saw the winners announced: that the judges might have looked at the slides/presentations first and used those to come up with a short list of top candidates. When you think about it, it actually makes a lot of sense to do it that way, because reading through entire kernels takes a lot of time and the slides were the one part of the competition that couldn't be copied by someone else and that will be seen by a wider audience. Still, it feels disappointing to pour so much time into a kernel that may never have even been looked at by the judges. It would be nice if they made the winning presentations public now that the competition is over.</p>\n\n<p>I agree with the others posts here that for future analytics competitions, the details of judging should be made more clear. I'm not sure that peer reviews should be taken into account in judging though, because people would definitely find ways to manipulate the system and plenty of good kernels don't get much attention from other users.</p>\n\n<p>All things considered, I don't regret the time I spent because I learned a lot and tried my best and I don't think trying your best is something to regret even if things don't turn out how you want.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459476,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/21/2019 20:35:04",
      "content": "<h3>Reward structure &amp; Lack of collaboration</h3>\n\n<p>I missed the open discussions and sharing that usually happens in other competitions.\nThe total prize pool was more than enough. Maybe too much...</p>\n\n<p>This was not even the first kaggle competition where judgdes decided the winners.\nJust two recent examples: </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/kaggle/kaggle-survey-2018/home\">2018 Kaggle ML &amp; DS Survey Challenge ($28k in prizes)</a></li>\n<li><a href=\"https://www.kaggle.com/kiva/data-science-for-good-kiva-crowdfunding/home\">Data Science for Good: Kiva Crowdfunding ($30k in prizes)</a></li>\n</ul>\n\n<p>Both competitions had the majority of the prize pool awarded by judges but they also had smaller (1000-2000) prizes to increase participation. Weekly awards, overall popularity awards, awards for the most used external datasets helped active collaboration and I think they improved the final results.\nThose competitions had ~1600 comments on the kernels and ~3500 total kernel upvotes each.</p>\n\n<p>Beside the financial rewards there are other ways to motivate people. \nI really liked the honorable mention idea in the other threads.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459517,
      "author_name": "ericfreeman",
      "author_url": "",
      "post_date": "01/21/2019 22:36:06",
      "content": "<p>I really like the idea of an analytics competition, but I wish it had been more \"kaggle\" like and emphasized analytics more.  Maybe have more focused, smaller kernels that people can vote and discuss.  There were some great visualizations and interactive graphics that I would have liked to have seen rewarded.  Best use of NGS data could have been another area for recognition.  Separating the kernels from the presentation would have been good.  Pick the 4 best kernels and then those 4 people make presentations.  </p>\n\n<p>This competition emphasized story telling.   It was geared towards the management consultant more than a data scientist.  Chris joked at the beginning about using a TI-83, but honestly, a management consultant with excel would crush a data scientist.  Story telling is a very important part of being a data scientist.  Many organizations force their data scientists to report to non-technical business managers, so it's important to be able  to tell persuasive stories.  My personal problem with that org structure is that it rewards well presented ideas over well founded ideas.  Left unchecked, data scientists end up hand waving the data science and become powerpoint engineers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459519,
      "author_name": "ericfreeman",
      "author_url": "",
      "post_date": "01/21/2019 22:41:59",
      "content": "<p>2 last thoughts.\n1. Is everyone else waiting for the 2018 concussion data to come out? <br>\n2. Given a mulligan, my rule proposal would be no new punt specific rules needed.  Continue working on making the overall game safer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 459598,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "01/22/2019 03:57:19",
          "content": "<p>Looking at concussion in another code of football - their points for primary prevention were - \nMinimise head contact - rule changes and enforcement of rules that limit contact\nMinimise impact energy - challenging\nMinimise impact forces - head linear and angular acceleration with helmets, head guards, etc.</p>\n\n<p>Quite  a lot of focus was on skills training, and to a certain extent, paying more attention to punts, tackles  or near collisions, you could observe the experienced, agile almost acrobatic players techniques that seemed to help avoid head contact.  So rule changes are just one aspect and maybe not the only or best.</p>\n\n<p>Thanks for your contributions to the competition Eric! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462139,
          "author_name": "someadityamandal",
          "author_url": "",
          "post_date": "01/27/2019 19:04:20",
          "content": "<p>2018 concussion data will be highly beneficial, keeping in mind the rule changes. We will be able to realise the impact of the \"lowering head' rule change. Most of the concussion in this data were caused when player used their helmet to tackle someone. When we have the 2018 data , we can see if the rule managed to lower the tackles with helmet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 459590,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "01/22/2019 03:43:04",
      "content": "<p>Rob - if there were an MVP or Walter Peyton award for the competition that would surely be yours. You made great contributions to the discussions, kernels, and keeping up the interest levels in an otherwise pretty quiet competition.  You should feel proud of your achievements and hopefully learning some things in the experience has some rewards now or in the future.</p>\n\n<p>On the flip side - if punt rules change and the public hate them, no trolling will come your way!    </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460032,
      "author_name": "mtodisco10",
      "author_url": "",
      "post_date": "01/22/2019 21:20:44",
      "content": "<p>Good to see that it's not just me upset with the outcome of the competition.  It's always tough when you pour so much time and effort into something, only to have it subjectively rejected.  Yet, I don't regret the time spent on the competition nor do I \"blame\" the judges for their decision.  They must've had an incredibly hard time judging very similar submissions.  I just wish there was a bit more transparency into how they came to their conclusions.  I spent a lot of time on my slides and was confident in them, so I would love to be able to compare my slides with the winning submissions.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460517,
      "author_name": "crawford",
      "author_url": "",
      "post_date": "01/23/2019 21:45:08",
      "content": "<p>Rob and everyone, thank you for the feedback. I know you've all been really active during this competition and that made it really exciting. I'm certainly taking notes for the future and right now it's <em>blaringly</em> obvious that we need to get some more feedback for everyone.  I'd be happy to respond to feedback if anyone wants (just tag me), otherwise I'll sit back and keep taking notes.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 460851,
          "author_name": "ericfreeman",
          "author_url": "",
          "post_date": "01/24/2019 14:46:27",
          "content": "<p>One problem was that there was no public leaderboard.  Obviously, we couldn't ask someone to manually score every kernel every time that is submitted.  That would be insane.  Two way we could simulate a leaderboard:\n1. Self-Scoring -- post a scoring rubric at the beginning of the competition.  Then we can self-grade as we go.  It wouldn't be perfect, but it would give us an idea if we are hitting all the right points.\n2. Kaggle competition to score analytics competition -- Honestly, not quite sure how this would work.  Like a real data science project, collecting and labeling the data would be very hard.  I just think it would be very cool to use Kaggle to improve Kaggle.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460859,
          "author_name": "ericfreeman",
          "author_url": "",
          "post_date": "01/24/2019 15:00:14",
          "content": "<p>I also had a thought about the nature of this competition.  In hindsight, this might have been a very good test case for a co-operative solution instead of a competitive solution.  There was a limited solution space, which we saw by the heavy overlap in proposed solutions.  It might have been better to work on the proposal as a crowd.  IIRC, the first kernel was just a wall of text that said \"58% of concussions happen on returns of less than 5 yards.  Therefore, give 5 yards for a fair catch.\"  Hours of analysis later using NGS data, making graphs, doing statistics, I came back to that very same conclusion.  </p>\n\n<p>I think we could all have benefited from early feedback from others.  \"Have you thought of this?\", \"That graph doesn't make sense\",  etc.  Instead of us all parsing the play information, one person could do it and we could all use it.  One person makes a graph, the next person makes it better.   </p>\n\n<p>I don't know how to score that kind of thing, and maybe there's no special scoring, just regular kernel and discussion points.  I've forked a number of the kernels that won or that I liked, and I learned a lot by working through their code.  </p>\n\n<p>Which brings me back to the above post.  Creating and scoring enough kernels to have data to develop a ML scorer would be a good candidate for a co-operative project.  We might even (if we are ambitious) develop a project to create and score kernels.  I've been wanting to play with GAN's and that could be a good use case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460943,
          "author_name": "hamelg",
          "author_url": "",
          "post_date": "01/24/2019 19:35:08",
          "content": "<p>I made a post <a href=\"https://hamelg.blogspot.com/2019/01/thoughts-on-kaggle-analytics.html\">on my blog</a> with some extended thoughts on the competition if you care to read it and can stand some criticism. It is a little long so I won't post the whole thing here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460979,
          "author_name": "crawford",
          "author_url": "",
          "post_date": "01/24/2019 23:15:58",
          "content": "<p><a href=\"/ericfreeman\">@ericfreeman</a> You might be interested in the Data Science for Good competitions. We aren't running one at the moment but they're the same kind of format and are <em>maybe, possibly, kind of</em> a little more cooperative.  </p>\n\n<p><a href=\"/hamelg\">@hamelg</a> Nice blog post. Thanks for the feedback and criticism. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460984,
          "author_name": "garlsham",
          "author_url": "",
          "post_date": "01/24/2019 23:40:22",
          "content": "<p>Hi Chris, can you give us a rough timeframe for when the next Data science for Good competition will start, just out of curiosity?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460993,
          "author_name": "crawford",
          "author_url": "",
          "post_date": "01/25/2019 00:28:20",
          "content": "<p><a href=\"/garlsham\">@garlsham</a> We don't have a date yet but probably in the next couple of weeks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460994,
          "author_name": "garlsham",
          "author_url": "",
          "post_date": "01/25/2019 00:31:40",
          "content": "<p>Ok thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461024,
          "author_name": "sp007c",
          "author_url": "",
          "post_date": "01/25/2019 02:47:55",
          "content": "<p>\" I'd be happy to respond to feedback if anyone wants (just tag me)\"</p>\n\n<p><a href=\"/crawford\">@crawford</a> I would love to hear some feedback, please.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461399,
          "author_name": "crawford",
          "author_url": "",
          "post_date": "01/25/2019 22:52:13",
          "content": "<p><a href=\"/sp007c\">@sp007c</a> I meant that I would be happy to respond to any comments or complaints about the competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461431,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "01/26/2019 02:35:54",
          "content": "<p>With respect to public LB, scoring, etc.  There could be a peer review for a given scoring rubric submission file format - e.g. if teams needed to review 2-5 kernels or whatever is reasonable and of course not their own!  Ideally the same points to be used for evaluation.  Then at least it would give some feedback and create a pseudo LB.  </p>\n\n<p>It would probably be good if competition kernels could be made public only to those in the competition until it is over, maybe a generic competition collaborator could be added when ready for reviews rather than made public.   Along with that, perhaps having a timeline that made the kernels finish a week or two prior to presentation pdf/ppt submissions could be good.  Would give some time for reviews, feedback, comments, etc. to refine presentations, and to focus on the different requirements. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 460593,
      "author_name": "ambarish",
      "author_url": "",
      "post_date": "01/24/2019 03:45:09",
      "content": "<p>Very nice post. I did not participate in the competition due to my lack of domain knowledge.</p>\n\n<p>Rob, I have lost many competitions where I have <code>a lot of hard work</code> and as you say <code>not smart work</code>, but learned a lot which helped me in future competitions and otherwise.</p>\n\n<p>You have done yourself good because writing things out is extremely therapeutic. You have also done a major work for the community by writing the <strong>Lessons Learnt</strong>.</p>\n\n<p>Thank you for your great post!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 466722,
      "author_name": "rajshivraj",
      "author_url": "",
      "post_date": "02/05/2019 21:31:48",
      "content": "<p>Dear Rob,\n                   I can understand your emotions. Although many people deserve more than me who actually happened to submit their work unfortunately I missed my submission due to lack of track of where to start and end at the last minute and I missed my Hail Mary pass. Nevertheless I would not consider that I regret my time and efforts as I got an exclusive opportunity to review those NFL datasets, understand internal details of how games and teams function and learn pythons/R platform that gave boost to my career for sure. I would love to mention this in my resume and present like a feather in a cap. I do agree that you have spend enormous time and efforts looking at your kernel overall but I did have a nervous feeling due to loaded details with visualization and lines and curves. The way I looked at it i had feeling that a non-techie would have it like an overwhelming series of bouts. \nT.M.I my friend. T.M.I\nNext time I would see myself in judges shoes and review with someone and very important do some time management for myself which i missed from time to time.\nKaggle wins too !!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "459312": "I’m very thankful for Kaggle and the NFL for hosting this competition, it was a lot of fun and I learned a lot. The winners clearly did an amazing job and should be applauded for their hard work. As with most competitions, there are more losers than winners- and I'm sure many of us wish we were chosen. Regardless I do feel like it’s important to look back and grow from the experience. It’s been a few days since the winners were announced and I’ve had a chance to collect my thoughts and put them on paper. This may be more of a therapeutic way for me to finally close the book on this competition than anything :)\n\nIf I could do it all over again:\n\n- I would have focused more on the slides and presentation. Most of my work was spent digging into every angle of the data. Trying to understand and visualize the NGS data, and generally gaining an understanding the results and formations of punting plays. I wanted to make sure I investigated any possible area that could reduce the possibility of concussions. I assumed that the slides would come naturally as an extension of this analysis. Now that it’s over my gut feeling is that the judges used the slides as a starting point to evaluate submissions and dove into the kernels to confirm that it supported their conclusions. If I had known this going in I would’ve worked backwards from a polished power-point and then supported it with my kernel.\n- I would have made my presentation much, much shorter and to the point. I had assumed that some of non-technical judges might only look at the slides slides only- so I tried to include a lot of my analysis in my slides to show all the word I had done. The unintended consequence of this was that it was way too much to digest in a 5-10 minute presentation. At the end of the day this is for a PR event where the NFL wanted to hear quick pitches about proposals backed by data- not digest an entire research report.\n- Similar to my last point I would’ve made my analysis less exhaustive, more focused, and concise. I added a lot of analysis to my kernel that was more exploratory and didn’t impact my final rule proposals directly- I don’t think adding this extra analysis helped to keep things clear. I noticed that ‘concise’ was specifically mentioned when describing the winning submissions. Again, this makes sense with the whole 5-10 minute presentation thing.\n- I would've focused more on telling a cohesive story. The winning submissions did an excellent job of this in their kernels and I'm sure also in their slides. This is always the hardest part.\n- I would have tried to think more outside the box with my rule suggestions to differentiate it from others. It does appear that the winners each had their own angle on the rule proposal and the judges selected them in a way that there wasn’t a lot of overlap. As frustrating as it is to very similar proposals being selected, but the judges obviously thought they did a better job reaching their conclusions in a more unique, and clearer way than I did. I would have never thought of suggesting something like putting sensors in helmets as a rule proposal- it was very clever and was rewarded for being so.\n- I would have worked smarter, not harder. Especially when it came to the NGS data. Plotting and reviewing player routes was a big part of my approach, but in the end wasn’t digestible for a 5 minute presentation- so it didn’t matter. The winning submissions focused less on the NGS data and instead on supporting their rule changes in a way that anyone could understand. My approach to quantify each play’s risk level was not the type of thing I was going to be able to pitch to a general audience. I will say, after going back and reviewing @awainger submission I was blown away by the clean and clear way he used the NGS data. Really outstanding work and I learned a lot.\n- I would've made my preprocessing all run within a kernel. By the time I found out this was a requirement I had already gone too far into preprocessing offline. It probably wasn't a deciding factor in the end, but it's something I regret not doing it and can't help but notice all of the winning submissions ran everything in kernels.\n- Would've shared my EDA kernel publicly early on. I took the approach that many others did of not wanting my analysis to get out, but I think I could've benefited from receiving feedback earlier. Share more, share early, everyone benefits.\n- Probably the biggest for me: I would’ve not believed the hype. No matter how many people tell you you’re a ‘shoe-in’ for something, it doesn’t mean anything until the results are in. I appreciated all the kind comments I got about my kernel, but I also knew that it was setting myself for disappointment- and made missing the cut hurt even more. I guess that’s just a general life lesson.\n\nSuggestions and feedback to Kaggle for future Analytics competitions:\n\n- This competition had a fairly quick turnaround (which I really liked). I appreciate the responsiveness and feedback we did receive, but since things were moving relatively fast, I think the message board needed quick responses to questions from participants. This is especially true with the post detailing what content should be in the slides/kernel which that came out about a week before the deadline. If we had known this information up front I think many of us would’ve approached things differently.\n- Be clear about how the submissions are being judged and who will be reviewing them. Knowing the audience turned out to be really important. The evaluation section of the competition said the submissions would be judged on Solution Efficacy and Game Integrity- but it didn’t explain who was deciding what is or isn’t an effective solution. The audience was loosely clarified in a discussion post but should have been stated up front. Ideally this type of competition would say “We have a panel of X judges with positions A, B, and C who will each vote on their top 4 and the submissions”. This would also help with transparency and make it feel like submissions weren’t cherry picked by a company to support their preconceived beliefs (I’m not saying that’s the case here, but transparency couldn’t hurt). It also might help the participants who didn't win at least see if they were close to making the cut by seeing how many votes they received.\n- Take into account peer reviews. I think Kaggle should seriously consider this in future competitions. If a large portion of the participants seem to appreciate a submission, it should bear some weight in the review process. If it doesn't then either we are being told: (1) many of the participants didn't actually understand what the objective of the competition actually was or (2) The judges know better than those who had spent time working on the competition. I could see there being a valid argument for #2 - but isn't that the whole point of crowdsourcing for ideas? I know Kaggle comments/upvotes can be easily doctored or manipulated- so that can be a concern. Still, there must be some be clever way to include peer reviews (even if only as a superlative) in future competitions while minimizing the possibility for manipulation.\n\n....and by the length of this post it's clear I haven't learned how to be brief and concise yet :D\n\nSo in the end, knowing what I do now- Do I regret having spending all the time I did working on this competition? Yes.  If the NFL and Kaggle announced another competition today on a different topic would I participate? Absolutely yes- but I’d take an entirely different approach.\n\nBest of luck to all the winners in Atlanta!",
    "459415": "### Our main goals\nGiven that this competition forum is not that active I do not want to start a new thread.\nI agree with most points of Rob. I will simply flood this topic with my additional thoughts.\nI try to add several comments each have the best intent to improve similar further competitions.\nObviously I am not unbiased, as Rob (and sure others as well) we were hoping/expected to go to Atlanta. \n \nWe read the [Evaluation](https://www.kaggle.com/c/NFL-Punt-Analytics-Competition#evaluation) page very carefully and always kept in mind each and every point.\n\nOur main goals were\n\n* Create an excellent quality presentation as many of the judges might not care about the source code\n* Explore all the provided data do not leave rocks unturned\n* Collect more data to have more solid, statistically significant results\n* Expect broad audience and do not go into too much technical details in the slides and summary\n* Try to suggest several different rule modifications but highlight the one with the biggest impact",
    "459418": "### Anonimity \n\nThere was a strange question in the beginning of the competition about bias and discrimination.\n\n&gt; **Chris Crawford wrote**\n&gt; \n&gt; &gt; The answer is yes and no. We do a little to prevent bias. For example, we don't show your username or email address to the hosts when we give them the final list of submissions, but they'll obviously see who you are when they see your kernel. \n&gt; \n&gt; If you're worried about it, I give everyone permission to use my avatar so we all look the same :)\n\nI believe we had a compelling story (everyone likes to cheer to the small team who does not really has a chance, right? :))\nWe decided to put our work into focus and only shared who we are after the results were finalized. \nGiven that we are talking about PR event that was probably a mistake.",
    "459423": "### USA &amp; Domain Knowledge\n\nWe live in Hungary where soccer is way more popular than any other sports (imho way more than it should be). \nEven though my friend is a huge NFL fan and knows the game quite well we had to work hard to understand the fine details of the rules.\n\nActually we were more afraid of professional sport analysts/ PhD students with relevant research area. \n\nI am not saying this to claim some kind of consolation prize.\nThis is a competition and let the best team win!\nWe certainly had more chance here than would had against Prof. Keld Helsgaun and William Cook in the other \n[Traveling Santa Competition](https://www.kaggle.com/c/traveling-santa-2018-prime-paths/discussion/77134) :)",
    "459436": "### Transparency of the decision\nIn the other thread many others raised good points about how to make the selection process more transparent. I won't repeat them just wanted to emphasize the importance of it. We learned a lot during the competition but we don't have a clue what we missed or what to improve next. \n\nWe talked about single and double coverage, it feels that we had an additional [**maximum coverage problem**](https://en.wikipedia.org/wiki/Maximum_coverage_problem).\n![](https://s3-eu-west-1.amazonaws.com/nfl-punt-analytics/MaxCov.png)\n\nIt is rational from the host. We tried to mitigate the risk by suggesting more rules and having unique methods and suggestions. Unfortunately the best submissions had quite a few overlapping rule sets. I have to agree with Rob from the other thread:\n\n&gt; **Rob Mulla wrote**\n&gt; I had an idea for my rule change proposals fairly early on and was pretty devisdtated when I saw many others with the same ideas. But in the end, just like you said it’s pretty cool to see people coming to the same conclusions from different angles.",
    "459445": "### Ambiguity of requirements\n\nEven the last minute [clarification](https://www.kaggle.com/c/NFL-Punt-Analytics-Competition/discussion/76622) was not clear about what is required from the presentation. 5-10 minutes presentation with 1-50 slides is a bit vague. We aimed to be somewhere in the middle and thanks to my teammate we had an excellent deck with ~24 slides. Maybe it was to long, but we knew we could present the main part easily to any audience who already read our summary report.\n\nNot knowing the audience (e.g. former players, coaches, journalists, doctors, analysts, data scientists etc.) made it very difficult to balance between being too shallow and going too deep.",
    "459451": "RobMulla, let me remind you that you  did a great job. Your risk normalization was a brilliant idea, data augmentation when it was really needed.  Plus you were one of the few who kept alive the almost desert forums.\n\nAbout votes in this competition, I think there was some distortion because of the competition format. There are many kernels that possibly under \"normal\" circumstances would have had some votes that are almost unnoticed (Funnily, I had not even read the elegant solution of @awainger before winners were announced. 135 kernels are a lot...)\n\nAbout actual optimal format of presentation... who knows. I agree that it would be good to have more clear guidelines, or even a template,so that efforts can be better directed. But in almost all Kaggle competition I've taken part in there's noise and some degree of surprises...\nI find a huge problem the fact that kernels can not be shared before and can not be digested after... Maybe too \npesimistic but I'm sure early sharing would lead to other kind of problems.\n\nI personally enjoyed, I had zero domain knowledge before beginning, a \"martian\" seeing people crash following a ball. This, I think, was a huge dissadvantage in this competition. But disciplined process led to some signal in the noise what was a big satisfaction. Also to experience this format was interesting for me.\n\nAbout the format, I think it probably achieves expected result for the organizers but it is much harder for kagglers  that take part on it;  ironically there's a high chance your hard worked kernel will get a fraction of the votes a quality kernel would get in a \"normal\" competition. And no ranking feedback, or points...",
    "459473": "I guess this is the post-competition support group, so I'll share some of my thoughts here as well. \n\nFor my part,  I focused on coming up with a unique solution that did not function by reducing punts or punt returns, as such rules seemed too obvious and I feared they might be considered to conflict with the integrity of the game. I figured that even if they did select a couple projects with \"obvious\" rule recommendations, they'd maybe choose one with a more unique solution.\n\nI had the same hunch as Rob when I saw the winners announced: that the judges might have looked at the slides/presentations first and used those to come up with a short list of top candidates. When you think about it, it actually makes a lot of sense to do it that way, because reading through entire kernels takes a lot of time and the slides were the one part of the competition that couldn't be copied by someone else and that will be seen by a wider audience. Still, it feels disappointing to pour so much time into a kernel that may never have even been looked at by the judges. It would be nice if they made the winning presentations public now that the competition is over.\n\nI agree with the others posts here that for future analytics competitions, the details of judging should be made more clear. I'm not sure that peer reviews should be taken into account in judging though, because people would definitely find ways to manipulate the system and plenty of good kernels don't get much attention from other users.\n\nAll things considered, I don't regret the time I spent because I learned a lot and tried my best and I don't think trying your best is something to regret even if things don't turn out how you want.",
    "459476": "### Reward structure &amp; Lack of collaboration\n\nI missed the open discussions and sharing that usually happens in other competitions.\nThe total prize pool was more than enough. Maybe too much...\n\nThis was not even the first kaggle competition where judgdes decided the winners.\nJust two recent examples: \n\n* [2018 Kaggle ML &amp; DS Survey Challenge ($28k in prizes)](https://www.kaggle.com/kaggle/kaggle-survey-2018/home)\n* [Data Science for Good: Kiva Crowdfunding ($30k in prizes)](https://www.kaggle.com/kiva/data-science-for-good-kiva-crowdfunding/home)\n\nBoth competitions had the majority of the prize pool awarded by judges but they also had smaller (1000-2000) prizes to increase participation. Weekly awards, overall popularity awards, awards for the most used external datasets helped active collaboration and I think they improved the final results.\nThose competitions had ~1600 comments on the kernels and ~3500 total kernel upvotes each.\n\nBeside the financial rewards there are other ways to motivate people. \nI really liked the honorable mention idea in the other threads.",
    "459517": "I really like the idea of an analytics competition, but I wish it had been more \"kaggle\" like and emphasized analytics more.  Maybe have more focused, smaller kernels that people can vote and discuss.  There were some great visualizations and interactive graphics that I would have liked to have seen rewarded.  Best use of NGS data could have been another area for recognition.  Separating the kernels from the presentation would have been good.  Pick the 4 best kernels and then those 4 people make presentations.  \n\nThis competition emphasized story telling.   It was geared towards the management consultant more than a data scientist.  Chris joked at the beginning about using a TI-83, but honestly, a management consultant with excel would crush a data scientist.  Story telling is a very important part of being a data scientist.  Many organizations force their data scientists to report to non-technical business managers, so it's important to be able  to tell persuasive stories.  My personal problem with that org structure is that it rewards well presented ideas over well founded ideas.  Left unchecked, data scientists end up hand waving the data science and become powerpoint engineers.",
    "459519": "2 last thoughts.\n1. Is everyone else waiting for the 2018 concussion data to come out?  \n2. Given a mulligan, my rule proposal would be no new punt specific rules needed.  Continue working on making the overall game safer.",
    "459583": "Did wonder in the rules or evaluation if some preference against or exclusion applied to outside USA teams. Certainly in the Big Data Bowl the requirement to go to Indianapolis to present was stated in a way that if a team could not commit to that then they would not be eligible. Also it seemed to have some recruitment aspects.  It was not stated anywhere I could see, but from the point of view of travel time, expense, etc. maybe they leant toward those entrants that they could tell were already in the States and short flights away.   To some extent this goes to your point on anonymity as well.  Being a first time event for them, maybe they'd rather select a management consultant they can see from their profile that works at",
    "459590": "Rob - if there were an MVP or Walter Peyton award for the competition that would surely be yours. You made great contributions to the discussions, kernels, and keeping up the interest levels in an otherwise pretty quiet competition.  You should feel proud of your achievements and hopefully learning some things in the experience has some rewards now or in the future.\n\nOn the flip side - if punt rules change and the public hate them, no trolling will come your way!",
    "459598": "Looking at concussion in another code of football - their points for primary prevention were - \nMinimise head contact - rule changes and enforcement of rules that limit contact\nMinimise impact energy - challenging\nMinimise impact forces - head linear and angular acceleration with helmets, head guards, etc.\n\nQuite  a lot of focus was on skills training, and to a certain extent, paying more attention to punts, tackles  or near collisions, you could observe the experienced, agile almost acrobatic players techniques that seemed to help avoid head contact.  So rule changes are just one aspect and maybe not the only or best.\n\nThanks for your contributions to the competition Eric!",
    "460032": "Good to see that it's not just me upset with the outcome of the competition.  It's always tough when you pour so much time and effort into something, only to have it subjectively rejected.  Yet, I don't regret the time spent on the competition nor do I \"blame\" the judges for their decision.  They must've had an incredibly hard time judging very similar submissions.  I just wish there was a bit more transparency into how they came to their conclusions.  I spent a lot of time on my slides and was confident in them, so I would love to be able to compare my slides with the winning submissions.",
    "460517": "Rob and everyone, thank you for the feedback. I know you've all been really active during this competition and that made it really exciting. I'm certainly taking notes for the future and right now it's *blaringly* obvious that we need to get some more feedback for everyone.  I'd be happy to respond to feedback if anyone wants (just tag me), otherwise I'll sit back and keep taking notes.",
    "460593": "Very nice post. I did not participate in the competition due to my lack of domain knowledge.\n\nRob, I have lost many competitions where I have ` a lot of hard work` and as you say `not smart work`, but learned a lot which helped me in future competitions and otherwise.\n\nYou have done yourself good because writing things out is extremely therapeutic. You have also done a major work for the community by writing the **Lessons Learnt**.\n\nThank you for your great post!",
    "460851": "One problem was that there was no public leaderboard.  Obviously, we couldn't ask someone to manually score every kernel every time that is submitted.  That would be insane.  Two way we could simulate a leaderboard:\n1. Self-Scoring -- post a scoring rubric at the beginning of the competition.  Then we can self-grade as we go.  It wouldn't be perfect, but it would give us an idea if we are hitting all the right points.\n2. Kaggle competition to score analytics competition -- Honestly, not quite sure how this would work.  Like a real data science project, collecting and labeling the data would be very hard.  I just think it would be very cool to use Kaggle to improve Kaggle.",
    "460859": "I also had a thought about the nature of this competition.  In hindsight, this might have been a very good test case for a co-operative solution instead of a competitive solution.  There was a limited solution space, which we saw by the heavy overlap in proposed solutions.  It might have been better to work on the proposal as a crowd.  IIRC, the first kernel was just a wall of text that said \"58% of concussions happen on returns of less than 5 yards.  Therefore, give 5 yards for a fair catch.\"  Hours of analysis later using NGS data, making graphs, doing statistics, I came back to that very same conclusion.  \n\nI think we could all have benefited from early feedback from others.  \"Have you thought of this?\", \"That graph doesn't make sense\",  etc.  Instead of us all parsing the play information, one person could do it and we could all use it.  One person makes a graph, the next person makes it better.   \n\nI don't know how to score that kind of thing, and maybe there's no special scoring, just regular kernel and discussion points.  I've forked a number of the kernels that won or that I liked, and I learned a lot by working through their code.  \n\nWhich brings me back to the above post.  Creating and scoring enough kernels to have data to develop a ML scorer would be a good candidate for a co-operative project.  We might even (if we are ambitious) develop a project to create and score kernels.  I've been wanting to play with GAN's and that could be a good use case.",
    "460943": "I made a post [on my blog](https://hamelg.blogspot.com/2019/01/thoughts-on-kaggle-analytics.html) with some extended thoughts on the competition if you care to read it and can stand some criticism. It is a little long so I won't post the whole thing here.",
    "460979": "ericfreeman You might be interested in the Data Science for Good competitions. We aren't running one at the moment but they're the same kind of format and are *maybe, possibly, kind of* a little more cooperative.  \n\n\n@hamelg Nice blog post. Thanks for the feedback and criticism.",
    "460984": "Hi Chris, can you give us a rough timeframe for when the next Data science for Good competition will start, just out of curiosity?",
    "460993": "garlsham We don't have a date yet but probably in the next couple of weeks.",
    "460994": "Ok thanks.",
    "461024": "\" I'd be happy to respond to feedback if anyone wants (just tag me)\"\n\n@crawford I would love to hear some feedback, please.",
    "461399": "sp007c I meant that I would be happy to respond to any comments or complaints about the competition",
    "461431": "With respect to public LB, scoring, etc.  There could be a peer review for a given scoring rubric submission file format - e.g. if teams needed to review 2-5 kernels or whatever is reasonable and of course not their own!  Ideally the same points to be used for evaluation.  Then at least it would give some feedback and create a pseudo LB.  \n\nIt would probably be good if competition kernels could be made public only to those in the competition until it is over, maybe a generic competition collaborator could be added when ready for reviews rather than made public.   Along with that, perhaps having a timeline that made the kernels finish a week or two prior to presentation pdf/ppt submissions could be good.  Would give some time for reviews, feedback, comments, etc. to refine presentations, and to focus on the different requirements.",
    "462139": "2018 concussion data will be highly beneficial, keeping in mind the rule changes. We will be able to realise the impact of the \"lowering head' rule change. Most of the concussion in this data were caused when player used their helmet to tackle someone. When we have the 2018 data , we can see if the rule managed to lower the tackles with helmet.",
    "466722": "Dear Rob,\n                   I can understand your emotions. Although many people deserve more than me who actually happened to submit their work unfortunately I missed my submission due to lack of track of where to start and end at the last minute and I missed my Hail Mary pass. Nevertheless I would not consider that I regret my time and efforts as I got an exclusive opportunity to review those NFL datasets, understand internal details of how games and teams function and learn pythons/R platform that gave boost to my career for sure. I would love to mention this in my resume and present like a feather in a cap. I do agree that you have spend enormous time and efforts looking at your kernel overall but I did have a nervous feeling due to loaded details with visualization and lines and curves. The way I looked at it i had feeling that a non-techie would have it like an overwhelming series of bouts. \nT.M.I my friend. T.M.I\nNext time I would see myself in judges shoes and review with someone and very important do some time management for myself which i missed from time to time.\nKaggle wins too !!!!"
  },
  "source": "meta"
}