{
  "id": 76766,
  "title": "Beware of analysis overfitting ",
  "url": "/competitions/NFL-Punt-Analytics-Competition/discussion/76766",
  "author_name": "",
  "post_date": "2019-01-06T12:59:29.397718200Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Frequent Kagglers have already realized this competition is extremely easy to overfit. Maybe this heads up will help some newer ones. </p>\n\n<p>Looks like there are some sports+data enthusiasts jumping on their first challenge here, what is great.  Granted, good insights not always require super high coding skills or a PhD in statistis. But <strong>with this kind of dataset potentially high dimensional, NGS included, and with very few \"events\" -concussive plays- either you take measures not to overfit or you overfit</strong>. </p>\n\n<p>In this case there's no public or private leaderboard but analytics can also be overfit just as usual by finding non existent patterns that look good in sample data.</p>\n\n<p>That's dangerous, if undetected -worst case scenario- actions based on those patterns would misdirect time and efforts with no effect on concussions.</p>\n\n<p>What to do? Depends on your resources but extensive cross validation of models, simulation, proper statistical tests -if you are good at statistics- or all of them are effective measures to take.</p>\n\n<p>Tl, dr: <strong>No matter how good a patern looks, make sure to cross validate it</strong>.</p>",
  "messages": [
    {
      "id": "451128",
      "postDate": "01/06/2019 12:59:29",
      "content": "<p>Frequent Kagglers have already realized this competition is extremely easy to overfit. Maybe this heads up will help some newer ones. </p>\n\n<p>Looks like there are some sports+data enthusiasts jumping on their first challenge here, what is great.  Granted, good insights not always require super high coding skills or a PhD in statistis. But <strong>with this kind of dataset potentially high dimensional, NGS included, and with very few \"events\" -concussive plays- either you take measures not to overfit or you overfit</strong>. </p>\n\n<p>In this case there's no public or private leaderboard but analytics can also be overfit just as usual by finding non existent patterns that look good in sample data.</p>\n\n<p>That's dangerous, if undetected -worst case scenario- actions based on those patterns would misdirect time and efforts with no effect on concussions.</p>\n\n<p>What to do? Depends on your resources but extensive cross validation of models, simulation, proper statistical tests -if you are good at statistics- or all of them are effective measures to take.</p>\n\n<p>Tl, dr: <strong>No matter how good a patern looks, make sure to cross validate it</strong>.</p>",
      "rawMarkdown": "Frequent Kagglers have already realized this competition is extremely easy to overfit. Maybe this heads up will help some newer ones. \n\nLooks like there are some sports+data enthusiasts jumping on their first challenge here, what is great.  Granted, good insights not always require super high coding skills or a PhD in statistis. But **with this kind of dataset potentially high dimensional, NGS included, and with very few \"events\" -concussive plays- either you take measures not to overfit or you overfit**. \n\nIn this case there's no public or private leaderboard but analytics can also be overfit just as usual by finding non existent patterns that look good in sample data.\n\nThat's dangerous, if undetected -worst case scenario- actions based on those patterns would misdirect time and efforts with no effect on concussions.\n\nWhat to do? Depends on your resources but extensive cross validation of models, simulation, proper statistical tests -if you are good at statistics- or all of them are effective measures to take.\n\nTl, dr: **No matter how good a patern looks, make sure to cross validate it**.",
      "votes": null
    },
    {
      "id": "451433",
      "postDate": "01/07/2019 05:16:05",
      "content": "<p>Great points. I'd go as far as saying that even with cross validation - that fitting a model purely based on the 37 out of 6600 + plays to make causal claims will be suspect. I think that's where fully understanding the problem, framing it in the right context, and thinking outside of the box is what makes this competition so interesting and different than the purely ML types.</p>",
      "rawMarkdown": "Great points. I'd go as far as saying that even with cross validation - that fitting a model purely based on the 37 out of 6600 + plays to make causal claims will be suspect. I think that's where fully understanding the problem, framing it in the right context, and thinking outside of the box is what makes this competition so interesting and different than the purely ML types.",
      "votes": null
    },
    {
      "id": "451475",
      "postDate": "01/07/2019 07:02:57",
      "content": "<p>The 37 control events is really low sample CV/ statistical tests are good but I am afraid they won't help to avoid overfitting. Especially with selection bias. </p>",
      "rawMarkdown": "The 37 control events is really low sample CV/ statistical tests are good but I am afraid they won't help to avoid overfitting. Especially with selection bias.",
      "votes": null
    },
    {
      "id": "451497",
      "postDate": "01/07/2019 07:55:02",
      "content": "<p>@RobMulla, <a href=\"/beluga\">@beluga</a>, I agree with your both comments. And, of course, no matter how much resampling 37 cases is not enough for almost any degree of certainty... </p>\n\n<p>My point is more that awareness of the problem can help to:</p>\n\n<ul>\n<li>discard 90% of apparent patterns and,</li>\n<li>more importantly, assess how highly uncertain the other 10% of the patterns are</li>\n</ul>\n\n<p>So not so much about completely avoiding overfitting as about \"data science humbleness\" in this case :-)</p>",
      "rawMarkdown": "RobMulla, @beluga, I agree with your both comments. And, of course, no matter how much resampling 37 cases is not enough for almost any degree of certainty... \n\nMy point is more that awareness of the problem can help to:\n\n- discard 90% of apparent patterns and,\n- more importantly, assess how highly uncertain the other 10% of the patterns are\n\nSo not so much about completely avoiding overfitting as about \"data science humbleness\" in this case :-)",
      "votes": null
    },
    {
      "id": "451568",
      "postDate": "01/07/2019 09:27:09",
      "content": "<p>I am pretty sure with optimum regularization the overfitting can be avoided. However to what extent this idea of generalization goes, depends solely on the task.</p>",
      "rawMarkdown": "I am pretty sure with optimum regularization the overfitting can be avoided. However to what extent this idea of generalization goes, depends solely on the task.",
      "votes": null
    },
    {
      "id": "451595",
      "postDate": "01/07/2019 10:27:04",
      "content": "<p>@ShahbazKhan, realize that what we have here is a very specific situation, a \"rare event\" dataset with very low number of events. Regularization by itself is, in my opinion, of limited usefulness. </p>\n\n<p>Anyway, great to see a bit of discussion begin in this -up to this date- almost desert forum!</p>",
      "rawMarkdown": "ShahbazKhan, realize that what we have here is a very specific situation, a \"rare event\" dataset with very low number of events. Regularization by itself is, in my opinion, of limited usefulness. \n\nAnyway, great to see a bit of discussion begin in this -up to this date- almost desert forum!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 451433,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "01/07/2019 05:16:05",
      "content": "<p>Great points. I'd go as far as saying that even with cross validation - that fitting a model purely based on the 37 out of 6600 + plays to make causal claims will be suspect. I think that's where fully understanding the problem, framing it in the right context, and thinking outside of the box is what makes this competition so interesting and different than the purely ML types.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 451475,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "01/07/2019 07:02:57",
      "content": "<p>The 37 control events is really low sample CV/ statistical tests are good but I am afraid they won't help to avoid overfitting. Especially with selection bias. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 451497,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "01/07/2019 07:55:02",
      "content": "<p>@RobMulla, <a href=\"/beluga\">@beluga</a>, I agree with your both comments. And, of course, no matter how much resampling 37 cases is not enough for almost any degree of certainty... </p>\n\n<p>My point is more that awareness of the problem can help to:</p>\n\n<ul>\n<li>discard 90% of apparent patterns and,</li>\n<li>more importantly, assess how highly uncertain the other 10% of the patterns are</li>\n</ul>\n\n<p>So not so much about completely avoiding overfitting as about \"data science humbleness\" in this case :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 451568,
      "author_name": "thehawkgriffith",
      "author_url": "",
      "post_date": "01/07/2019 09:27:09",
      "content": "<p>I am pretty sure with optimum regularization the overfitting can be avoided. However to what extent this idea of generalization goes, depends solely on the task.</p>",
      "votes": null,
      "replies": [
        {
          "id": 451595,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "01/07/2019 10:27:04",
          "content": "<p>@ShahbazKhan, realize that what we have here is a very specific situation, a \"rare event\" dataset with very low number of events. Regularization by itself is, in my opinion, of limited usefulness. </p>\n\n<p>Anyway, great to see a bit of discussion begin in this -up to this date- almost desert forum!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "451128": "Frequent Kagglers have already realized this competition is extremely easy to overfit. Maybe this heads up will help some newer ones. \n\nLooks like there are some sports+data enthusiasts jumping on their first challenge here, what is great.  Granted, good insights not always require super high coding skills or a PhD in statistis. But **with this kind of dataset potentially high dimensional, NGS included, and with very few \"events\" -concussive plays- either you take measures not to overfit or you overfit**. \n\nIn this case there's no public or private leaderboard but analytics can also be overfit just as usual by finding non existent patterns that look good in sample data.\n\nThat's dangerous, if undetected -worst case scenario- actions based on those patterns would misdirect time and efforts with no effect on concussions.\n\nWhat to do? Depends on your resources but extensive cross validation of models, simulation, proper statistical tests -if you are good at statistics- or all of them are effective measures to take.\n\nTl, dr: **No matter how good a patern looks, make sure to cross validate it**.",
    "451433": "Great points. I'd go as far as saying that even with cross validation - that fitting a model purely based on the 37 out of 6600 + plays to make causal claims will be suspect. I think that's where fully understanding the problem, framing it in the right context, and thinking outside of the box is what makes this competition so interesting and different than the purely ML types.",
    "451475": "The 37 control events is really low sample CV/ statistical tests are good but I am afraid they won't help to avoid overfitting. Especially with selection bias.",
    "451497": "RobMulla, @beluga, I agree with your both comments. And, of course, no matter how much resampling 37 cases is not enough for almost any degree of certainty... \n\nMy point is more that awareness of the problem can help to:\n\n- discard 90% of apparent patterns and,\n- more importantly, assess how highly uncertain the other 10% of the patterns are\n\nSo not so much about completely avoiding overfitting as about \"data science humbleness\" in this case :-)",
    "451568": "I am pretty sure with optimum regularization the overfitting can be avoided. However to what extent this idea of generalization goes, depends solely on the task.",
    "451595": "ShahbazKhan, realize that what we have here is a very specific situation, a \"rare event\" dataset with very low number of events. Regularization by itself is, in my opinion, of limited usefulness. \n\nAnyway, great to see a bit of discussion begin in this -up to this date- almost desert forum!"
  },
  "source": "meta"
}