{
  "id": 372382,
  "title": "Extreme Overfitting with tracking data",
  "url": "/competitions/nfl-player-contact-detection/discussion/372382",
  "author_name": "",
  "post_date": "2022-12-15T18:00:00.195762900Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm seeing extreme reductions in matthews correlation between the train/val sets during training and the test set used for public score. I've narrowed down the performance issue to the tracking files, and get much better generalization exclusively using the helmet files. Here are some notebooks that help show what I'm running into. Extra pair of eyes would be great:</p>\n<ol>\n<li><p>Baseline where I define predicted contact as any player1/player2 pair within 1 yard of each other. Using the tracking x,y coordinates to calculate the distance of players gives me a validation matthews correlation around 0.5, but then a terrible public score. Utilizing the helmet data to calculate distance gets me a much better public score (0.238), and much more aligned with the matthews correlation found on the training data: <a href=\"https://www.kaggle.com/code/ryancaldwell/within-x-yards-distance-as-contact-baseline/notebook\" target=\"_blank\">https://www.kaggle.com/code/ryancaldwell/within-x-yards-distance-as-contact-baseline/notebook</a></p></li>\n<li><p>Xgboost model utilizing tracking data for player1/player2 distances and step_pct as features. This model shows a matthews correlation around 0.55 on the validation fold during training, but then scores 0.04 on the public leaderboard: <a href=\"https://www.kaggle.com/code/ryancaldwell/xgboost-model?scriptVersionId=113927847\" target=\"_blank\">https://www.kaggle.com/code/ryancaldwell/xgboost-model?scriptVersionId=113927847</a></p></li>\n<li><p>What's interesting is that I created a separate model with only features from the helmet file, which gave a leaderboard score of 0.311 (version 1), and then it dropped way down to 0.035 (version 2) when adding in the player distance and x,y positions from the tracking file: <a href=\"https://www.kaggle.com/code/ryancaldwell/helmet-bounding-box-iou\" target=\"_blank\">https://www.kaggle.com/code/ryancaldwell/helmet-bounding-box-iou</a></p></li>\n</ol>\n<p>So it appears maybe I'm doing something wrong when creating the distance features when joining up the test tracking file with submission. That's where I think it could be going wrong.</p>",
  "messages": [
    {
      "id": "2066451",
      "postDate": "12/15/2022 18:00:00",
      "content": "<p>I'm seeing extreme reductions in matthews correlation between the train/val sets during training and the test set used for public score. I've narrowed down the performance issue to the tracking files, and get much better generalization exclusively using the helmet files. Here are some notebooks that help show what I'm running into. Extra pair of eyes would be great:</p>\n<ol>\n<li><p>Baseline where I define predicted contact as any player1/player2 pair within 1 yard of each other. Using the tracking x,y coordinates to calculate the distance of players gives me a validation matthews correlation around 0.5, but then a terrible public score. Utilizing the helmet data to calculate distance gets me a much better public score (0.238), and much more aligned with the matthews correlation found on the training data: <a href=\"https://www.kaggle.com/code/ryancaldwell/within-x-yards-distance-as-contact-baseline/notebook\" target=\"_blank\">https://www.kaggle.com/code/ryancaldwell/within-x-yards-distance-as-contact-baseline/notebook</a></p></li>\n<li><p>Xgboost model utilizing tracking data for player1/player2 distances and step_pct as features. This model shows a matthews correlation around 0.55 on the validation fold during training, but then scores 0.04 on the public leaderboard: <a href=\"https://www.kaggle.com/code/ryancaldwell/xgboost-model?scriptVersionId=113927847\" target=\"_blank\">https://www.kaggle.com/code/ryancaldwell/xgboost-model?scriptVersionId=113927847</a></p></li>\n<li><p>What's interesting is that I created a separate model with only features from the helmet file, which gave a leaderboard score of 0.311 (version 1), and then it dropped way down to 0.035 (version 2) when adding in the player distance and x,y positions from the tracking file: <a href=\"https://www.kaggle.com/code/ryancaldwell/helmet-bounding-box-iou\" target=\"_blank\">https://www.kaggle.com/code/ryancaldwell/helmet-bounding-box-iou</a></p></li>\n</ol>\n<p>So it appears maybe I'm doing something wrong when creating the distance features when joining up the test tracking file with submission. That's where I think it could be going wrong.</p>",
      "rawMarkdown": "I'm seeing extreme reductions in matthews correlation between the train/val sets during training and the test set used for public score. I've narrowed down the performance issue to the tracking files, and get much better generalization exclusively using the helmet files. Here are some notebooks that help show what I'm running into. Extra pair of eyes would be great:\n\n1. Baseline where I define predicted contact as any player1/player2 pair within 1 yard of each other. Using the tracking x,y coordinates to calculate the distance of players gives me a validation matthews correlation around 0.5, but then a terrible public score. Utilizing the helmet data to calculate distance gets me a much better public score (0.238), and much more aligned with the matthews correlation found on the training data: https://www.kaggle.com/code/ryancaldwell/within-x-yards-distance-as-contact-baseline/notebook\n\n2. Xgboost model utilizing tracking data for player1/player2 distances and step_pct as features. This model shows a matthews correlation around 0.55 on the validation fold during training, but then scores 0.04 on the public leaderboard: https://www.kaggle.com/code/ryancaldwell/xgboost-model?scriptVersionId=113927847\n\n3. What's interesting is that I created a separate model with only features from the helmet file, which gave a leaderboard score of 0.311 (version 1), and then it dropped way down to 0.035 (version 2) when adding in the player distance and x,y positions from the tracking file: https://www.kaggle.com/code/ryancaldwell/helmet-bounding-box-iou\n\nSo it appears maybe I'm doing something wrong when creating the distance features when joining up the test tracking file with submission. That's where I think it could be going wrong.",
      "votes": null
    },
    {
      "id": "2068648",
      "postDate": "12/18/2022 07:40:23",
      "content": "<p>Sometimes messing the order of things screw submissions, be careful with that.</p>",
      "rawMarkdown": "Sometimes messing the order of things screw submissions, be careful with that.",
      "votes": null
    },
    {
      "id": "2068826",
      "postDate": "12/18/2022 10:46:25",
      "content": "<p>What is step_pct?</p>",
      "rawMarkdown": "What is step_pct?",
      "votes": null
    },
    {
      "id": "2068955",
      "postDate": "12/18/2022 12:55:36",
      "content": "<p>It just represents the percentage of the way through a given play is. So, if a play has 10 steps and is currently on step 1, then its step_pct would be 10%. It works as a key with the helmet file, but I don't think it does with the tracking data.</p>",
      "rawMarkdown": "It just represents the percentage of the way through a given play is. So, if a play has 10 steps and is currently on step 1, then its step_pct would be 10%. It works as a key with the helmet file, but I don't think it does with the tracking data.",
      "votes": null
    },
    {
      "id": "2068961",
      "postDate": "12/18/2022 13:00:35",
      "content": "<p>Yeah that makes sense, but I'm not sorting anything. I'm just joining on common keys where the left side of the join is the submission file. I'll do some digging on this, but it doesn't seem like the order of the predictions or submission rows should be changing.</p>",
      "rawMarkdown": "Yeah that makes sense, but I'm not sorting anything. I'm just joining on common keys where the left side of the join is the submission file. I'll do some digging on this, but it doesn't seem like the order of the predictions or submission rows should be changing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2068648,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "12/18/2022 07:40:23",
      "content": "<p>Sometimes messing the order of things screw submissions, be careful with that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2068961,
          "author_name": "ryancaldwell",
          "author_url": "",
          "post_date": "12/18/2022 13:00:35",
          "content": "<p>Yeah that makes sense, but I'm not sorting anything. I'm just joining on common keys where the left side of the join is the submission file. I'll do some digging on this, but it doesn't seem like the order of the predictions or submission rows should be changing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2068826,
      "author_name": "louisbunuel",
      "author_url": "",
      "post_date": "12/18/2022 10:46:25",
      "content": "<p>What is step_pct?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2068955,
          "author_name": "ryancaldwell",
          "author_url": "",
          "post_date": "12/18/2022 12:55:36",
          "content": "<p>It just represents the percentage of the way through a given play is. So, if a play has 10 steps and is currently on step 1, then its step_pct would be 10%. It works as a key with the helmet file, but I don't think it does with the tracking data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2066451": "I'm seeing extreme reductions in matthews correlation between the train/val sets during training and the test set used for public score. I've narrowed down the performance issue to the tracking files, and get much better generalization exclusively using the helmet files. Here are some notebooks that help show what I'm running into. Extra pair of eyes would be great:\n\n1. Baseline where I define predicted contact as any player1/player2 pair within 1 yard of each other. Using the tracking x,y coordinates to calculate the distance of players gives me a validation matthews correlation around 0.5, but then a terrible public score. Utilizing the helmet data to calculate distance gets me a much better public score (0.238), and much more aligned with the matthews correlation found on the training data: https://www.kaggle.com/code/ryancaldwell/within-x-yards-distance-as-contact-baseline/notebook\n\n2. Xgboost model utilizing tracking data for player1/player2 distances and step_pct as features. This model shows a matthews correlation around 0.55 on the validation fold during training, but then scores 0.04 on the public leaderboard: https://www.kaggle.com/code/ryancaldwell/xgboost-model?scriptVersionId=113927847\n\n3. What's interesting is that I created a separate model with only features from the helmet file, which gave a leaderboard score of 0.311 (version 1), and then it dropped way down to 0.035 (version 2) when adding in the player distance and x,y positions from the tracking file: https://www.kaggle.com/code/ryancaldwell/helmet-bounding-box-iou\n\nSo it appears maybe I'm doing something wrong when creating the distance features when joining up the test tracking file with submission. That's where I think it could be going wrong.",
    "2068648": "Sometimes messing the order of things screw submissions, be careful with that.",
    "2068826": "What is step_pct?",
    "2068955": "It just represents the percentage of the way through a given play is. So, if a play has 10 steps and is currently on step 1, then its step_pct would be 10%. It works as a key with the helmet file, but I don't think it does with the tracking data.",
    "2068961": "Yeah that makes sense, but I'm not sorting anything. I'm just joining on common keys where the left side of the join is the submission file. I'll do some digging on this, but it doesn't seem like the order of the predictions or submission rows should be changing."
  },
  "source": "meta"
}