{
  "id": 391723,
  "title": "6th Place Solution (Qishen & Bo)",
  "url": "/competitions/nfl-player-contact-detection/discussion/391723",
  "author_name": "",
  "post_date": "2023-03-02T13:05:04.162136400Z",
  "votes": 31,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to the organizers and congrats to all the winners and my wonderful teammates <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> </p>\n<p>It is an interesting competition with very rich data types, we can train the model on video, or on table data, or a combination of both.</p>\n<p>So I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.</p>\n<p>This was also confirmed after we merge the team. When we teamed up with another team ( <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> ) we found that they and I surprisingly used a completely different training pipeline. After spending some time we found that our two pipelines even can hardly borrow ideas from each other. So in the end our choice was to optimize the pipelines individually and finally do a simple average ensemble.</p>\n<p><a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> part have already been post here: <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391620\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391620</a></p>\n<p>Next I'm going to describe our part of solution.</p>\n<h1>Summary</h1>\n<p>We designed a 2-stage pipeline. </p>\n<ul>\n<li>Stage 1: Single player sequence 2.5D CNN w/ LSTM classification model.</li>\n<li>Stage 2: Player pairs classification model (2.5D CNN and GBDT).</li>\n</ul>\n<h1>Stage 1: Single player CNN w/ LSTM</h1>\n<p>For a single player in one play, crop his boxes from all frames for both 2 views, using helmet bb extend height and width to 5x</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F39f624f6ffab585d65adb434ddb604df%2FScreen%20Shot%202023-03-02%20at%2010.04.04%20PM.png?generation=1677762260999335&amp;alt=media\" alt=\"\"></p>\n<p>This is an example where the top 6 imgs are Endzone and the bottom 6 imgs are Sideline, of the same player.</p>\n<p>From left to right, each img is one step (6 frames exactly, we don't do 59.94hz…).</p>\n<p>For a single crop, it crop through 5 frames so it's 2.5D CNN</p>\n<p>For training 2.5 CNN w/ LSTM, we use seq_len=128.</p>\n<p>Which means if a video is more than 128 steps, the rest are dropped. (around 0.2% pos samples  in training data are dropped here.)</p>\n<p>The CNN w/ LSTM model take 2 view sequences as input at the same time.</p>\n<p>We concat 2 view pairs features after backbone, then feed into LSTM.</p>\n<p>The output of this model is:</p>\n<ul>\n<li><code>prob_p</code> : whether the player contact with anyone</li>\n<li><code>prob_g</code> : whether the player on the ground </li>\n</ul>\n<p>After stage 1 models (5folds) training done, we predict OOF for all training data. Here we get contact with ground prediction which directly used for submission.</p>\n<p>And contact with people prediction for stage 2.</p>\n<h1>Stage 2: Player pairs classification</h1>\n<p>We selected the samples with <code>prob_p &gt; 0.145</code>, and matched the pairs of players whose distance between players was less than <code>1.6</code> as stage2 training data.</p>\n<p>For a single pair, we crop through 9 frames like this ↓</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F8d760c85635cac713e19729e5852495b%2FScreen%20Shot%202023-03-02%20at%2010.10.38%20PM.png?generation=1677762664204997&amp;alt=media\" alt=\"\"></p>\n<p>Take two helmet bboxes and find the center point of them, expand height and width then crop out.</p>\n<p>We also draw the two bboxes of the player pair on the picture, so that the model knows which player pair is needed when there are many people in the images.</p>\n<p>In this way we are able to train 2.5D CNN to obtain the contact prediction between player pairs extracted from stage 1.</p>\n<p>As for GBDT, we simple extract some features from <code>tracking.csv</code>, nothing special there.</p>\n<p>The final CV score of our part is <code>0.796</code></p>\n<p>After ensemble with <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> we got CV <code>0.804</code></p>\n<h1>Acknowledge</h1>\n<p>I trained many models in this competition using Z by HP Z8G4 Workstation with dual A6000 GPU as usual, it's stable and never had a problem with the training process. Thanks to Z by HP for sponsoring!</p>",
  "messages": [
    {
      "id": "2165784",
      "postDate": "03/02/2023 13:05:04",
      "content": "<p>Thanks to the organizers and congrats to all the winners and my wonderful teammates <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> </p>\n<p>It is an interesting competition with very rich data types, we can train the model on video, or on table data, or a combination of both.</p>\n<p>So I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.</p>\n<p>This was also confirmed after we merge the team. When we teamed up with another team ( <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> ) we found that they and I surprisingly used a completely different training pipeline. After spending some time we found that our two pipelines even can hardly borrow ideas from each other. So in the end our choice was to optimize the pipelines individually and finally do a simple average ensemble.</p>\n<p><a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> part have already been post here: <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391620\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391620</a></p>\n<p>Next I'm going to describe our part of solution.</p>\n<h1>Summary</h1>\n<p>We designed a 2-stage pipeline. </p>\n<ul>\n<li>Stage 1: Single player sequence 2.5D CNN w/ LSTM classification model.</li>\n<li>Stage 2: Player pairs classification model (2.5D CNN and GBDT).</li>\n</ul>\n<h1>Stage 1: Single player CNN w/ LSTM</h1>\n<p>For a single player in one play, crop his boxes from all frames for both 2 views, using helmet bb extend height and width to 5x</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F39f624f6ffab585d65adb434ddb604df%2FScreen%20Shot%202023-03-02%20at%2010.04.04%20PM.png?generation=1677762260999335&amp;alt=media\" alt=\"\"></p>\n<p>This is an example where the top 6 imgs are Endzone and the bottom 6 imgs are Sideline, of the same player.</p>\n<p>From left to right, each img is one step (6 frames exactly, we don't do 59.94hz…).</p>\n<p>For a single crop, it crop through 5 frames so it's 2.5D CNN</p>\n<p>For training 2.5 CNN w/ LSTM, we use seq_len=128.</p>\n<p>Which means if a video is more than 128 steps, the rest are dropped. (around 0.2% pos samples  in training data are dropped here.)</p>\n<p>The CNN w/ LSTM model take 2 view sequences as input at the same time.</p>\n<p>We concat 2 view pairs features after backbone, then feed into LSTM.</p>\n<p>The output of this model is:</p>\n<ul>\n<li><code>prob_p</code> : whether the player contact with anyone</li>\n<li><code>prob_g</code> : whether the player on the ground </li>\n</ul>\n<p>After stage 1 models (5folds) training done, we predict OOF for all training data. Here we get contact with ground prediction which directly used for submission.</p>\n<p>And contact with people prediction for stage 2.</p>\n<h1>Stage 2: Player pairs classification</h1>\n<p>We selected the samples with <code>prob_p &gt; 0.145</code>, and matched the pairs of players whose distance between players was less than <code>1.6</code> as stage2 training data.</p>\n<p>For a single pair, we crop through 9 frames like this ↓</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F8d760c85635cac713e19729e5852495b%2FScreen%20Shot%202023-03-02%20at%2010.10.38%20PM.png?generation=1677762664204997&amp;alt=media\" alt=\"\"></p>\n<p>Take two helmet bboxes and find the center point of them, expand height and width then crop out.</p>\n<p>We also draw the two bboxes of the player pair on the picture, so that the model knows which player pair is needed when there are many people in the images.</p>\n<p>In this way we are able to train 2.5D CNN to obtain the contact prediction between player pairs extracted from stage 1.</p>\n<p>As for GBDT, we simple extract some features from <code>tracking.csv</code>, nothing special there.</p>\n<p>The final CV score of our part is <code>0.796</code></p>\n<p>After ensemble with <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> we got CV <code>0.804</code></p>\n<h1>Acknowledge</h1>\n<p>I trained many models in this competition using Z by HP Z8G4 Workstation with dual A6000 GPU as usual, it's stable and never had a problem with the training process. Thanks to Z by HP for sponsoring!</p>",
      "rawMarkdown": "Thanks to the organizers and congrats to all the winners and my wonderful teammates @boliu0 @tanakar @ryotayoshinobu \n\nIt is an interesting competition with very rich data types, we can train the model on video, or on table data, or a combination of both.\n\nSo I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.\n\nThis was also confirmed after we merge the team. When we teamed up with another team ( @tanakar @ryotayoshinobu ) we found that they and I surprisingly used a completely different training pipeline. After spending some time we found that our two pipelines even can hardly borrow ideas from each other. So in the end our choice was to optimize the pipelines individually and finally do a simple average ensemble.\n\n@tanakar @ryotayoshinobu part have already been post here: https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391620\n\nNext I'm going to describe our part of solution.\n\n# Summary\n\nWe designed a 2-stage pipeline. \n\n* Stage 1: Single player sequence 2.5D CNN w/ LSTM classification model.\n* Stage 2: Player pairs classification model (2.5D CNN and GBDT).\n\n# Stage 1: Single player CNN w/ LSTM\n\nFor a single player in one play, crop his boxes from all frames for both 2 views, using helmet bb extend height and width to 5x\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F39f624f6ffab585d65adb434ddb604df%2FScreen%20Shot%202023-03-02%20at%2010.04.04%20PM.png?generation=1677762260999335&alt=media)\n\n\nThis is an example where the top 6 imgs are Endzone and the bottom 6 imgs are Sideline, of the same player.\n\nFrom left to right, each img is one step (6 frames exactly, we don't do 59.94hz...).\n\nFor a single crop, it crop through 5 frames so it's 2.5D CNN\n\nFor training 2.5 CNN w/ LSTM, we use seq_len=128.\n\nWhich means if a video is more than 128 steps, the rest are dropped. (around 0.2% pos samples  in training data are dropped here.)\n\nThe CNN w/ LSTM model take 2 view sequences as input at the same time.\n\nWe concat 2 view pairs features after backbone, then feed into LSTM.\n\nThe output of this model is:\n\n* `prob_p` : whether the player contact with anyone\n* `prob_g` : whether the player on the ground \n \n\nAfter stage 1 models (5folds) training done, we predict OOF for all training data. Here we get contact with ground prediction which directly used for submission.\n\nAnd contact with people prediction for stage 2.\n\n# Stage 2: Player pairs classification\n\nWe selected the samples with `prob_p > 0.145`, and matched the pairs of players whose distance between players was less than `1.6` as stage2 training data.\n\nFor a single pair, we crop through 9 frames like this ↓\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F8d760c85635cac713e19729e5852495b%2FScreen%20Shot%202023-03-02%20at%2010.10.38%20PM.png?generation=1677762664204997&alt=media)\n\nTake two helmet bboxes and find the center point of them, expand height and width then crop out.\n\nWe also draw the two bboxes of the player pair on the picture, so that the model knows which player pair is needed when there are many people in the images.\n\nIn this way we are able to train 2.5D CNN to obtain the contact prediction between player pairs extracted from stage 1.\n\nAs for GBDT, we simple extract some features from `tracking.csv`, nothing special there.\n\nThe final CV score of our part is `0.796`\n\nAfter ensemble with @tanakar @ryotayoshinobu we got CV `0.804`\n\n\n\n\n# Acknowledge\n\nI trained many models in this competition using Z by HP Z8G4 Workstation with dual A6000 GPU as usual, it's stable and never had a problem with the training process. Thanks to Z by HP for sponsoring!",
      "votes": null
    },
    {
      "id": "2168955",
      "postDate": "03/04/2023 16:56:25",
      "content": "<p>Thanks for posting your solution. I 100% agree about what you said \"So I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.\"</p>\n<p>Congrats on your finish!</p>",
      "rawMarkdown": "Thanks for posting your solution. I 100% agree about what you said \"So I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.\"\n\nCongrats on your finish!",
      "votes": null
    },
    {
      "id": "2173649",
      "postDate": "03/08/2023 14:35:17",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <br>\nCongrats on your gold medal!<br>\nOne question, is the shape of the input to the 1st stage model, for example, as follows?</p>\n<pre><code>(batch_size, 2, 128, 5, h, w)\n</code></pre>\n<p>I mean,</p>\n<pre><code>2: 2views  \n128: 128steps  \n5: 5frames(+-2frames?)\n</code></pre>\n<p>And did you input (bs*2*128, 5, h, w) into the backbone?</p>",
      "rawMarkdown": "haqishen \nCongrats on your gold medal!\nOne question, is the shape of the input to the 1st stage model, for example, as follows?\n```\n(batch_size, 2, 128, 5, h, w)\n```\nI mean,\n```\n2: 2views  \n128: 128steps  \n5: 5frames(+-2frames?)\n```\nAnd did you input (bs\\*2\\*128, 5, h, w) into the backbone?",
      "votes": null
    },
    {
      "id": "2174517",
      "postDate": "03/09/2023 07:22:38",
      "content": "<p>yes that's correct!</p>",
      "rawMarkdown": "yes that's correct!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168955,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:56:25",
      "content": "<p>Thanks for posting your solution. I 100% agree about what you said \"So I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.\"</p>\n<p>Congrats on your finish!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2173649,
      "author_name": "yujiariyasu",
      "author_url": "",
      "post_date": "03/08/2023 14:35:17",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <br>\nCongrats on your gold medal!<br>\nOne question, is the shape of the input to the 1st stage model, for example, as follows?</p>\n<pre><code>(batch_size, 2, 128, 5, h, w)\n</code></pre>\n<p>I mean,</p>\n<pre><code>2: 2views  \n128: 128steps  \n5: 5frames(+-2frames?)\n</code></pre>\n<p>And did you input (bs*2*128, 5, h, w) into the backbone?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174517,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/09/2023 07:22:38",
          "content": "<p>yes that's correct!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2165784": "Thanks to the organizers and congrats to all the winners and my wonderful teammates @boliu0 @tanakar @ryotayoshinobu \n\nIt is an interesting competition with very rich data types, we can train the model on video, or on table data, or a combination of both.\n\nSo I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.\n\nThis was also confirmed after we merge the team. When we teamed up with another team ( @tanakar @ryotayoshinobu ) we found that they and I surprisingly used a completely different training pipeline. After spending some time we found that our two pipelines even can hardly borrow ideas from each other. So in the end our choice was to optimize the pipelines individually and finally do a simple average ensemble.\n\n@tanakar @ryotayoshinobu part have already been post here: https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391620\n\nNext I'm going to describe our part of solution.\n\n# Summary\n\nWe designed a 2-stage pipeline. \n\n* Stage 1: Single player sequence 2.5D CNN w/ LSTM classification model.\n* Stage 2: Player pairs classification model (2.5D CNN and GBDT).\n\n# Stage 1: Single player CNN w/ LSTM\n\nFor a single player in one play, crop his boxes from all frames for both 2 views, using helmet bb extend height and width to 5x\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F39f624f6ffab585d65adb434ddb604df%2FScreen%20Shot%202023-03-02%20at%2010.04.04%20PM.png?generation=1677762260999335&alt=media)\n\n\nThis is an example where the top 6 imgs are Endzone and the bottom 6 imgs are Sideline, of the same player.\n\nFrom left to right, each img is one step (6 frames exactly, we don't do 59.94hz...).\n\nFor a single crop, it crop through 5 frames so it's 2.5D CNN\n\nFor training 2.5 CNN w/ LSTM, we use seq_len=128.\n\nWhich means if a video is more than 128 steps, the rest are dropped. (around 0.2% pos samples  in training data are dropped here.)\n\nThe CNN w/ LSTM model take 2 view sequences as input at the same time.\n\nWe concat 2 view pairs features after backbone, then feed into LSTM.\n\nThe output of this model is:\n\n* `prob_p` : whether the player contact with anyone\n* `prob_g` : whether the player on the ground \n \n\nAfter stage 1 models (5folds) training done, we predict OOF for all training data. Here we get contact with ground prediction which directly used for submission.\n\nAnd contact with people prediction for stage 2.\n\n# Stage 2: Player pairs classification\n\nWe selected the samples with `prob_p > 0.145`, and matched the pairs of players whose distance between players was less than `1.6` as stage2 training data.\n\nFor a single pair, we crop through 9 frames like this ↓\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2F8d760c85635cac713e19729e5852495b%2FScreen%20Shot%202023-03-02%20at%2010.10.38%20PM.png?generation=1677762664204997&alt=media)\n\nTake two helmet bboxes and find the center point of them, expand height and width then crop out.\n\nWe also draw the two bboxes of the player pair on the picture, so that the model knows which player pair is needed when there are many people in the images.\n\nIn this way we are able to train 2.5D CNN to obtain the contact prediction between player pairs extracted from stage 1.\n\nAs for GBDT, we simple extract some features from `tracking.csv`, nothing special there.\n\nThe final CV score of our part is `0.796`\n\nAfter ensemble with @tanakar @ryotayoshinobu we got CV `0.804`\n\n\n\n\n# Acknowledge\n\nI trained many models in this competition using Z by HP Z8G4 Workstation with dual A6000 GPU as usual, it's stable and never had a problem with the training process. Thanks to Z by HP for sponsoring!",
    "2168955": "Thanks for posting your solution. I 100% agree about what you said \"So I believe it's very likely that each team uses a different approach. That's the biggest thing I'm looking forward to about this competition - because I'll be able to learn a lot of different approachs.\"\n\nCongrats on your finish!",
    "2173649": "haqishen \nCongrats on your gold medal!\nOne question, is the shape of the input to the 1st stage model, for example, as follows?\n```\n(batch_size, 2, 128, 5, h, w)\n```\nI mean,\n```\n2: 2views  \n128: 128steps  \n5: 5frames(+-2frames?)\n```\nAnd did you input (bs\\*2\\*128, 5, h, w) into the backbone?",
    "2174517": "yes that's correct!"
  },
  "source": "meta"
}