{
  "id": 392290,
  "title": "5th place solution",
  "url": "/competitions/nfl-player-contact-detection/writeups/m-t-s-s-5th-place-solution",
  "author_name": "",
  "post_date": "2023-03-04T15:13:35.507425900Z",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to the host and kaggle for hosting such an interesting competition.<br>\nI would also like to thank all of the participants and teammates( <a href=\"https://www.kaggle.com/takashisomeya\" target=\"_blank\">@takashisomeya</a> <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@nomorevotch</a> <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">@fuumin621</a> ) for a great time.</p>\n<p>Our solution consists of two stages: NN and GBDT. We will show you how in detail.</p>\n<h2>■stage1 NN part overview</h2>\n<ul>\n<li>tracking data and images as input(player-player distance &lt; 2 and player-ground)  </li>\n<li>inference of sequential frames at once  </li>\n<li>CNN + LSTM  </li>\n</ul>\n<h2>Input to NN</h2>\n<h3>[1]tracking data</h3>\n<p>Use the following tracking data.</p>\n<ul>\n<li>distance</li>\n<li>distance_1(player1)</li>\n<li>distance_2(player2)</li>\n<li>speed_1</li>\n<li>speed_2</li>\n<li>acceleration_1</li>\n<li>acceleration_2</li>\n<li>same_team(bool)</li>\n<li>different_team(bool)</li>\n<li>G_flag(bool)</li>\n</ul>\n<p>If player is G, fill distance and XXXX_2 values with -1.<br>\nsame_team and different_team are flags for whether the players are belong to the same/different team.<br>\nG_flag means the player-ground pair flag.</p>\n<h3>[2]Images + Bbox</h3>\n<ul>\n<li>Concat the following three in the channel direction<ul>\n<li>video frames of +-1 frame cropped around the helmet. </li>\n<li>helmet bbox mask</li></ul></li>\n<li>Image size<ul>\n<li>player-player pair   ：crop size = max(average bbox width, average bbox height) * 3</li>\n<li>player-ground pair ：crop size = max(bbox width, bbox height) * 3</li>\n<li>Resize the cropped image to 128x128.</li></ul></li>\n</ul>\n<p>We used sequential frames containing at least one frame with a distance &lt; 2.\n(At this time the data may contain frames of distance &gt; 2.)</p>\n<ul>\n<li>[1]：B x N x 10  </li>\n<li>[2]：B x N x 3 x 128 x 128  <br>\n(B:batch_size, N:Sequential frames (e,g. 16,32,48,64))  </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2F4e835e68beeb3243da667319ac771c14%2Fcnn_input.jpg?generation=1677942599176145&amp;alt=media\" alt=\"\"></p>\n<p>Sequential frames (N) are cut out with different strides during training and inference.   <br>\ntraining: No duplicate frames  (stride == N)  <br>\ninference: Duplicate frames(stride &lt; N, Duplicate frame results are averaged.)  </p>\n<h2>Augmentations during training</h2>\n<p>Use the following augmentations.</p>\n<ul>\n<li>HorizontalFlip</li>\n<li>RandomBrightnessContrast</li>\n<li>OneOf<ul>\n<li>MotionBlur</li>\n<li>Blur</li>\n<li>GaussianBlur  </li></ul></li>\n<li>Ramdom frame dropout (40-60% for images and 20-60% for tracking data)</li>\n</ul>\n<h2>NN Model</h2>\n<p>The overall NN model architecture is as follows  </p>\n<ul>\n<li>Endzone/sideline images go through a shared CNN backbone.  </li>\n<li>The CNN backbone uses the TSM module.  <br>\n　<a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403</a>  </li>\n<li>Concatenate features extracted by CNN with tracking features  </li>\n<li>BiLSTM layers + FC layer infer sequential frames at once  </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2Fcc3cb2c68d98704faeb57d48d11ecea4%2Fcnn_model.jpg?generation=1677942632702738&amp;alt=media\" alt=\"\"></p>\n<h2>■stage2 GBDT part overview</h2>\n<p>The key feature in this model is the logit from stage1.<br>\nThe goal is to further improve the score by combining logit with tracking data and other data to create a binary classification model.</p>\n<h2>Data</h2>\n<ul>\n<li>distance &lt;= 2</li>\n<li>swap player1 and player2 features then concatenate them vertically to the original data.</li>\n<li>average swap and original  features for final prediction</li>\n</ul>\n<h2>Features</h2>\n<h3>Raw value</h3>\n<ul>\n<li>x_position, y_position, speed, distance, orientation, acceleration, direction, sa, jersey_number of each player</li>\n<li>distance between players</li>\n<li>frame number</li>\n<li>nn_pred</li>\n</ul>\n<h3>Helmet</h3>\n<p><a href=\"https://www.kaggle.com/code/ahmedelfazouan/nfl-player-contact-detection-helmet-track-ftrs\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/nfl-player-contact-detection-helmet-track-ftrs</a></p>\n<h3>Simple computational features</h3>\n<p>The following are calculated for x_position, y_position, speed, distance, orientation, acceleration, direction, sa</p>\n<ul>\n<li>Absolute difference between players, multiplied by</li>\n<li>Difference from the average of all players in each frame</li>\n</ul>\n<h3>Aggregate features</h3>\n<p>For distance, nn_pred, sa, distance, speed</p>\n<ul>\n<li>Aggregate features for (game_play, position), (game_play, player), (game_play, team), (game_play, step)</li>\n<li>Aggregate features for each (game_play, player_1, player_2)</li>\n<li>shift, diff(-3~3) for each (game_play, player_1, player_2).</li>\n</ul>\n<h2>model</h2>\n<ul>\n<li>lgbm</li>\n<li>xgboost</li>\n</ul>\n<h2>■Ensemble</h2>\n<h3>stage1 (NN part)</h3>\n<p>Created models on different backbones and different sequence lengths as follows</p>\n<ul>\n<li>backbone<ul>\n<li>resnet18,34,50</li>\n<li>resnext50</li>\n<li>efficientnet b0,b1</li></ul></li>\n<li>sequence length<ul>\n<li>16,32,48,64</li></ul></li>\n</ul>\n<h3>stage2 (GBDT part)</h3>\n<p>Two models were created with the same features</p>\n<ul>\n<li>LightGBM</li>\n<li>XGBoost</li>\n</ul>\n<h3>Forward Selection</h3>\n<p>Created models for (almost) all combinations of the above, and use Forward Selection </p>\n<ul>\n<li>Forward Selection was based on the excellent kernel by chris here.<br>\n   <a href=\"https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private/notebook</a></li>\n<li>It is a simple method. so we expected to avoid overfit.</li>\n<li>The following models were finally selected by Forward Selection</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>sequence length</th>\n<th>backbone</th>\n<th>gbdt</th>\n<th>cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>xgb</td>\n<td>0.7918</td>\n</tr>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>lgb</td>\n<td>0.7906</td>\n</tr>\n<tr>\n<td>64</td>\n<td>effib0</td>\n<td>lgb</td>\n<td>0.79</td>\n</tr>\n<tr>\n<td>32</td>\n<td>resnext50</td>\n<td>lgb</td>\n<td>0.7935</td>\n</tr>\n<tr>\n<td>32</td>\n<td>effib0</td>\n<td>lgb</td>\n<td>0.7881</td>\n</tr>\n<tr>\n<td>16</td>\n<td>resnext50</td>\n<td>xgb</td>\n<td>0.7906</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Final submit is CV:0.8016 ,LB : 0.7902, PB : 0.7913</li>\n</ul>\n<h2>Threshold</h2>\n<p>We simply blend predictions of selected models (x5fold), and determined by a single threshold.</p>\n<ul>\n<li>We used two threshold. <ul>\n<li>predictions themselves</li>\n<li>percentile of the predictions</li></ul></li>\n<li>We also tried voting ensemble , but decided not to use it because the LB score was better with a single threshold.</li>\n</ul>\n<h2>Other tips</h2>\n<p>In the inference notebook, the following were introduced to avoid OOM and timeout.</p>\n<ul>\n<li>using lru_cache for read image at high speed</li>\n<li>PyTurboJPEG loads images faster than OpenCV</li>\n<li>Polars helps reducing submission time.</li>\n</ul>\n<h2>Acknowledgments</h2>\n<p>zzy's excellent kernel is very helpful in our pipeline.  <br>\n<a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference</a></p>",
  "messages": [
    {
      "id": "2168838",
      "postDate": "03/04/2023 15:13:35",
      "content": "<p>Thanks to the host and kaggle for hosting such an interesting competition.<br>\nI would also like to thank all of the participants and teammates( <a href=\"https://www.kaggle.com/takashisomeya\" target=\"_blank\">@takashisomeya</a> <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@nomorevotch</a> <a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">@fuumin621</a> ) for a great time.</p>\n<p>Our solution consists of two stages: NN and GBDT. We will show you how in detail.</p>\n<h2>■stage1 NN part overview</h2>\n<ul>\n<li>tracking data and images as input(player-player distance &lt; 2 and player-ground)  </li>\n<li>inference of sequential frames at once  </li>\n<li>CNN + LSTM  </li>\n</ul>\n<h2>Input to NN</h2>\n<h3>[1]tracking data</h3>\n<p>Use the following tracking data.</p>\n<ul>\n<li>distance</li>\n<li>distance_1(player1)</li>\n<li>distance_2(player2)</li>\n<li>speed_1</li>\n<li>speed_2</li>\n<li>acceleration_1</li>\n<li>acceleration_2</li>\n<li>same_team(bool)</li>\n<li>different_team(bool)</li>\n<li>G_flag(bool)</li>\n</ul>\n<p>If player is G, fill distance and XXXX_2 values with -1.<br>\nsame_team and different_team are flags for whether the players are belong to the same/different team.<br>\nG_flag means the player-ground pair flag.</p>\n<h3>[2]Images + Bbox</h3>\n<ul>\n<li>Concat the following three in the channel direction<ul>\n<li>video frames of +-1 frame cropped around the helmet. </li>\n<li>helmet bbox mask</li></ul></li>\n<li>Image size<ul>\n<li>player-player pair   ：crop size = max(average bbox width, average bbox height) * 3</li>\n<li>player-ground pair ：crop size = max(bbox width, bbox height) * 3</li>\n<li>Resize the cropped image to 128x128.</li></ul></li>\n</ul>\n<p>We used sequential frames containing at least one frame with a distance &lt; 2.\n(At this time the data may contain frames of distance &gt; 2.)</p>\n<ul>\n<li>[1]：B x N x 10  </li>\n<li>[2]：B x N x 3 x 128 x 128  <br>\n(B:batch_size, N:Sequential frames (e,g. 16,32,48,64))  </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2F4e835e68beeb3243da667319ac771c14%2Fcnn_input.jpg?generation=1677942599176145&amp;alt=media\" alt=\"\"></p>\n<p>Sequential frames (N) are cut out with different strides during training and inference.   <br>\ntraining: No duplicate frames  (stride == N)  <br>\ninference: Duplicate frames(stride &lt; N, Duplicate frame results are averaged.)  </p>\n<h2>Augmentations during training</h2>\n<p>Use the following augmentations.</p>\n<ul>\n<li>HorizontalFlip</li>\n<li>RandomBrightnessContrast</li>\n<li>OneOf<ul>\n<li>MotionBlur</li>\n<li>Blur</li>\n<li>GaussianBlur  </li></ul></li>\n<li>Ramdom frame dropout (40-60% for images and 20-60% for tracking data)</li>\n</ul>\n<h2>NN Model</h2>\n<p>The overall NN model architecture is as follows  </p>\n<ul>\n<li>Endzone/sideline images go through a shared CNN backbone.  </li>\n<li>The CNN backbone uses the TSM module.  <br>\n　<a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403</a>  </li>\n<li>Concatenate features extracted by CNN with tracking features  </li>\n<li>BiLSTM layers + FC layer infer sequential frames at once  </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2Fcc3cb2c68d98704faeb57d48d11ecea4%2Fcnn_model.jpg?generation=1677942632702738&amp;alt=media\" alt=\"\"></p>\n<h2>■stage2 GBDT part overview</h2>\n<p>The key feature in this model is the logit from stage1.<br>\nThe goal is to further improve the score by combining logit with tracking data and other data to create a binary classification model.</p>\n<h2>Data</h2>\n<ul>\n<li>distance &lt;= 2</li>\n<li>swap player1 and player2 features then concatenate them vertically to the original data.</li>\n<li>average swap and original  features for final prediction</li>\n</ul>\n<h2>Features</h2>\n<h3>Raw value</h3>\n<ul>\n<li>x_position, y_position, speed, distance, orientation, acceleration, direction, sa, jersey_number of each player</li>\n<li>distance between players</li>\n<li>frame number</li>\n<li>nn_pred</li>\n</ul>\n<h3>Helmet</h3>\n<p><a href=\"https://www.kaggle.com/code/ahmedelfazouan/nfl-player-contact-detection-helmet-track-ftrs\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/nfl-player-contact-detection-helmet-track-ftrs</a></p>\n<h3>Simple computational features</h3>\n<p>The following are calculated for x_position, y_position, speed, distance, orientation, acceleration, direction, sa</p>\n<ul>\n<li>Absolute difference between players, multiplied by</li>\n<li>Difference from the average of all players in each frame</li>\n</ul>\n<h3>Aggregate features</h3>\n<p>For distance, nn_pred, sa, distance, speed</p>\n<ul>\n<li>Aggregate features for (game_play, position), (game_play, player), (game_play, team), (game_play, step)</li>\n<li>Aggregate features for each (game_play, player_1, player_2)</li>\n<li>shift, diff(-3~3) for each (game_play, player_1, player_2).</li>\n</ul>\n<h2>model</h2>\n<ul>\n<li>lgbm</li>\n<li>xgboost</li>\n</ul>\n<h2>■Ensemble</h2>\n<h3>stage1 (NN part)</h3>\n<p>Created models on different backbones and different sequence lengths as follows</p>\n<ul>\n<li>backbone<ul>\n<li>resnet18,34,50</li>\n<li>resnext50</li>\n<li>efficientnet b0,b1</li></ul></li>\n<li>sequence length<ul>\n<li>16,32,48,64</li></ul></li>\n</ul>\n<h3>stage2 (GBDT part)</h3>\n<p>Two models were created with the same features</p>\n<ul>\n<li>LightGBM</li>\n<li>XGBoost</li>\n</ul>\n<h3>Forward Selection</h3>\n<p>Created models for (almost) all combinations of the above, and use Forward Selection </p>\n<ul>\n<li>Forward Selection was based on the excellent kernel by chris here.<br>\n   <a href=\"https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private/notebook</a></li>\n<li>It is a simple method. so we expected to avoid overfit.</li>\n<li>The following models were finally selected by Forward Selection</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>sequence length</th>\n<th>backbone</th>\n<th>gbdt</th>\n<th>cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>xgb</td>\n<td>0.7918</td>\n</tr>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>lgb</td>\n<td>0.7906</td>\n</tr>\n<tr>\n<td>64</td>\n<td>effib0</td>\n<td>lgb</td>\n<td>0.79</td>\n</tr>\n<tr>\n<td>32</td>\n<td>resnext50</td>\n<td>lgb</td>\n<td>0.7935</td>\n</tr>\n<tr>\n<td>32</td>\n<td>effib0</td>\n<td>lgb</td>\n<td>0.7881</td>\n</tr>\n<tr>\n<td>16</td>\n<td>resnext50</td>\n<td>xgb</td>\n<td>0.7906</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Final submit is CV:0.8016 ,LB : 0.7902, PB : 0.7913</li>\n</ul>\n<h2>Threshold</h2>\n<p>We simply blend predictions of selected models (x5fold), and determined by a single threshold.</p>\n<ul>\n<li>We used two threshold. <ul>\n<li>predictions themselves</li>\n<li>percentile of the predictions</li></ul></li>\n<li>We also tried voting ensemble , but decided not to use it because the LB score was better with a single threshold.</li>\n</ul>\n<h2>Other tips</h2>\n<p>In the inference notebook, the following were introduced to avoid OOM and timeout.</p>\n<ul>\n<li>using lru_cache for read image at high speed</li>\n<li>PyTurboJPEG loads images faster than OpenCV</li>\n<li>Polars helps reducing submission time.</li>\n</ul>\n<h2>Acknowledgments</h2>\n<p>zzy's excellent kernel is very helpful in our pipeline.  <br>\n<a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference</a></p>",
      "rawMarkdown": "Thanks to the host and kaggle for hosting such an interesting competition.\nI would also like to thank all of the participants and teammates( @takashisomeya @nomorevotch @fuumin621 ) for a great time.\n\nOur solution consists of two stages: NN and GBDT. We will show you how in detail.\n\n## ■stage1 NN part overview\n- tracking data and images as input(player-player distance < 2 and player-ground)  \n- inference of sequential frames at once  \n- CNN + LSTM  \n\n## Input to NN\n### [1]tracking data\nUse the following tracking data.\n- distance\n- distance_1(player1)\n- distance_2(player2)\n- speed_1\n- speed_2\n- acceleration_1\n- acceleration_2\n- same_team(bool)\n- different_team(bool)\n- G_flag(bool)\n\n\nIf player is G, fill distance and XXXX_2 values with -1.\nsame_team and different_team are flags for whether the players are belong to the same/different team.\nG_flag means the player-ground pair flag.\n\n### [2]Images + Bbox\n- Concat the following three in the channel direction\n    - video frames of +-1 frame cropped around the helmet. \n    - helmet bbox mask\n- Image size\n    - player-player pair   ：crop size = max(average bbox width, average bbox height) * 3\n    - player-ground pair ：crop size = max(bbox width, bbox height) * 3\n    - Resize the cropped image to 128x128.\n\nWe used sequential frames containing at least one frame with a distance < 2.\n(At this time the data may contain frames of distance > 2.)\n- [1]：B x N x 10  \n- [2]：B x N x 3 x 128 x 128  \n(B:batch_size, N:Sequential frames (e,g. 16,32,48,64))  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2F4e835e68beeb3243da667319ac771c14%2Fcnn_input.jpg?generation=1677942599176145&alt=media)\n\nSequential frames (N) are cut out with different strides during training and inference.   \ntraining: No duplicate frames  (stride == N)  \ninference: Duplicate frames(stride < N, Duplicate frame results are averaged.)  \n\n\n## Augmentations during training\nUse the following augmentations.\n- HorizontalFlip\n- RandomBrightnessContrast\n- OneOf\n    - MotionBlur\n    - Blur\n    - GaussianBlur  \n- Ramdom frame dropout (40-60% for images and 20-60% for tracking data)\n\n## NN Model\nThe overall NN model architecture is as follows  \n- Endzone/sideline images go through a shared CNN backbone.  \n- The CNN backbone uses the TSM module.  \n　https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403  \n- Concatenate features extracted by CNN with tracking features  \n- BiLSTM layers + FC layer infer sequential frames at once  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2Fcc3cb2c68d98704faeb57d48d11ecea4%2Fcnn_model.jpg?generation=1677942632702738&alt=media)\n\n\n## ■stage2 GBDT part overview\n\nThe key feature in this model is the logit from stage1.\nThe goal is to further improve the score by combining logit with tracking data and other data to create a binary classification model.\n\n## Data\n\n- distance <= 2\n- swap player1 and player2 features then concatenate them vertically to the original data.\n- average swap and original  features for final prediction\n\n## Features\n### Raw value\n- x_position, y_position, speed, distance, orientation, acceleration, direction, sa, jersey_number of each player\n- distance between players\n- frame number\n- nn_pred\n\n### Helmet\nhttps://www.kaggle.com/code/ahmedelfazouan/nfl-player-contact-detection-helmet-track-ftrs\n\n\n### Simple computational features\n\nThe following are calculated for x_position, y_position, speed, distance, orientation, acceleration, direction, sa\n- Absolute difference between players, multiplied by\n- Difference from the average of all players in each frame\n\n\n### Aggregate features\nFor distance, nn_pred, sa, distance, speed\n- Aggregate features for (game_play, position), (game_play, player), (game_play, team), (game_play, step)\n- Aggregate features for each (game_play, player_1, player_2)\n- shift, diff(-3~3) for each (game_play, player_1, player_2).\n\n## model\n- lgbm\n- xgboost\n\n## ■Ensemble \n### stage1 (NN part)\nCreated models on different backbones and different sequence lengths as follows\n* backbone\n  * resnet18,34,50\n  * resnext50\n  * efficientnet b0,b1\n* sequence length\n  * 16,32,48,64\n### stage2 (GBDT part)\nTwo models were created with the same features\n* LightGBM\n* XGBoost\n\n### Forward Selection\nCreated models for (almost) all combinations of the above, and use Forward Selection \n* Forward Selection was based on the excellent kernel by chris here.\n       https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private/notebook\n* It is a simple method. so we expected to avoid overfit.\n* The following models were finally selected by Forward Selection\n\n| sequence length | backbone | gbdt | cv |\n| --- | --- | --- | --- |\n| 64 | resnext50 | xgb | 0.7918 |\n| 64 | resnext50 | lgb | 0.7906 |\n| 64 | effib0 | lgb | 0.79 |\n| 32 | resnext50 | lgb | 0.7935 |\n| 32 | effib0 | lgb | 0.7881 |\n| 16 | resnext50 | xgb | 0.7906 |\n* Final submit is CV:0.8016 ,LB : 0.7902, PB : 0.7913\n\n## Threshold\nWe simply blend predictions of selected models (x5fold), and determined by a single threshold.\n* We used two threshold. \n   * predictions themselves\n   * percentile of the predictions\n* We also tried voting ensemble , but decided not to use it because the LB score was better with a single threshold.\n\n## Other tips\nIn the inference notebook, the following were introduced to avoid OOM and timeout.\n  * using lru_cache for read image at high speed\n  * PyTurboJPEG loads images faster than OpenCV\n  * Polars helps reducing submission time.\n\n## Acknowledgments\nzzy's excellent kernel is very helpful in our pipeline.  \nhttps://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference",
      "votes": null
    },
    {
      "id": "2168940",
      "postDate": "03/04/2023 16:45:05",
      "content": "<p>Great work and thanks for sharing the solution writeup. It seems many teams had good results with a 2-stage approach, but it's fun to see so much diversity in the architectures used. Do you happen to know the CV gain from stage 1 vs stage 2?</p>\n<p>Looking forward to learning more about your solution!</p>",
      "rawMarkdown": "Great work and thanks for sharing the solution writeup. It seems many teams had good results with a 2-stage approach, but it's fun to see so much diversity in the architectures used. Do you happen to know the CV gain from stage 1 vs stage 2?\n\nLooking forward to learning more about your solution!",
      "votes": null
    },
    {
      "id": "2169680",
      "postDate": "03/05/2023 10:13:37",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a></p>\n<p>The stage 2 gain depends mainly on the sequential length.   <br>\nIf the sequential length is small, the gain is large, but if the sequential length is large, the gain is almost none.</p>\n<table>\n<thead>\n<tr>\n<th>sequence length</th>\n<th>backbone</th>\n<th>stage1 cv</th>\n<th>stage2 cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>16</td>\n<td>resnext50</td>\n<td>0.7868</td>\n<td>0.7906(xgb)</td>\n</tr>\n<tr>\n<td>32</td>\n<td>resnext50</td>\n<td>0.7929</td>\n<td>0.7935(lgb)</td>\n</tr>\n<tr>\n<td>32</td>\n<td>effib0</td>\n<td>0.785</td>\n<td>0.7881(lgb)</td>\n</tr>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>0.7923</td>\n<td>0.7918(xgb)</td>\n</tr>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>0.7923</td>\n<td>0.7906(lgb)</td>\n</tr>\n<tr>\n<td>64</td>\n<td>effib0</td>\n<td>0.7893</td>\n<td>0.79(lgb)</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thank you, @robikscube\n\nThe stage 2 gain depends mainly on the sequential length.   \nIf the sequential length is small, the gain is large, but if the sequential length is large, the gain is almost none.\n\n| sequence length | backbone | stage1 cv | stage2 cv |\n| --- | --- | --- | --- |\n| 16 | resnext50 | 0.7868 | 0.7906(xgb) |\n| 32 | resnext50 | 0.7929 | 0.7935(lgb) |\n| 32 | effib0 | 0.785 | 0.7881(lgb) |\n| 64 | resnext50 | 0.7923 | 0.7918(xgb) |\n| 64 | resnext50 | 0.7923 | 0.7906(lgb) |\n| 64 | effib0 | 0.7893 | 0.79(lgb) |",
      "votes": null
    },
    {
      "id": "2169762",
      "postDate": "03/05/2023 12:09:00",
      "content": "<p>Thanks for sharing great solution.</p>\n<p>I have a question about the input to the CNN.<br>\nDoes the Sequential frames(N) stack up on the channel axis?<br>\nOr do you input each frame to the CNN and concatenate the output of the CNN?<br>\nThe input to the CNN is usually B*C*H*W, but I would like to know how you have handled the \"N\" in your solutions that deal with B*N*C*H*W.</p>",
      "rawMarkdown": "Thanks for sharing great solution.\n\nI have a question about the input to the CNN.\nDoes the Sequential frames(N) stack up on the channel axis?\nOr do you input each frame to the CNN and concatenate the output of the CNN?\nThe input to the CNN is usually B\\*C\\*H\\*W, but I would like to know how you have handled the \"N\" in your solutions that deal with B\\*N\\*C\\*H\\*W.",
      "votes": null
    },
    {
      "id": "2169811",
      "postDate": "03/05/2023 13:05:31",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/yururoi\" target=\"_blank\">@yururoi</a> <br>\nThe input to the CNN is the latter, that is (BxN)xCxHxW.</p>",
      "rawMarkdown": "Thanks, @yururoi \nThe input to the CNN is the latter, that is (BxN)xCxHxW.",
      "votes": null
    },
    {
      "id": "2170036",
      "postDate": "03/05/2023 16:27:06",
      "content": "<p>Thanks for sharing great solution and explaination. your Forward Selection and ensemble solutions are nice ~</p>",
      "rawMarkdown": "Thanks for sharing great solution and explaination. your Forward Selection and ensemble solutions are nice ~",
      "votes": null
    },
    {
      "id": "2170089",
      "postDate": "03/05/2023 17:10:53",
      "content": "<p>Understood! Thank you!</p>",
      "rawMarkdown": "Understood! Thank you!",
      "votes": null
    },
    {
      "id": "2189022",
      "postDate": "03/20/2023 07:08:34",
      "content": "<p>Hi Tamo, thank you for sharing this great solution!</p>\n<p>Your explanation is very straight forward, but I still got one question about the sequential frames N.</p>\n<p>Are they extracted by 6 frames interval?<br>\nWhen you say N=16, so there are 16 sample steps with contact label?<br>\nI would very appreciate if you can explain my question above!</p>",
      "rawMarkdown": "Hi Tamo, thank you for sharing this great solution!\n\nYour explanation is very straight forward, but I still got one question about the sequential frames N.\n\nAre they extracted by 6 frames interval?\nWhen you say N=16, so there are 16 sample steps with contact label?\nI would very appreciate if you can explain my question above!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168940,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:45:05",
      "content": "<p>Great work and thanks for sharing the solution writeup. It seems many teams had good results with a 2-stage approach, but it's fun to see so much diversity in the architectures used. Do you happen to know the CV gain from stage 1 vs stage 2?</p>\n<p>Looking forward to learning more about your solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169680,
          "author_name": "yuyuki11235",
          "author_url": "",
          "post_date": "03/05/2023 10:13:37",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a></p>\n<p>The stage 2 gain depends mainly on the sequential length.   <br>\nIf the sequential length is small, the gain is large, but if the sequential length is large, the gain is almost none.</p>\n<table>\n<thead>\n<tr>\n<th>sequence length</th>\n<th>backbone</th>\n<th>stage1 cv</th>\n<th>stage2 cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>16</td>\n<td>resnext50</td>\n<td>0.7868</td>\n<td>0.7906(xgb)</td>\n</tr>\n<tr>\n<td>32</td>\n<td>resnext50</td>\n<td>0.7929</td>\n<td>0.7935(lgb)</td>\n</tr>\n<tr>\n<td>32</td>\n<td>effib0</td>\n<td>0.785</td>\n<td>0.7881(lgb)</td>\n</tr>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>0.7923</td>\n<td>0.7918(xgb)</td>\n</tr>\n<tr>\n<td>64</td>\n<td>resnext50</td>\n<td>0.7923</td>\n<td>0.7906(lgb)</td>\n</tr>\n<tr>\n<td>64</td>\n<td>effib0</td>\n<td>0.7893</td>\n<td>0.79(lgb)</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 2170036,
              "author_name": "chg0901",
              "author_url": "",
              "post_date": "03/05/2023 16:27:06",
              "content": "<p>Thanks for sharing great solution and explaination. your Forward Selection and ensemble solutions are nice ~</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2169762,
      "author_name": "yururoi",
      "author_url": "",
      "post_date": "03/05/2023 12:09:00",
      "content": "<p>Thanks for sharing great solution.</p>\n<p>I have a question about the input to the CNN.<br>\nDoes the Sequential frames(N) stack up on the channel axis?<br>\nOr do you input each frame to the CNN and concatenate the output of the CNN?<br>\nThe input to the CNN is usually B*C*H*W, but I would like to know how you have handled the \"N\" in your solutions that deal with B*N*C*H*W.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169811,
          "author_name": "yuyuki11235",
          "author_url": "",
          "post_date": "03/05/2023 13:05:31",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/yururoi\" target=\"_blank\">@yururoi</a> <br>\nThe input to the CNN is the latter, that is (BxN)xCxHxW.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2170089,
              "author_name": "yururoi",
              "author_url": "",
              "post_date": "03/05/2023 17:10:53",
              "content": "<p>Understood! Thank you!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2189022,
      "author_name": "traptinblur",
      "author_url": "",
      "post_date": "03/20/2023 07:08:34",
      "content": "<p>Hi Tamo, thank you for sharing this great solution!</p>\n<p>Your explanation is very straight forward, but I still got one question about the sequential frames N.</p>\n<p>Are they extracted by 6 frames interval?<br>\nWhen you say N=16, so there are 16 sample steps with contact label?<br>\nI would very appreciate if you can explain my question above!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2168838": "Thanks to the host and kaggle for hosting such an interesting competition.\nI would also like to thank all of the participants and teammates( @takashisomeya @nomorevotch @fuumin621 ) for a great time.\n\nOur solution consists of two stages: NN and GBDT. We will show you how in detail.\n\n## ■stage1 NN part overview\n- tracking data and images as input(player-player distance < 2 and player-ground)  \n- inference of sequential frames at once  \n- CNN + LSTM  \n\n## Input to NN\n### [1]tracking data\nUse the following tracking data.\n- distance\n- distance_1(player1)\n- distance_2(player2)\n- speed_1\n- speed_2\n- acceleration_1\n- acceleration_2\n- same_team(bool)\n- different_team(bool)\n- G_flag(bool)\n\n\nIf player is G, fill distance and XXXX_2 values with -1.\nsame_team and different_team are flags for whether the players are belong to the same/different team.\nG_flag means the player-ground pair flag.\n\n### [2]Images + Bbox\n- Concat the following three in the channel direction\n    - video frames of +-1 frame cropped around the helmet. \n    - helmet bbox mask\n- Image size\n    - player-player pair   ：crop size = max(average bbox width, average bbox height) * 3\n    - player-ground pair ：crop size = max(bbox width, bbox height) * 3\n    - Resize the cropped image to 128x128.\n\nWe used sequential frames containing at least one frame with a distance < 2.\n(At this time the data may contain frames of distance > 2.)\n- [1]：B x N x 10  \n- [2]：B x N x 3 x 128 x 128  \n(B:batch_size, N:Sequential frames (e,g. 16,32,48,64))  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2F4e835e68beeb3243da667319ac771c14%2Fcnn_input.jpg?generation=1677942599176145&alt=media)\n\nSequential frames (N) are cut out with different strides during training and inference.   \ntraining: No duplicate frames  (stride == N)  \ninference: Duplicate frames(stride < N, Duplicate frame results are averaged.)  \n\n\n## Augmentations during training\nUse the following augmentations.\n- HorizontalFlip\n- RandomBrightnessContrast\n- OneOf\n    - MotionBlur\n    - Blur\n    - GaussianBlur  \n- Ramdom frame dropout (40-60% for images and 20-60% for tracking data)\n\n## NN Model\nThe overall NN model architecture is as follows  \n- Endzone/sideline images go through a shared CNN backbone.  \n- The CNN backbone uses the TSM module.  \n　https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403  \n- Concatenate features extracted by CNN with tracking features  \n- BiLSTM layers + FC layer infer sequential frames at once  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3584397%2Fcc3cb2c68d98704faeb57d48d11ecea4%2Fcnn_model.jpg?generation=1677942632702738&alt=media)\n\n\n## ■stage2 GBDT part overview\n\nThe key feature in this model is the logit from stage1.\nThe goal is to further improve the score by combining logit with tracking data and other data to create a binary classification model.\n\n## Data\n\n- distance <= 2\n- swap player1 and player2 features then concatenate them vertically to the original data.\n- average swap and original  features for final prediction\n\n## Features\n### Raw value\n- x_position, y_position, speed, distance, orientation, acceleration, direction, sa, jersey_number of each player\n- distance between players\n- frame number\n- nn_pred\n\n### Helmet\nhttps://www.kaggle.com/code/ahmedelfazouan/nfl-player-contact-detection-helmet-track-ftrs\n\n\n### Simple computational features\n\nThe following are calculated for x_position, y_position, speed, distance, orientation, acceleration, direction, sa\n- Absolute difference between players, multiplied by\n- Difference from the average of all players in each frame\n\n\n### Aggregate features\nFor distance, nn_pred, sa, distance, speed\n- Aggregate features for (game_play, position), (game_play, player), (game_play, team), (game_play, step)\n- Aggregate features for each (game_play, player_1, player_2)\n- shift, diff(-3~3) for each (game_play, player_1, player_2).\n\n## model\n- lgbm\n- xgboost\n\n## ■Ensemble \n### stage1 (NN part)\nCreated models on different backbones and different sequence lengths as follows\n* backbone\n  * resnet18,34,50\n  * resnext50\n  * efficientnet b0,b1\n* sequence length\n  * 16,32,48,64\n### stage2 (GBDT part)\nTwo models were created with the same features\n* LightGBM\n* XGBoost\n\n### Forward Selection\nCreated models for (almost) all combinations of the above, and use Forward Selection \n* Forward Selection was based on the excellent kernel by chris here.\n       https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private/notebook\n* It is a simple method. so we expected to avoid overfit.\n* The following models were finally selected by Forward Selection\n\n| sequence length | backbone | gbdt | cv |\n| --- | --- | --- | --- |\n| 64 | resnext50 | xgb | 0.7918 |\n| 64 | resnext50 | lgb | 0.7906 |\n| 64 | effib0 | lgb | 0.79 |\n| 32 | resnext50 | lgb | 0.7935 |\n| 32 | effib0 | lgb | 0.7881 |\n| 16 | resnext50 | xgb | 0.7906 |\n* Final submit is CV:0.8016 ,LB : 0.7902, PB : 0.7913\n\n## Threshold\nWe simply blend predictions of selected models (x5fold), and determined by a single threshold.\n* We used two threshold. \n   * predictions themselves\n   * percentile of the predictions\n* We also tried voting ensemble , but decided not to use it because the LB score was better with a single threshold.\n\n## Other tips\nIn the inference notebook, the following were introduced to avoid OOM and timeout.\n  * using lru_cache for read image at high speed\n  * PyTurboJPEG loads images faster than OpenCV\n  * Polars helps reducing submission time.\n\n## Acknowledgments\nzzy's excellent kernel is very helpful in our pipeline.  \nhttps://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference",
    "2168940": "Great work and thanks for sharing the solution writeup. It seems many teams had good results with a 2-stage approach, but it's fun to see so much diversity in the architectures used. Do you happen to know the CV gain from stage 1 vs stage 2?\n\nLooking forward to learning more about your solution!",
    "2169680": "Thank you, @robikscube\n\nThe stage 2 gain depends mainly on the sequential length.   \nIf the sequential length is small, the gain is large, but if the sequential length is large, the gain is almost none.\n\n| sequence length | backbone | stage1 cv | stage2 cv |\n| --- | --- | --- | --- |\n| 16 | resnext50 | 0.7868 | 0.7906(xgb) |\n| 32 | resnext50 | 0.7929 | 0.7935(lgb) |\n| 32 | effib0 | 0.785 | 0.7881(lgb) |\n| 64 | resnext50 | 0.7923 | 0.7918(xgb) |\n| 64 | resnext50 | 0.7923 | 0.7906(lgb) |\n| 64 | effib0 | 0.7893 | 0.79(lgb) |",
    "2169762": "Thanks for sharing great solution.\n\nI have a question about the input to the CNN.\nDoes the Sequential frames(N) stack up on the channel axis?\nOr do you input each frame to the CNN and concatenate the output of the CNN?\nThe input to the CNN is usually B\\*C\\*H\\*W, but I would like to know how you have handled the \"N\" in your solutions that deal with B\\*N\\*C\\*H\\*W.",
    "2169811": "Thanks, @yururoi \nThe input to the CNN is the latter, that is (BxN)xCxHxW.",
    "2170036": "Thanks for sharing great solution and explaination. your Forward Selection and ensemble solutions are nice ~",
    "2170089": "Understood! Thank you!",
    "2189022": "Hi Tamo, thank you for sharing this great solution!\n\nYour explanation is very straight forward, but I still got one question about the sequential frames N.\n\nAre they extracted by 6 frames interval?\nWhen you say N=16, so there are 16 sample steps with contact label?\nI would very appreciate if you can explain my question above!"
  },
  "source": "meta"
}