{
  "id": 391761,
  "title": "4th place solution Overall pipeline & tabular part - Osaka Tigers",
  "url": "/competitions/nfl-player-contact-detection/writeups/osaka-tigers-4th-place-solution-overall-pipeline-t",
  "author_name": "",
  "post_date": "2023-03-03T01:12:47.513Z",
  "votes": 37,
  "comment_count": 7,
  "views": 0,
  "content": "<p>We really appreciated the hosts and the kaggle team for organizing the competition. Moreover, we would also like to thank all the participants who joined. We could enjoy this competition and write up our solutions. </p>\n<p>I would like to thank team members, <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>, <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> and <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a>, who have the top talent to analyze the task. I could discuss and enjoy the competition. </p>\n<h1>Overview</h1>\n<p>Simple solution outline is attached pic.<br>\n<a href=\"https://postimg.cc/VJ6Rkh2p\" target=\"_blank\"><img src=\"https://i.postimg.cc/pLQ1qbCW/pipeline.png\" alt=\"pipeline.png\"></a></p>\n<p>In the 1st stage we predict the contact by multiple CNN. In the 2nd stage CNN prediction(s), tracking and helmet data are aggregated and created features to input GDBT.  Lastly we compute 5 models averaged value and optimize threshold for both player-player and player-ground contact.</p>\n<h1>1st stage CNN</h1>\n<h2>k mat model</h2>\n<p>Details are written in <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719</a>.<br>\nWe can obtain both Endzone and Sideline prediction values. </p>\n<h2>camaro model</h2>\n<p>will come up soon</p>\n<h1>2nd stage aggregation &amp; binary classification models</h1>\n<p>We excluded player-player pairs with distance &gt; 3, and the remaining ~880K rows were used to train 2nd stage models. During inference time, we assigned 0 to pair with distance &gt; 3 and predicted only the remaining data.</p>\n<h2>Created features</h2>\n<p>Because our CNN predictions are so strong, more than 90% of the top 30 important features were CNN-related features. Below are part of the features we have created.</p>\n<h3>Tracking</h3>\n<ul>\n<li>distance between two players</li>\n<li>distance/x_position/y_position from step0</li>\n<li>distance from around player (full/same team/different team )</li>\n<li>distance between team center</li>\n<li>distance to second nearest player</li>\n<li>current step / max step</li>\n<li>lag / lead of acc, speed, sa etc</li>\n<li>max/min/mean of x, y, speed, acc, sa, distance group by (play, step), (play, step, team) and (play, player1, player2) x/y positon diff from step=0</li>\n<li>”interceptor” features<ul>\n<li>find playerC who meet the following conditions and add distance(A, C) and ∠BAC to the features of playerA-playerB (to detect that C intercepts between A-B)<ul>\n<li>∠BAC &lt; 30deg</li>\n<li>distance(A, C) &lt; distance(A, B) and distance(B, C) &lt; distance(A, B)</li></ul></li></ul></li>\n</ul>\n<h3>Helmet</h3>\n<ul>\n<li>bbox aspect ratio</li>\n<li>bbox overlap</li>\n<li>lag / lead of bbox coordinates</li>\n<li>bbox center x,y std/shift/diff</li>\n<li>distance of bbox centers</li>\n</ul>\n<h3>CNN prediction and  meta-features</h3>\n<ul>\n<li>oof predictions of 1st stage CNNs</li>\n<li>max/min/std of predictions group by (play, step) and (play, player1, player2)</li>\n<li>5/11/21 rolling features<ul>\n<li>to complement CNN predictions on frames without helmets</li></ul></li>\n<li>lag / diff</li>\n<li>around players’ player-ground prediction value</li>\n</ul>\n<h4>Combinations</h4>\n<ul>\n<li>registration errors from helmet-tracking coordinate transform (similar to 6th place solution, and previous NFL’s 1st place solution by K_mat)</li>\n</ul>\n<h3>Models</h3>\n<p>We trained four GBDT models with different combinations of 1st stage CNNs. We also added one NN model (\"camaro2\" in the figure above) and calculated the simple average of these 5 models. Predictions were binarized with separate thresholds optimized for player-player and player-ground respectively.</p>\n<ul>\n<li>LightGBM<ul>\n<li>K_mat A + Camaro1 Public 0.795/Private 0.792</li>\n<li>K_mat B + Camaro 1</li>\n<li>K_mat B</li></ul></li>\n<li>xgboost<ul>\n<li>K_mat B + Camaro 1</li></ul></li>\n<li>Camaro 2</li>\n</ul>\n<h3>tips</h3>\n<ul>\n<li>rolling features for CNN prediction values are most important in our models.</li>\n<li>judging from permutation feature importance, ‘minimum distance between players in the game_play’,  ‘distance between away team mean and home team mean’ and ‘player-player distance’ are important tracking features to increase score.</li>\n<li>We did not use early-stopping to train the GBDTs because the optimal number of rounds for MCC is always longer than AUC.</li>\n</ul>\n<h3>not wroked for models</h3>\n<ul>\n<li>Catboost</li>\n<li>Residual fit</li>\n<li>Meta Features by non CNN (e.g. logistic regression prediction values/ k-means clustering feature)</li>\n<li>Separate player-player and player-ground model</li>\n<li>1DCNN</li>\n<li>External NFL data</li>\n<li>Focal loss</li>\n</ul>\n<h1>not worked for overall</h1>\n<ul>\n<li>Adding previous competition pseudo labeling data</li>\n<li>Removing noisy label</li>\n<li>all29 assignment and its prediction</li>\n<li>2.5D or 3D CNN, but should have dug more..</li>\n<li>Aggregate near frame information</li>\n</ul>",
  "messages": [
    {
      "id": "2166037",
      "postDate": "03/02/2023 15:36:10",
      "content": "<p>We really appreciated the hosts and the kaggle team for organizing the competition. Moreover, we would also like to thank all the participants who joined. We could enjoy this competition and write up our solutions. </p>\n<p>I would like to thank team members, <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>, <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> and <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a>, who have the top talent to analyze the task. I could discuss and enjoy the competition. </p>\n<h1>Overview</h1>\n<p>Simple solution outline is attached pic.<br>\n<a href=\"https://postimg.cc/VJ6Rkh2p\" target=\"_blank\"><img src=\"https://i.postimg.cc/pLQ1qbCW/pipeline.png\" alt=\"pipeline.png\"></a></p>\n<p>In the 1st stage we predict the contact by multiple CNN. In the 2nd stage CNN prediction(s), tracking and helmet data are aggregated and created features to input GDBT.  Lastly we compute 5 models averaged value and optimize threshold for both player-player and player-ground contact.</p>\n<h1>1st stage CNN</h1>\n<h2>k mat model</h2>\n<p>Details are written in <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719</a>.<br>\nWe can obtain both Endzone and Sideline prediction values. </p>\n<h2>camaro model</h2>\n<p>will come up soon</p>\n<h1>2nd stage aggregation &amp; binary classification models</h1>\n<p>We excluded player-player pairs with distance &gt; 3, and the remaining ~880K rows were used to train 2nd stage models. During inference time, we assigned 0 to pair with distance &gt; 3 and predicted only the remaining data.</p>\n<h2>Created features</h2>\n<p>Because our CNN predictions are so strong, more than 90% of the top 30 important features were CNN-related features. Below are part of the features we have created.</p>\n<h3>Tracking</h3>\n<ul>\n<li>distance between two players</li>\n<li>distance/x_position/y_position from step0</li>\n<li>distance from around player (full/same team/different team )</li>\n<li>distance between team center</li>\n<li>distance to second nearest player</li>\n<li>current step / max step</li>\n<li>lag / lead of acc, speed, sa etc</li>\n<li>max/min/mean of x, y, speed, acc, sa, distance group by (play, step), (play, step, team) and (play, player1, player2) x/y positon diff from step=0</li>\n<li>”interceptor” features<ul>\n<li>find playerC who meet the following conditions and add distance(A, C) and ∠BAC to the features of playerA-playerB (to detect that C intercepts between A-B)<ul>\n<li>∠BAC &lt; 30deg</li>\n<li>distance(A, C) &lt; distance(A, B) and distance(B, C) &lt; distance(A, B)</li></ul></li></ul></li>\n</ul>\n<h3>Helmet</h3>\n<ul>\n<li>bbox aspect ratio</li>\n<li>bbox overlap</li>\n<li>lag / lead of bbox coordinates</li>\n<li>bbox center x,y std/shift/diff</li>\n<li>distance of bbox centers</li>\n</ul>\n<h3>CNN prediction and  meta-features</h3>\n<ul>\n<li>oof predictions of 1st stage CNNs</li>\n<li>max/min/std of predictions group by (play, step) and (play, player1, player2)</li>\n<li>5/11/21 rolling features<ul>\n<li>to complement CNN predictions on frames without helmets</li></ul></li>\n<li>lag / diff</li>\n<li>around players’ player-ground prediction value</li>\n</ul>\n<h4>Combinations</h4>\n<ul>\n<li>registration errors from helmet-tracking coordinate transform (similar to 6th place solution, and previous NFL’s 1st place solution by K_mat)</li>\n</ul>\n<h3>Models</h3>\n<p>We trained four GBDT models with different combinations of 1st stage CNNs. We also added one NN model (\"camaro2\" in the figure above) and calculated the simple average of these 5 models. Predictions were binarized with separate thresholds optimized for player-player and player-ground respectively.</p>\n<ul>\n<li>LightGBM<ul>\n<li>K_mat A + Camaro1 Public 0.795/Private 0.792</li>\n<li>K_mat B + Camaro 1</li>\n<li>K_mat B</li></ul></li>\n<li>xgboost<ul>\n<li>K_mat B + Camaro 1</li></ul></li>\n<li>Camaro 2</li>\n</ul>\n<h3>tips</h3>\n<ul>\n<li>rolling features for CNN prediction values are most important in our models.</li>\n<li>judging from permutation feature importance, ‘minimum distance between players in the game_play’,  ‘distance between away team mean and home team mean’ and ‘player-player distance’ are important tracking features to increase score.</li>\n<li>We did not use early-stopping to train the GBDTs because the optimal number of rounds for MCC is always longer than AUC.</li>\n</ul>\n<h3>not wroked for models</h3>\n<ul>\n<li>Catboost</li>\n<li>Residual fit</li>\n<li>Meta Features by non CNN (e.g. logistic regression prediction values/ k-means clustering feature)</li>\n<li>Separate player-player and player-ground model</li>\n<li>1DCNN</li>\n<li>External NFL data</li>\n<li>Focal loss</li>\n</ul>\n<h1>not worked for overall</h1>\n<ul>\n<li>Adding previous competition pseudo labeling data</li>\n<li>Removing noisy label</li>\n<li>all29 assignment and its prediction</li>\n<li>2.5D or 3D CNN, but should have dug more..</li>\n<li>Aggregate near frame information</li>\n</ul>",
      "rawMarkdown": "We really appreciated the hosts and the kaggle team for organizing the competition. Moreover, we would also like to thank all the participants who joined. We could enjoy this competition and write up our solutions. \n\nI would like to thank team members, @bamps53, @nyanpn and @kmat2019, who have the top talent to analyze the task. I could discuss and enjoy the competition. \n\n# Overview\nSimple solution outline is attached pic.\n[![pipeline.png](https://i.postimg.cc/pLQ1qbCW/pipeline.png)](https://postimg.cc/VJ6Rkh2p)\n\nIn the 1st stage we predict the contact by multiple CNN. In the 2nd stage CNN prediction(s), tracking and helmet data are aggregated and created features to input GDBT.  Lastly we compute 5 models averaged value and optimize threshold for both player-player and player-ground contact.\n \n\n# 1st stage CNN\n## k mat model\nDetails are written in https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719.\nWe can obtain both Endzone and Sideline prediction values. \n\n## camaro model\nwill come up soon\n\n# 2nd stage aggregation & binary classification models\n\nWe excluded player-player pairs with distance > 3, and the remaining ~880K rows were used to train 2nd stage models. During inference time, we assigned 0 to pair with distance > 3 and predicted only the remaining data.\n\n## Created features\nBecause our CNN predictions are so strong, more than 90% of the top 30 important features were CNN-related features. Below are part of the features we have created.\n### Tracking \n - distance between two players\n - distance/x_position/y_position from step0\n - distance from around player (full/same team/different team )\n - distance between team center\n - distance to second nearest player\n - current step / max step\n - lag / lead of acc, speed, sa etc\n - max/min/mean of x, y, speed, acc, sa, distance group by (play, step), (play, step, team) and (play, player1, player2) x/y positon diff from step=0\n - ”interceptor” features\n  - find playerC who meet the following conditions and add distance(A, C) and ∠BAC to the features of playerA-playerB (to detect that C intercepts between A-B)\n     - ∠BAC < 30deg\n     - distance(A, C) < distance(A, B) and distance(B, C) < distance(A, B)\n### Helmet\n - bbox aspect ratio\n - bbox overlap\n - lag / lead of bbox coordinates\n - bbox center x,y std/shift/diff\n - distance of bbox centers\n### CNN prediction and  meta-features\n - oof predictions of 1st stage CNNs\n - max/min/std of predictions group by (play, step) and (play, player1, player2)\n - 5/11/21 rolling features\n  - to complement CNN predictions on frames without helmets\n - lag / diff\n - around players’ player-ground prediction value\n#### Combinations\n - registration errors from helmet-tracking coordinate transform (similar to 6th place solution, and previous NFL’s 1st place solution by K_mat)\n\n### Models\nWe trained four GBDT models with different combinations of 1st stage CNNs. We also added one NN model (\"camaro2\" in the figure above) and calculated the simple average of these 5 models. Predictions were binarized with separate thresholds optimized for player-player and player-ground respectively.\n\n - LightGBM\n   - K_mat A + Camaro1 Public 0.795/Private 0.792\n   - K_mat B + Camaro 1\n   - K_mat B\n - xgboost\n   - K_mat B + Camaro 1\n - Camaro 2\n### tips\n- rolling features for CNN prediction values are most important in our models.\n- judging from permutation feature importance, ‘minimum distance between players in the game_play’,  ‘distance between away team mean and home team mean’ and ‘player-player distance’ are important tracking features to increase score.\n- We did not use early-stopping to train the GBDTs because the optimal number of rounds for MCC is always longer than AUC.\n\n### not wroked for models\n- Catboost\n- Residual fit\n- Meta Features by non CNN (e.g. logistic regression prediction values/ k-means clustering feature)\n- Separate player-player and player-ground model\n- 1DCNN\n- External NFL data\n- Focal loss\n\n\n\n\n\n# not worked for overall\n- Adding previous competition pseudo labeling data\n- Removing noisy label\n- all29 assignment and its prediction\n- 2.5D or 3D CNN, but should have dug more..\n- Aggregate near frame information",
      "votes": null
    },
    {
      "id": "2167364",
      "postDate": "03/03/2023 12:51:16",
      "content": "<h1>Camaro part</h1>\n<h2>1st stage</h2>\n<p>Kind of Object Detection like archiarchitecture. Predict all combination pairs' contact and ground contact at once.  (Should I name as YOLO?)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F87f137cab92a5afa4827695aa2ff72ad%2F1st_stage.png?generation=1677847249397624&amp;alt=media\" alt=\"\"></p>\n<h3>Details</h3>\n<ol>\n<li>Extract feature map by Yolox FPN.</li>\n<li>RoIAlign around the player by using helmet coordinates.<br>\na. Region size is adjusted by average helmet size in the frame<br>\nb. Helmet is located in the same position in the crop area</li>\n<li>Concatenate tracking features after linear transformation</li>\n<li>For inter contact<br>\na. Creare pairwise player matrix and concat p1 and p2 features<br>\nb. Multiply distance features<br>\nc. Linear</li>\n<li>For ground contact, simply linear</li>\n<li>Additionally predict the player is in contact or not with any player</li>\n</ol>\n<h3>Others</h3>\n<p>As you can see there is no temporal information aggregation here.<br>\nI tried 2.5d or 3d modeling like other teams, but somehow failed.<br>\nOne of the reasons is that our team table feature engineering includes many roll, shift or diff features, so the benefit of temporal modeling was may be less than other teams.<br>\nAnd my model architecture is a way different from my teammates’ <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> models.<br>\nI guess this is one of the many reasons why we finished in the prize zone, 2 diverse cnn engines:)<br>\n<a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719</a></p>\n<hr>\n<h2>2nd stage</h2>\n<p>Other than the GBDT models, I built a simple 1d UNet as a 2nd stage model.   <br>\nThe main motivation of this model is to predict for no helmet player’s contact.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F98583c7e5063325426feb845193745f1%2F2nd_stage.png?generation=1677847508845712&amp;alt=media\" alt=\"\"></p>\n<p>It’s not so good model compared to LightGBM or XGBoost.(CV0.780/LB0.764)  <br>\nBut its prediction was very unique to other models, so it shined when ensembling.</p>\n<h2>Where I failed?</h2>\n<p>When we analyze failure cases, there are a lot of label noises.  <br>\nSo we wasted a lot of time cleaning up labels, by pseudo labeling, manual error removal and so on…   <br>\nBut all failed in overfitting to the validation set. No successful result in LB.  <br>\nI should have much more focus on pure CNN modeling, like other top teams!</p>",
      "rawMarkdown": "# Camaro part\n\n## 1st stage\n\nKind of Object Detection like archiarchitecture. Predict all combination pairs' contact and ground contact at once.  (Should I name as YOLO?)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F87f137cab92a5afa4827695aa2ff72ad%2F1st_stage.png?generation=1677847249397624&alt=media)\n\n### Details\n1. Extract feature map by Yolox FPN.\n2. RoIAlign around the player by using helmet coordinates.\n  a. Region size is adjusted by average helmet size in the frame\n  b. Helmet is located in the same position in the crop area\n3. Concatenate tracking features after linear transformation\n4. For inter contact\n  a. Creare pairwise player matrix and concat p1 and p2 features\n  b. Multiply distance features\n  c. Linear\n5. For ground contact, simply linear\n6. Additionally predict the player is in contact or not with any player\n\n### Others\nAs you can see there is no temporal information aggregation here.\nI tried 2.5d or 3d modeling like other teams, but somehow failed.\nOne of the reasons is that our team table feature engineering includes many roll, shift or diff features, so the benefit of temporal modeling was may be less than other teams.\nAnd my model architecture is a way different from my teammates’ @kmat2019 models.\nI guess this is one of the many reasons why we finished in the prize zone, 2 diverse cnn engines:)\nhttps://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719\n\n--- \n\n## 2nd stage\nOther than the GBDT models, I built a simple 1d UNet as a 2nd stage model.   \nThe main motivation of this model is to predict for no helmet player’s contact.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F98583c7e5063325426feb845193745f1%2F2nd_stage.png?generation=1677847508845712&alt=media)\n\nIt’s not so good model compared to LightGBM or XGBoost.(CV0.780/LB0.764)  \nBut its prediction was very unique to other models, so it shined when ensembling.\n\n## Where I failed?\nWhen we analyze failure cases, there are a lot of label noises.  \nSo we wasted a lot of time cleaning up labels, by pseudo labeling, manual error removal and so on…   \nBut all failed in overfitting to the validation set. No successful result in LB.  \nI should have much more focus on pure CNN modeling, like other top teams!",
      "votes": null
    },
    {
      "id": "2168918",
      "postDate": "03/04/2023 16:28:31",
      "content": "<p>Really great solution. Thanks for the writeup. It's interesting to see how you were able to ensemble individual team member's solutions so effectively.</p>\n<p>Do you know the CV score of the individual stage 1 models prior to stage 2? What type of improvement did you see from adding the 2nd stage with the tracking data?</p>\n<p>Did you have the same pipeline for P-P and P-G with the only difference being the difference in threshold?</p>\n<p>Appreciate you sharing this solution.</p>",
      "rawMarkdown": "Really great solution. Thanks for the writeup. It's interesting to see how you were able to ensemble individual team member's solutions so effectively.\n\nDo you know the CV score of the individual stage 1 models prior to stage 2? What type of improvement did you see from adding the 2nd stage with the tracking data?\n\nDid you have the same pipeline for P-P and P-G with the only difference being the difference in threshold?\n\nAppreciate you sharing this solution.",
      "votes": null
    },
    {
      "id": "2168921",
      "postDate": "03/04/2023 16:33:48",
      "content": "<p>Really interesting writeup on your stage 1 approach <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> - thanks for sharing. I'm not as familiar with Yolox FPN - was there a specific implementation you used? Was the feature map you extracted just for players?</p>\n<p>Great work.</p>",
      "rawMarkdown": "Really interesting writeup on your stage 1 approach @bamps53 - thanks for sharing. I'm not as familiar with Yolox FPN - was there a specific implementation you used? Was the feature map you extracted just for players?\n\nGreat work.",
      "votes": null
    },
    {
      "id": "2169274",
      "postDate": "03/05/2023 01:08:55",
      "content": "<p>Thanks! It's just a FPN which extracts 1/8 scale feature map. I used the one used in Yolox object detection model. Specifically this implementation.<br>\n<a href=\"https://github.com/Megvii-BaseDetection/YOLOX/blob/main/yolox/models/yolo_fpn.py\" target=\"_blank\">https://github.com/Megvii-BaseDetection/YOLOX/blob/main/yolox/models/yolo_fpn.py</a></p>\n<blockquote>\n  <p>Was the feature map you extracted just for players?  </p>\n</blockquote>\n<p>No, it's extracted for entire image.</p>",
      "rawMarkdown": "Thanks! It's just a FPN which extracts 1/8 scale feature map. I used the one used in Yolox object detection model. Specifically this implementation.\nhttps://github.com/Megvii-BaseDetection/YOLOX/blob/main/yolox/models/yolo_fpn.py\n\n> Was the feature map you extracted just for players?  \n\nNo, it's extracted for entire image.",
      "votes": null
    },
    {
      "id": "2169435",
      "postDate": "03/05/2023 05:33:38",
      "content": "<p>Since we started to team up at a very early stage of the competition, our solution is designed to be ensembled later. (ex. 2nd stage LGBM takes cnn prediction outputs and it can be null, cnn doesn’t care about missing helmet case.)<br>\nBut about my 1st stage CNN models typically score only around cv0.750~0.760. And if I feed it to LGBM with tabular data, it is boosted around cv0.780~0.790 area.</p>\n<p>As for my NN 2nd stage model, it is actually separately trained for P-P and P-G. Other GBDT model doesn’t have any difference other than threshold, it’s better than building separate models somehow.</p>",
      "rawMarkdown": "Since we started to team up at a very early stage of the competition, our solution is designed to be ensembled later. (ex. 2nd stage LGBM takes cnn prediction outputs and it can be null, cnn doesn’t care about missing helmet case.)\nBut about my 1st stage CNN models typically score only around cv0.750~0.760. And if I feed it to LGBM with tabular data, it is boosted around cv0.780~0.790 area.\n\nAs for my NN 2nd stage model, it is actually separately trained for P-P and P-G. Other GBDT model doesn’t have any difference other than threshold, it’s better than building separate models somehow.",
      "votes": null
    },
    {
      "id": "2170466",
      "postDate": "03/06/2023 03:02:19",
      "content": "<p>many thanks! waiting for your camaro model~</p>",
      "rawMarkdown": "many thanks! waiting for your camaro model~",
      "votes": null
    },
    {
      "id": "2170537",
      "postDate": "03/06/2023 04:45:59",
      "content": "<p>Thank you for your comment. Camaro model detail is written in comment.</p>",
      "rawMarkdown": "Thank you for your comment. Camaro model detail is written in comment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2167364,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "03/03/2023 12:51:16",
      "content": "<h1>Camaro part</h1>\n<h2>1st stage</h2>\n<p>Kind of Object Detection like archiarchitecture. Predict all combination pairs' contact and ground contact at once.  (Should I name as YOLO?)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F87f137cab92a5afa4827695aa2ff72ad%2F1st_stage.png?generation=1677847249397624&amp;alt=media\" alt=\"\"></p>\n<h3>Details</h3>\n<ol>\n<li>Extract feature map by Yolox FPN.</li>\n<li>RoIAlign around the player by using helmet coordinates.<br>\na. Region size is adjusted by average helmet size in the frame<br>\nb. Helmet is located in the same position in the crop area</li>\n<li>Concatenate tracking features after linear transformation</li>\n<li>For inter contact<br>\na. Creare pairwise player matrix and concat p1 and p2 features<br>\nb. Multiply distance features<br>\nc. Linear</li>\n<li>For ground contact, simply linear</li>\n<li>Additionally predict the player is in contact or not with any player</li>\n</ol>\n<h3>Others</h3>\n<p>As you can see there is no temporal information aggregation here.<br>\nI tried 2.5d or 3d modeling like other teams, but somehow failed.<br>\nOne of the reasons is that our team table feature engineering includes many roll, shift or diff features, so the benefit of temporal modeling was may be less than other teams.<br>\nAnd my model architecture is a way different from my teammates’ <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> models.<br>\nI guess this is one of the many reasons why we finished in the prize zone, 2 diverse cnn engines:)<br>\n<a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719</a></p>\n<hr>\n<h2>2nd stage</h2>\n<p>Other than the GBDT models, I built a simple 1d UNet as a 2nd stage model.   <br>\nThe main motivation of this model is to predict for no helmet player’s contact.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F98583c7e5063325426feb845193745f1%2F2nd_stage.png?generation=1677847508845712&amp;alt=media\" alt=\"\"></p>\n<p>It’s not so good model compared to LightGBM or XGBoost.(CV0.780/LB0.764)  <br>\nBut its prediction was very unique to other models, so it shined when ensembling.</p>\n<h2>Where I failed?</h2>\n<p>When we analyze failure cases, there are a lot of label noises.  <br>\nSo we wasted a lot of time cleaning up labels, by pseudo labeling, manual error removal and so on…   <br>\nBut all failed in overfitting to the validation set. No successful result in LB.  <br>\nI should have much more focus on pure CNN modeling, like other top teams!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2168921,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/04/2023 16:33:48",
          "content": "<p>Really interesting writeup on your stage 1 approach <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> - thanks for sharing. I'm not as familiar with Yolox FPN - was there a specific implementation you used? Was the feature map you extracted just for players?</p>\n<p>Great work.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2169274,
              "author_name": "bamps53",
              "author_url": "",
              "post_date": "03/05/2023 01:08:55",
              "content": "<p>Thanks! It's just a FPN which extracts 1/8 scale feature map. I used the one used in Yolox object detection model. Specifically this implementation.<br>\n<a href=\"https://github.com/Megvii-BaseDetection/YOLOX/blob/main/yolox/models/yolo_fpn.py\" target=\"_blank\">https://github.com/Megvii-BaseDetection/YOLOX/blob/main/yolox/models/yolo_fpn.py</a></p>\n<blockquote>\n  <p>Was the feature map you extracted just for players?  </p>\n</blockquote>\n<p>No, it's extracted for entire image.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2168918,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:28:31",
      "content": "<p>Really great solution. Thanks for the writeup. It's interesting to see how you were able to ensemble individual team member's solutions so effectively.</p>\n<p>Do you know the CV score of the individual stage 1 models prior to stage 2? What type of improvement did you see from adding the 2nd stage with the tracking data?</p>\n<p>Did you have the same pipeline for P-P and P-G with the only difference being the difference in threshold?</p>\n<p>Appreciate you sharing this solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169435,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "03/05/2023 05:33:38",
          "content": "<p>Since we started to team up at a very early stage of the competition, our solution is designed to be ensembled later. (ex. 2nd stage LGBM takes cnn prediction outputs and it can be null, cnn doesn’t care about missing helmet case.)<br>\nBut about my 1st stage CNN models typically score only around cv0.750~0.760. And if I feed it to LGBM with tabular data, it is boosted around cv0.780~0.790 area.</p>\n<p>As for my NN 2nd stage model, it is actually separately trained for P-P and P-G. Other GBDT model doesn’t have any difference other than threshold, it’s better than building separate models somehow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2170466,
      "author_name": "chg0901",
      "author_url": "",
      "post_date": "03/06/2023 03:02:19",
      "content": "<p>many thanks! waiting for your camaro model~</p>",
      "votes": null,
      "replies": [
        {
          "id": 2170537,
          "author_name": "hattan0523",
          "author_url": "",
          "post_date": "03/06/2023 04:45:59",
          "content": "<p>Thank you for your comment. Camaro model detail is written in comment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2166037": "We really appreciated the hosts and the kaggle team for organizing the competition. Moreover, we would also like to thank all the participants who joined. We could enjoy this competition and write up our solutions. \n\nI would like to thank team members, @bamps53, @nyanpn and @kmat2019, who have the top talent to analyze the task. I could discuss and enjoy the competition. \n\n# Overview\nSimple solution outline is attached pic.\n[![pipeline.png](https://i.postimg.cc/pLQ1qbCW/pipeline.png)](https://postimg.cc/VJ6Rkh2p)\n\nIn the 1st stage we predict the contact by multiple CNN. In the 2nd stage CNN prediction(s), tracking and helmet data are aggregated and created features to input GDBT.  Lastly we compute 5 models averaged value and optimize threshold for both player-player and player-ground contact.\n \n\n# 1st stage CNN\n## k mat model\nDetails are written in https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719.\nWe can obtain both Endzone and Sideline prediction values. \n\n## camaro model\nwill come up soon\n\n# 2nd stage aggregation & binary classification models\n\nWe excluded player-player pairs with distance > 3, and the remaining ~880K rows were used to train 2nd stage models. During inference time, we assigned 0 to pair with distance > 3 and predicted only the remaining data.\n\n## Created features\nBecause our CNN predictions are so strong, more than 90% of the top 30 important features were CNN-related features. Below are part of the features we have created.\n### Tracking \n - distance between two players\n - distance/x_position/y_position from step0\n - distance from around player (full/same team/different team )\n - distance between team center\n - distance to second nearest player\n - current step / max step\n - lag / lead of acc, speed, sa etc\n - max/min/mean of x, y, speed, acc, sa, distance group by (play, step), (play, step, team) and (play, player1, player2) x/y positon diff from step=0\n - ”interceptor” features\n  - find playerC who meet the following conditions and add distance(A, C) and ∠BAC to the features of playerA-playerB (to detect that C intercepts between A-B)\n     - ∠BAC < 30deg\n     - distance(A, C) < distance(A, B) and distance(B, C) < distance(A, B)\n### Helmet\n - bbox aspect ratio\n - bbox overlap\n - lag / lead of bbox coordinates\n - bbox center x,y std/shift/diff\n - distance of bbox centers\n### CNN prediction and  meta-features\n - oof predictions of 1st stage CNNs\n - max/min/std of predictions group by (play, step) and (play, player1, player2)\n - 5/11/21 rolling features\n  - to complement CNN predictions on frames without helmets\n - lag / diff\n - around players’ player-ground prediction value\n#### Combinations\n - registration errors from helmet-tracking coordinate transform (similar to 6th place solution, and previous NFL’s 1st place solution by K_mat)\n\n### Models\nWe trained four GBDT models with different combinations of 1st stage CNNs. We also added one NN model (\"camaro2\" in the figure above) and calculated the simple average of these 5 models. Predictions were binarized with separate thresholds optimized for player-player and player-ground respectively.\n\n - LightGBM\n   - K_mat A + Camaro1 Public 0.795/Private 0.792\n   - K_mat B + Camaro 1\n   - K_mat B\n - xgboost\n   - K_mat B + Camaro 1\n - Camaro 2\n### tips\n- rolling features for CNN prediction values are most important in our models.\n- judging from permutation feature importance, ‘minimum distance between players in the game_play’,  ‘distance between away team mean and home team mean’ and ‘player-player distance’ are important tracking features to increase score.\n- We did not use early-stopping to train the GBDTs because the optimal number of rounds for MCC is always longer than AUC.\n\n### not wroked for models\n- Catboost\n- Residual fit\n- Meta Features by non CNN (e.g. logistic regression prediction values/ k-means clustering feature)\n- Separate player-player and player-ground model\n- 1DCNN\n- External NFL data\n- Focal loss\n\n\n\n\n\n# not worked for overall\n- Adding previous competition pseudo labeling data\n- Removing noisy label\n- all29 assignment and its prediction\n- 2.5D or 3D CNN, but should have dug more..\n- Aggregate near frame information",
    "2167364": "# Camaro part\n\n## 1st stage\n\nKind of Object Detection like archiarchitecture. Predict all combination pairs' contact and ground contact at once.  (Should I name as YOLO?)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F87f137cab92a5afa4827695aa2ff72ad%2F1st_stage.png?generation=1677847249397624&alt=media)\n\n### Details\n1. Extract feature map by Yolox FPN.\n2. RoIAlign around the player by using helmet coordinates.\n  a. Region size is adjusted by average helmet size in the frame\n  b. Helmet is located in the same position in the crop area\n3. Concatenate tracking features after linear transformation\n4. For inter contact\n  a. Creare pairwise player matrix and concat p1 and p2 features\n  b. Multiply distance features\n  c. Linear\n5. For ground contact, simply linear\n6. Additionally predict the player is in contact or not with any player\n\n### Others\nAs you can see there is no temporal information aggregation here.\nI tried 2.5d or 3d modeling like other teams, but somehow failed.\nOne of the reasons is that our team table feature engineering includes many roll, shift or diff features, so the benefit of temporal modeling was may be less than other teams.\nAnd my model architecture is a way different from my teammates’ @kmat2019 models.\nI guess this is one of the many reasons why we finished in the prize zone, 2 diverse cnn engines:)\nhttps://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391719\n\n--- \n\n## 2nd stage\nOther than the GBDT models, I built a simple 1d UNet as a 2nd stage model.   \nThe main motivation of this model is to predict for no helmet player’s contact.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554318%2F98583c7e5063325426feb845193745f1%2F2nd_stage.png?generation=1677847508845712&alt=media)\n\nIt’s not so good model compared to LightGBM or XGBoost.(CV0.780/LB0.764)  \nBut its prediction was very unique to other models, so it shined when ensembling.\n\n## Where I failed?\nWhen we analyze failure cases, there are a lot of label noises.  \nSo we wasted a lot of time cleaning up labels, by pseudo labeling, manual error removal and so on…   \nBut all failed in overfitting to the validation set. No successful result in LB.  \nI should have much more focus on pure CNN modeling, like other top teams!",
    "2168918": "Really great solution. Thanks for the writeup. It's interesting to see how you were able to ensemble individual team member's solutions so effectively.\n\nDo you know the CV score of the individual stage 1 models prior to stage 2? What type of improvement did you see from adding the 2nd stage with the tracking data?\n\nDid you have the same pipeline for P-P and P-G with the only difference being the difference in threshold?\n\nAppreciate you sharing this solution.",
    "2168921": "Really interesting writeup on your stage 1 approach @bamps53 - thanks for sharing. I'm not as familiar with Yolox FPN - was there a specific implementation you used? Was the feature map you extracted just for players?\n\nGreat work.",
    "2169274": "Thanks! It's just a FPN which extracts 1/8 scale feature map. I used the one used in Yolox object detection model. Specifically this implementation.\nhttps://github.com/Megvii-BaseDetection/YOLOX/blob/main/yolox/models/yolo_fpn.py\n\n> Was the feature map you extracted just for players?  \n\nNo, it's extracted for entire image.",
    "2169435": "Since we started to team up at a very early stage of the competition, our solution is designed to be ensembled later. (ex. 2nd stage LGBM takes cnn prediction outputs and it can be null, cnn doesn’t care about missing helmet case.)\nBut about my 1st stage CNN models typically score only around cv0.750~0.760. And if I feed it to LGBM with tabular data, it is boosted around cv0.780~0.790 area.\n\nAs for my NN 2nd stage model, it is actually separately trained for P-P and P-G. Other GBDT model doesn’t have any difference other than threshold, it’s better than building separate models somehow.",
    "2170466": "many thanks! waiting for your camaro model~",
    "2170537": "Thank you for your comment. Camaro model detail is written in comment."
  },
  "source": "meta"
}