{
  "id": 391620,
  "title": "6th Place Solution (TK&penguin46 part)",
  "url": "/competitions/nfl-player-contact-detection/discussion/391620",
  "author_name": "",
  "post_date": "2023-03-02T01:54:32.257799100Z",
  "votes": 37,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First, I would like to thank the hosts for organizing the competition.<br>\nCongratulations to the winning teams. Thanks also to <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a>, <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>, and <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> for competing with me.</p>\n<p>In this post, we will explain TK(@tanakar) &amp; penguin46 part of the whole solution.<br>\n[update] Qishen &amp; Bo part was posted <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723\" target=\"_blank\">here</a></p>\n<h1>Overview</h1>\n<p>Features created using train.csv, tracking.csv, bbox.csv and pretrained models (yolov7, mmpose) for the video data is input to xgboost along with predictions of the cnn part. CV with csv data only is 0.766, CV with pretrained models added is 0.774, and 0.7955 with the CNN part added.</p>\n<h1>CNN Part</h1>\n<p>2.5d CNN is used for contact prediction. In this part, <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> in especially made a great contribution.</p>\n<h2>Input Image</h2>\n<p>Grayscale images are cropped around the helmet's bbox and resized to 224x224.<br>\nInputs are different for player &amp; player contacts and player &amp; ground contacts.</p>\n<ul>\n<li>player &amp; player<ul>\n<li>crop range up to 2 times the rectangle surrounding the two bboxes</li>\n<li>input bbox mask as an additional channel</li>\n<li>±2step</li></ul></li>\n<li>player &amp; ground<ul>\n<li>crop 6x around the bbox</li>\n<li>no mask</li>\n<li>±2step<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F4a505595b255bfda550e795046c0df38%2Fcnn.png?generation=1677721348768987&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Downsampling</h2>\n<ul>\n<li>player &amp; player<ul>\n<li>use samples with predictions of 1st stage lightgbm greater than 0.01</li></ul></li>\n<li>player &amp; ground<ul>\n<li>remove samples with predictions of Qishen&amp;Bo’s stage1 less than 0.005</li></ul></li>\n</ul>\n<h2>Backbones</h2>\n<ul>\n<li>player &amp; player<ul>\n<li>resnet18d</li>\n<li>swin_tiny_patch4_window7_224</li>\n<li>tf_efficientnet_b0_ns</li></ul></li>\n<li>player &amp; ground<ul>\n<li>resnet18d</li>\n<li>swin_tiny_patch4_window7_224</li>\n<li>efficientnetv2_rw_t</li></ul></li>\n</ul>\n<p>The simple average of multiple backbones is used as the features of next xgb. </p>\n<h1>Table &amp; Pretrained Models Part</h1>\n<h2>Player Detection</h2>\n<p>Player was detected using pretrained YOLOv7 and detected boxes were tracked using <a href=\"https://github.com/AlbertoSabater/Robust-and-efficient-post-processing-for-video-object-detection\" target=\"_blank\">REPP</a>, and matched with helmets. We inferred in steps instead of frames to speed up the submission time. In the matching process, bboxes were extracted in the order of the longest tracked length and greedily matched to the helmets of the players that matched best. The features created are as follows</p>\n<ul>\n<li>size of the bbox</li>\n<li>aspect ratio of the bbox</li>\n<li>distance between the center coordinates of the two players' bboxes (dx, dy, dist, dist/width)</li>\n<li>IoU of the bbox of the two players</li>\n<li>area of overlap with the bbox of the helmet</li>\n<li>tracked length</li>\n<li>matching score (sum of overlapped area)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F929fef6a9bf240a7b4faa87bb38e8d9f%2Ftracking.png?generation=1677721445396871&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Pose Estimation</h2>\n<p>We used mmpose to estimate the players' posture. mmpose saved inference time by using the players' bboxes that were object detected by yolov7. In addition, inference was performed only once in 2steps because this part was not significant for the computation time.</p>\n<ul>\n<li>distance between head and foot, hand and foot</li>\n<li>knee, hip angle</li>\n<li>number of keypoints detected and sum of confiences</li>\n<li>distance between head and center of helmet bbox</li>\n</ul>\n<h2>Coordinates Gap Between Field and Camera View</h2>\n<p>tracking.csv does not include a vertical component. Optimize the homographic transformation from field coordinates to camera coordinates and calculate the displacement of the helmet position. This misalignment can be attributed to the misalignment between the measured position of the field coordinates and the position of the helmet, and the vertical coordinates of the player, from which we can extract the vertical component. This idea is based on <a href=\"https://www.kaggle.com/competitions/nfl-health-and-safety-helmet-assignment/discussion/285112\" target=\"_blank\">this solution</a> by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> from the last year's solution. The homography transformation was optimized using cv2.findHomography.</p>\n<ul>\n<li>Coordinates (x, y) after projection</li>\n<li>Number of players used to optimize the homography matrix</li>\n<li>Distance between the coordinates after projection and the center coordinates of the helmet (dx, dy, dist, dx/width, dy/width dist/width)</li>\n<li>Total distance between the post-projection coordinates and the center coordinates of the helmet in frames<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F0b6d775701b80ea65fb25115692ed5e4%2Fgap.png?generation=1677721429648162&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Other Features</h2>\n<ul>\n<li>step, max(step), step / max(step)</li>\n<li>amount of movement from the start/end of the match (dx, dy, dist)</li>\n<li>distance to the nth nearest player (n=1, 2, 5, 10, 15)</li>\n<li>distance from the center of gravity coordinates of all players (dx, dy, dist)</li>\n<li>difference, product, and sum of velocity, acceleration, distance, and angle of two players</li>\n<li>aspect ratio of the bboxes of the helmets</li>\n<li>distance between the center coordinates of the bboxes of the two helmets</li>\n<li>IoU of the tow helmets</li>\n<li>position</li>\n<li>same team or not (0/1)</li>\n<li>lag features for the most important features among the above (diff, shift, lag=±1, 2, 5, 10, 20)</li>\n</ul>\n<h1>Other Tips</h1>\n<ul>\n<li>To make the experiment more efficient, we trained lightgbm, which predicts contact faster with fewer features, and excluded samples that were obviously unnecessary from penguin&amp;tk pipeline.</li>\n<li>xgboost was about 0.002~4 better than lightgbm.</li>\n<li>When using cutoff with qishen&amp;bo's predictions, add a mask feature to shows it has been cutoff.</li>\n</ul>\n<h1>Not Worked</h1>\n<ul>\n<li>LSTM, Transformer, 1DCNN</li>\n<li>data augmentation by swapping nfl_player_id_1, nfl_player_id_2 in table part</li>\n<li>instance segmentation</li>\n<li>seed averaging improved by less than 0.001</li>\n</ul>",
  "messages": [
    {
      "id": "2165106",
      "postDate": "03/02/2023 01:54:32",
      "content": "<p>First, I would like to thank the hosts for organizing the competition.<br>\nCongratulations to the winning teams. Thanks also to <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a>, <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>, and <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> for competing with me.</p>\n<p>In this post, we will explain TK(@tanakar) &amp; penguin46 part of the whole solution.<br>\n[update] Qishen &amp; Bo part was posted <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723\" target=\"_blank\">here</a></p>\n<h1>Overview</h1>\n<p>Features created using train.csv, tracking.csv, bbox.csv and pretrained models (yolov7, mmpose) for the video data is input to xgboost along with predictions of the cnn part. CV with csv data only is 0.766, CV with pretrained models added is 0.774, and 0.7955 with the CNN part added.</p>\n<h1>CNN Part</h1>\n<p>2.5d CNN is used for contact prediction. In this part, <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> in especially made a great contribution.</p>\n<h2>Input Image</h2>\n<p>Grayscale images are cropped around the helmet's bbox and resized to 224x224.<br>\nInputs are different for player &amp; player contacts and player &amp; ground contacts.</p>\n<ul>\n<li>player &amp; player<ul>\n<li>crop range up to 2 times the rectangle surrounding the two bboxes</li>\n<li>input bbox mask as an additional channel</li>\n<li>±2step</li></ul></li>\n<li>player &amp; ground<ul>\n<li>crop 6x around the bbox</li>\n<li>no mask</li>\n<li>±2step<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F4a505595b255bfda550e795046c0df38%2Fcnn.png?generation=1677721348768987&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Downsampling</h2>\n<ul>\n<li>player &amp; player<ul>\n<li>use samples with predictions of 1st stage lightgbm greater than 0.01</li></ul></li>\n<li>player &amp; ground<ul>\n<li>remove samples with predictions of Qishen&amp;Bo’s stage1 less than 0.005</li></ul></li>\n</ul>\n<h2>Backbones</h2>\n<ul>\n<li>player &amp; player<ul>\n<li>resnet18d</li>\n<li>swin_tiny_patch4_window7_224</li>\n<li>tf_efficientnet_b0_ns</li></ul></li>\n<li>player &amp; ground<ul>\n<li>resnet18d</li>\n<li>swin_tiny_patch4_window7_224</li>\n<li>efficientnetv2_rw_t</li></ul></li>\n</ul>\n<p>The simple average of multiple backbones is used as the features of next xgb. </p>\n<h1>Table &amp; Pretrained Models Part</h1>\n<h2>Player Detection</h2>\n<p>Player was detected using pretrained YOLOv7 and detected boxes were tracked using <a href=\"https://github.com/AlbertoSabater/Robust-and-efficient-post-processing-for-video-object-detection\" target=\"_blank\">REPP</a>, and matched with helmets. We inferred in steps instead of frames to speed up the submission time. In the matching process, bboxes were extracted in the order of the longest tracked length and greedily matched to the helmets of the players that matched best. The features created are as follows</p>\n<ul>\n<li>size of the bbox</li>\n<li>aspect ratio of the bbox</li>\n<li>distance between the center coordinates of the two players' bboxes (dx, dy, dist, dist/width)</li>\n<li>IoU of the bbox of the two players</li>\n<li>area of overlap with the bbox of the helmet</li>\n<li>tracked length</li>\n<li>matching score (sum of overlapped area)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F929fef6a9bf240a7b4faa87bb38e8d9f%2Ftracking.png?generation=1677721445396871&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Pose Estimation</h2>\n<p>We used mmpose to estimate the players' posture. mmpose saved inference time by using the players' bboxes that were object detected by yolov7. In addition, inference was performed only once in 2steps because this part was not significant for the computation time.</p>\n<ul>\n<li>distance between head and foot, hand and foot</li>\n<li>knee, hip angle</li>\n<li>number of keypoints detected and sum of confiences</li>\n<li>distance between head and center of helmet bbox</li>\n</ul>\n<h2>Coordinates Gap Between Field and Camera View</h2>\n<p>tracking.csv does not include a vertical component. Optimize the homographic transformation from field coordinates to camera coordinates and calculate the displacement of the helmet position. This misalignment can be attributed to the misalignment between the measured position of the field coordinates and the position of the helmet, and the vertical coordinates of the player, from which we can extract the vertical component. This idea is based on <a href=\"https://www.kaggle.com/competitions/nfl-health-and-safety-helmet-assignment/discussion/285112\" target=\"_blank\">this solution</a> by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> from the last year's solution. The homography transformation was optimized using cv2.findHomography.</p>\n<ul>\n<li>Coordinates (x, y) after projection</li>\n<li>Number of players used to optimize the homography matrix</li>\n<li>Distance between the coordinates after projection and the center coordinates of the helmet (dx, dy, dist, dx/width, dy/width dist/width)</li>\n<li>Total distance between the post-projection coordinates and the center coordinates of the helmet in frames<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F0b6d775701b80ea65fb25115692ed5e4%2Fgap.png?generation=1677721429648162&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Other Features</h2>\n<ul>\n<li>step, max(step), step / max(step)</li>\n<li>amount of movement from the start/end of the match (dx, dy, dist)</li>\n<li>distance to the nth nearest player (n=1, 2, 5, 10, 15)</li>\n<li>distance from the center of gravity coordinates of all players (dx, dy, dist)</li>\n<li>difference, product, and sum of velocity, acceleration, distance, and angle of two players</li>\n<li>aspect ratio of the bboxes of the helmets</li>\n<li>distance between the center coordinates of the bboxes of the two helmets</li>\n<li>IoU of the tow helmets</li>\n<li>position</li>\n<li>same team or not (0/1)</li>\n<li>lag features for the most important features among the above (diff, shift, lag=±1, 2, 5, 10, 20)</li>\n</ul>\n<h1>Other Tips</h1>\n<ul>\n<li>To make the experiment more efficient, we trained lightgbm, which predicts contact faster with fewer features, and excluded samples that were obviously unnecessary from penguin&amp;tk pipeline.</li>\n<li>xgboost was about 0.002~4 better than lightgbm.</li>\n<li>When using cutoff with qishen&amp;bo's predictions, add a mask feature to shows it has been cutoff.</li>\n</ul>\n<h1>Not Worked</h1>\n<ul>\n<li>LSTM, Transformer, 1DCNN</li>\n<li>data augmentation by swapping nfl_player_id_1, nfl_player_id_2 in table part</li>\n<li>instance segmentation</li>\n<li>seed averaging improved by less than 0.001</li>\n</ul>",
      "rawMarkdown": "First, I would like to thank the hosts for organizing the competition.\nCongratulations to the winning teams. Thanks also to @tanakar, @haqishen, and @boliu0 for competing with me.\n\nIn this post, we will explain TK(@tanakar) & penguin46 part of the whole solution.\n[update] Qishen & Bo part was posted [here](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723)\n\n# Overview\nFeatures created using train.csv, tracking.csv, bbox.csv and pretrained models (yolov7, mmpose) for the video data is input to xgboost along with predictions of the cnn part. CV with csv data only is 0.766, CV with pretrained models added is 0.774, and 0.7955 with the CNN part added.\n\n# CNN Part\n2.5d CNN is used for contact prediction. In this part, @tanakar in especially made a great contribution.\n\n## Input Image\nGrayscale images are cropped around the helmet's bbox and resized to 224x224.\nInputs are different for player & player contacts and player & ground contacts.\n- player & player\n  - crop range up to 2 times the rectangle surrounding the two bboxes\n  - input bbox mask as an additional channel\n  - ±2step\n- player & ground\n  - crop 6x around the bbox\n  - no mask\n  - ±2step\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F4a505595b255bfda550e795046c0df38%2Fcnn.png?generation=1677721348768987&alt=media)\n\n\n## Downsampling\n- player & player\n  - use samples with predictions of 1st stage lightgbm greater than 0.01\n- player & ground\n  - remove samples with predictions of Qishen&Bo’s stage1 less than 0.005\n\n## Backbones\n- player & player\n  - resnet18d\n  - swin_tiny_patch4_window7_224\n  - tf_efficientnet_b0_ns\n- player & ground\n  - resnet18d\n  - swin_tiny_patch4_window7_224\n  - efficientnetv2_rw_t\n\nThe simple average of multiple backbones is used as the features of next xgb. \n\n# Table & Pretrained Models Part\n\n## Player Detection\nPlayer was detected using pretrained YOLOv7 and detected boxes were tracked using [REPP](https://github.com/AlbertoSabater/Robust-and-efficient-post-processing-for-video-object-detection), and matched with helmets. We inferred in steps instead of frames to speed up the submission time. In the matching process, bboxes were extracted in the order of the longest tracked length and greedily matched to the helmets of the players that matched best. The features created are as follows\n- size of the bbox\n- aspect ratio of the bbox\n- distance between the center coordinates of the two players' bboxes (dx, dy, dist, dist/width)\n- IoU of the bbox of the two players\n- area of overlap with the bbox of the helmet\n- tracked length\n- matching score (sum of overlapped area)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F929fef6a9bf240a7b4faa87bb38e8d9f%2Ftracking.png?generation=1677721445396871&alt=media)\n\n## Pose Estimation\nWe used mmpose to estimate the players' posture. mmpose saved inference time by using the players' bboxes that were object detected by yolov7. In addition, inference was performed only once in 2steps because this part was not significant for the computation time.\n- distance between head and foot, hand and foot\n- knee, hip angle\n- number of keypoints detected and sum of confiences\n- distance between head and center of helmet bbox\n\n## Coordinates Gap Between Field and Camera View\ntracking.csv does not include a vertical component. Optimize the homographic transformation from field coordinates to camera coordinates and calculate the displacement of the helmet position. This misalignment can be attributed to the misalignment between the measured position of the field coordinates and the position of the helmet, and the vertical coordinates of the player, from which we can extract the vertical component. This idea is based on [this solution](https://www.kaggle.com/competitions/nfl-health-and-safety-helmet-assignment/discussion/285112) by @its7171 from the last year's solution. The homography transformation was optimized using cv2.findHomography.\n- Coordinates (x, y) after projection\n- Number of players used to optimize the homography matrix\n- Distance between the coordinates after projection and the center coordinates of the helmet (dx, dy, dist, dx/width, dy/width dist/width)\n- Total distance between the post-projection coordinates and the center coordinates of the helmet in frames\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F0b6d775701b80ea65fb25115692ed5e4%2Fgap.png?generation=1677721429648162&alt=media)\n\n## Other Features\n- step, max(step), step / max(step)\n- amount of movement from the start/end of the match (dx, dy, dist)\n- distance to the nth nearest player (n=1, 2, 5, 10, 15)\n- distance from the center of gravity coordinates of all players (dx, dy, dist)\n- difference, product, and sum of velocity, acceleration, distance, and angle of two players\n- aspect ratio of the bboxes of the helmets\n- distance between the center coordinates of the bboxes of the two helmets\n- IoU of the tow helmets\n- position\n- same team or not (0/1)\n- lag features for the most important features among the above (diff, shift, lag=±1, 2, 5, 10, 20)\n\n# Other Tips\n- To make the experiment more efficient, we trained lightgbm, which predicts contact faster with fewer features, and excluded samples that were obviously unnecessary from penguin&tk pipeline.\n- xgboost was about 0.002~4 better than lightgbm.\n- When using cutoff with qishen&bo's predictions, add a mask feature to shows it has been cutoff.\n\n# Not Worked\n- LSTM, Transformer, 1DCNN\n- data augmentation by swapping nfl_player_id_1, nfl_player_id_2 in table part\n- instance segmentation\n- seed averaging improved by less than 0.001",
      "votes": null
    },
    {
      "id": "2165297",
      "postDate": "03/02/2023 05:10:12",
      "content": "<p>Nice Detailed Sharing and Congratulations!</p>",
      "rawMarkdown": "Nice Detailed Sharing and Congratulations!",
      "votes": null
    },
    {
      "id": "2165496",
      "postDate": "03/02/2023 08:37:58",
      "content": "<p>congratulations! Excellent work. specially Not Worked part is very informative. if you don't mind can you please tell me which yolo architecture did you use? yolov7l or yolov7x. because i used yolov8x for contact detection. but when i did the submission it took me 7.5 hours to execute only this model(my idea is just efficiently crop the area with this mode). and feed this data TSM.</p>",
      "rawMarkdown": "congratulations! Excellent work. specially Not Worked part is very informative. if you don't mind can you please tell me which yolo architecture did you use? yolov7l or yolov7x. because i used yolov8x for contact detection. but when i did the submission it took me 7.5 hours to execute only this model(my idea is just efficiently crop the area with this mode). and feed this data TSM.",
      "votes": null
    },
    {
      "id": "2165536",
      "postDate": "03/02/2023 09:06:18",
      "content": "<p>thanks. we used <code>yolov7.pt</code> and the inference should be finished within 10min when submission. we also tried <code>yolov8x.pt</code>, but the time was not much different. please note that we input images in step units, not frames. </p>",
      "rawMarkdown": "thanks. we used `yolov7.pt` and the inference should be finished within 10min when submission. we also tried `yolov8x.pt`, but the time was not much different. please note that we input images in step units, not frames.",
      "votes": null
    },
    {
      "id": "2168951",
      "postDate": "03/04/2023 16:53:37",
      "content": "<p>Great great work! There is a lot of interesting research knowledge gained from the different approaches you took. I believe you are one of the few top teams that used pose estimation in your solution. Calculating the displacement of the helmet position vertically is a smart idea, I'd guess it was really helpful for identifying player-ground contact.</p>\n<p>Of all the additional features you created (specifically pose and helmet displacement) do you have any idea for which were the most helpful to the model?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Great great work! There is a lot of interesting research knowledge gained from the different approaches you took. I believe you are one of the few top teams that used pose estimation in your solution. Calculating the displacement of the helmet position vertically is a smart idea, I'd guess it was really helpful for identifying player-ground contact.\n\nOf all the additional features you created (specifically pose and helmet displacement) do you have any idea for which were the most helpful to the model?\n\nThanks!",
      "votes": null
    },
    {
      "id": "2169524",
      "postDate": "03/05/2023 07:26:06",
      "content": "<p>Thank you! We have not analyzed the contribution of the features in the final submission, but below are the changes in cv before and after adding the key features. I hope this will be useful.</p>\n<ul>\n<li>player detection &amp; tracking (769 -&gt; 775)</li>\n<li>helmet displacement (774 -&gt; 778~779)</li>\n<li>pose estimation (785 -&gt; 786)</li>\n</ul>",
      "rawMarkdown": "Thank you! We have not analyzed the contribution of the features in the final submission, but below are the changes in cv before and after adding the key features. I hope this will be useful.\n\n- player detection & tracking (769 -> 775)\n- helmet displacement (774 -> 778~779)\n- pose estimation (785 -> 786)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2165297,
      "author_name": "chg0901",
      "author_url": "",
      "post_date": "03/02/2023 05:10:12",
      "content": "<p>Nice Detailed Sharing and Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2165496,
      "author_name": "nadhirhasan",
      "author_url": "",
      "post_date": "03/02/2023 08:37:58",
      "content": "<p>congratulations! Excellent work. specially Not Worked part is very informative. if you don't mind can you please tell me which yolo architecture did you use? yolov7l or yolov7x. because i used yolov8x for contact detection. but when i did the submission it took me 7.5 hours to execute only this model(my idea is just efficiently crop the area with this mode). and feed this data TSM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2165536,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "03/02/2023 09:06:18",
          "content": "<p>thanks. we used <code>yolov7.pt</code> and the inference should be finished within 10min when submission. we also tried <code>yolov8x.pt</code>, but the time was not much different. please note that we input images in step units, not frames. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2168951,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:53:37",
      "content": "<p>Great great work! There is a lot of interesting research knowledge gained from the different approaches you took. I believe you are one of the few top teams that used pose estimation in your solution. Calculating the displacement of the helmet position vertically is a smart idea, I'd guess it was really helpful for identifying player-ground contact.</p>\n<p>Of all the additional features you created (specifically pose and helmet displacement) do you have any idea for which were the most helpful to the model?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169524,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "03/05/2023 07:26:06",
          "content": "<p>Thank you! We have not analyzed the contribution of the features in the final submission, but below are the changes in cv before and after adding the key features. I hope this will be useful.</p>\n<ul>\n<li>player detection &amp; tracking (769 -&gt; 775)</li>\n<li>helmet displacement (774 -&gt; 778~779)</li>\n<li>pose estimation (785 -&gt; 786)</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2165106": "First, I would like to thank the hosts for organizing the competition.\nCongratulations to the winning teams. Thanks also to @tanakar, @haqishen, and @boliu0 for competing with me.\n\nIn this post, we will explain TK(@tanakar) & penguin46 part of the whole solution.\n[update] Qishen & Bo part was posted [here](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723)\n\n# Overview\nFeatures created using train.csv, tracking.csv, bbox.csv and pretrained models (yolov7, mmpose) for the video data is input to xgboost along with predictions of the cnn part. CV with csv data only is 0.766, CV with pretrained models added is 0.774, and 0.7955 with the CNN part added.\n\n# CNN Part\n2.5d CNN is used for contact prediction. In this part, @tanakar in especially made a great contribution.\n\n## Input Image\nGrayscale images are cropped around the helmet's bbox and resized to 224x224.\nInputs are different for player & player contacts and player & ground contacts.\n- player & player\n  - crop range up to 2 times the rectangle surrounding the two bboxes\n  - input bbox mask as an additional channel\n  - ±2step\n- player & ground\n  - crop 6x around the bbox\n  - no mask\n  - ±2step\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F4a505595b255bfda550e795046c0df38%2Fcnn.png?generation=1677721348768987&alt=media)\n\n\n## Downsampling\n- player & player\n  - use samples with predictions of 1st stage lightgbm greater than 0.01\n- player & ground\n  - remove samples with predictions of Qishen&Bo’s stage1 less than 0.005\n\n## Backbones\n- player & player\n  - resnet18d\n  - swin_tiny_patch4_window7_224\n  - tf_efficientnet_b0_ns\n- player & ground\n  - resnet18d\n  - swin_tiny_patch4_window7_224\n  - efficientnetv2_rw_t\n\nThe simple average of multiple backbones is used as the features of next xgb. \n\n# Table & Pretrained Models Part\n\n## Player Detection\nPlayer was detected using pretrained YOLOv7 and detected boxes were tracked using [REPP](https://github.com/AlbertoSabater/Robust-and-efficient-post-processing-for-video-object-detection), and matched with helmets. We inferred in steps instead of frames to speed up the submission time. In the matching process, bboxes were extracted in the order of the longest tracked length and greedily matched to the helmets of the players that matched best. The features created are as follows\n- size of the bbox\n- aspect ratio of the bbox\n- distance between the center coordinates of the two players' bboxes (dx, dy, dist, dist/width)\n- IoU of the bbox of the two players\n- area of overlap with the bbox of the helmet\n- tracked length\n- matching score (sum of overlapped area)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F929fef6a9bf240a7b4faa87bb38e8d9f%2Ftracking.png?generation=1677721445396871&alt=media)\n\n## Pose Estimation\nWe used mmpose to estimate the players' posture. mmpose saved inference time by using the players' bboxes that were object detected by yolov7. In addition, inference was performed only once in 2steps because this part was not significant for the computation time.\n- distance between head and foot, hand and foot\n- knee, hip angle\n- number of keypoints detected and sum of confiences\n- distance between head and center of helmet bbox\n\n## Coordinates Gap Between Field and Camera View\ntracking.csv does not include a vertical component. Optimize the homographic transformation from field coordinates to camera coordinates and calculate the displacement of the helmet position. This misalignment can be attributed to the misalignment between the measured position of the field coordinates and the position of the helmet, and the vertical coordinates of the player, from which we can extract the vertical component. This idea is based on [this solution](https://www.kaggle.com/competitions/nfl-health-and-safety-helmet-assignment/discussion/285112) by @its7171 from the last year's solution. The homography transformation was optimized using cv2.findHomography.\n- Coordinates (x, y) after projection\n- Number of players used to optimize the homography matrix\n- Distance between the coordinates after projection and the center coordinates of the helmet (dx, dy, dist, dx/width, dy/width dist/width)\n- Total distance between the post-projection coordinates and the center coordinates of the helmet in frames\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F0b6d775701b80ea65fb25115692ed5e4%2Fgap.png?generation=1677721429648162&alt=media)\n\n## Other Features\n- step, max(step), step / max(step)\n- amount of movement from the start/end of the match (dx, dy, dist)\n- distance to the nth nearest player (n=1, 2, 5, 10, 15)\n- distance from the center of gravity coordinates of all players (dx, dy, dist)\n- difference, product, and sum of velocity, acceleration, distance, and angle of two players\n- aspect ratio of the bboxes of the helmets\n- distance between the center coordinates of the bboxes of the two helmets\n- IoU of the tow helmets\n- position\n- same team or not (0/1)\n- lag features for the most important features among the above (diff, shift, lag=±1, 2, 5, 10, 20)\n\n# Other Tips\n- To make the experiment more efficient, we trained lightgbm, which predicts contact faster with fewer features, and excluded samples that were obviously unnecessary from penguin&tk pipeline.\n- xgboost was about 0.002~4 better than lightgbm.\n- When using cutoff with qishen&bo's predictions, add a mask feature to shows it has been cutoff.\n\n# Not Worked\n- LSTM, Transformer, 1DCNN\n- data augmentation by swapping nfl_player_id_1, nfl_player_id_2 in table part\n- instance segmentation\n- seed averaging improved by less than 0.001",
    "2165297": "Nice Detailed Sharing and Congratulations!",
    "2165496": "congratulations! Excellent work. specially Not Worked part is very informative. if you don't mind can you please tell me which yolo architecture did you use? yolov7l or yolov7x. because i used yolov8x for contact detection. but when i did the submission it took me 7.5 hours to execute only this model(my idea is just efficiently crop the area with this mode). and feed this data TSM.",
    "2165536": "thanks. we used `yolov7.pt` and the inference should be finished within 10min when submission. we also tried `yolov8x.pt`, but the time was not much different. please note that we input images in step units, not frames.",
    "2168951": "Great great work! There is a lot of interesting research knowledge gained from the different approaches you took. I believe you are one of the few top teams that used pose estimation in your solution. Calculating the displacement of the helmet position vertically is a smart idea, I'd guess it was really helpful for identifying player-ground contact.\n\nOf all the additional features you created (specifically pose and helmet displacement) do you have any idea for which were the most helpful to the model?\n\nThanks!",
    "2169524": "Thank you! We have not analyzed the contribution of the features in the final submission, but below are the changes in cv before and after adding the key features. I hope this will be useful.\n\n- player detection & tracking (769 -> 775)\n- helmet displacement (774 -> 778~779)\n- pose estimation (785 -> 786)"
  },
  "source": "meta"
}