{
  "id": 392402,
  "title": "9th place solution - Team JK",
  "url": "/competitions/nfl-player-contact-detection/writeups/jk-9th-place-solution-team-jk",
  "author_name": "",
  "post_date": "2023-03-06T15:35:51.833Z",
  "votes": 32,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I would like to thank the organizers for such an interesting competition!   <br>\nWe share the Team JK's solution.  <br>\nTeam Member: <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>, <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>, <a href=\"https://www.kaggle.com/anonamename\" target=\"_blank\">@anonamename</a>, <a href=\"https://www.kaggle.com/yururoi\" target=\"_blank\">@yururoi</a>, <a href=\"https://www.kaggle.com/tomo20180402\" target=\"_blank\">@tomo20180402</a><br>\n<br></p>\n<h1>Overview</h1>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126876%2F4daeeaa70bf25eda3c4bb2519bcac346%2Fjk_solution_image.png?generation=1677991277683737&amp;alt=media\" alt=\"\"></p>\n<h1>1st stage</h1>\n<hr>\n<h2>yuki part</h2>\n<ul>\n<li>(1) of fig.</li>\n<li>See <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392046\" target=\"_blank\">yuki's post</a>.</li>\n</ul>\n<h2>Vladislav part</h2>\n<ul>\n<li>(2) of fig.</li>\n<li>Features are mainly created from sensor data, but helmets bboxes information is also used.</li>\n<li>Trained XGB and LGBM models for P2P and P2G individually.</li>\n<li>P2P and P2G have different features. There are 133 features for pair contact and 119 features for ground contact.</li>\n<li>Here is explanation of some features:<ul>\n<li>Excluded speed, since it correlates to distance.</li>\n<li>Step (or frame_number), it boosts score a lot.</li>\n<li>Player position on field (defense, offense or special)</li>\n<li>Twist feature (direction-orientation)</li>\n<li>Same team feature</li>\n<li>Is home team feature</li>\n<li>Number of players/opponents in (1,3,5 meters) is quite good feature</li>\n<li>Number of players in opposite orientation</li>\n<li>Acceleration of player ratio to mean acceleration of all players per step</li>\n<li>Diff of features of same player (in time domain)</li>\n<li>Time features (just copy of previous and future steps features)</li>\n<li>Difference of features between two players</li>\n<li>Euclidean distance is the main feature and other features based on it as well</li>\n<li>Features from helmets dataframe (bboxes coordinates, bboxes height &amp; width for each view and perimeter)</li>\n<li>IoU helmets features</li></ul></li>\n<li>XGB/LGBM models were trained with common hyperperameters that can be seen on public notebooks. Only added reg_alpha = 0.1 for both models.</li>\n</ul>\n<h2>anonamename part - combined knowledge of team members</h2>\n<ul>\n<li>(3) of fig.</li>\n<li>2-stage model of 2.5D/3D CNN and GBDT (5fold CV:0.778/Public:0.775/Private:0.773)</li>\n<li>2.5D/3D CNN<ul>\n<li>based <a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">public notebook</a>.</li>\n<li>input<ul>\n<li>image<ul>\n<li>15frames (±7frame, skip_frame=1)</li>\n<li>use both view (Endzone and Sideline)</li></ul></li>\n<li>tracking data<ul>\n<li>64 features (created by <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>)</li></ul></li></ul></li>\n<li>model<ul>\n<li>based <a href=\"https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932\" target=\"_blank\">DFL competition 1st solution</a>.</li>\n<li>pipeline : 15frames 2.5D -&gt; Residual3DBlock -&gt; GeM (created by <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>)</li>\n<li>2.5D backborn : tf_mobilenetv3_small_minimal_100.in1k</li>\n<li>multi-label classification (created by <a href=\"https://www.kaggle.com/anonamename\" target=\"_blank\">@anonamename</a>)<ul>\n<li>num_classes=2(Player-Player contact(P2P) and Player-Ground contact(P2G)) + nn.BCEWithLogitsLoss</li></ul></li>\n<li>fold : StratifiedGroupKFold(n_splits=5).split(y=\"contact_org\", groups=\"game_id\") (created by <a href=\"https://www.kaggle.com/tomo20180402\" target=\"_blank\">@tomo20180402</a>)<ul>\n<li>Set different labels for contacts between same team, different teams and ground.</li>\n<li>train data under sampling : positive:negative = 1:5 (change under sampling data for each epoch)</li></ul></li></ul></li>\n<li>optimaizer : AdamW(lr=1e-3-&gt;1e-5 CosineAnnealingLR, weight_decay=1e-5)</li>\n<li>epoch : 15</li>\n<li>augmentation<ul>\n<li>HorizontalFlip, ShiftScaleRotate, MotionBlur, OpticalDistortion, CoarseDropout</li>\n<li>Mixup at the last layer (like a <a href=\"https://arxiv.org/abs/1806.05236\" target=\"_blank\">Manifold mixup</a>. created by <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>)</li></ul></li>\n<li>TTA : HorizontalFlip</li></ul></li>\n<li>GBDT<ul>\n<li>Create xgboost and lightgbm for P2P and P2G individually.</li>\n<li>tracking feature (created by <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>)</li>\n<li>2.5D/3D CNN prob feature<ul>\n<li>groupby([\"game_play\", \"nfl_player_id_1\", \"nfl_player_id_2\"]) : shift(), diff(), mean(), max(), min(), std()</li></ul></li></ul></li>\n</ul>\n<h2>tomo part</h2>\n<ul>\n<li>(4) of fig.</li>\n<li>single-stage NN model (3fold CV:0.771/Public:0.759/Private:0.760)</li>\n<li>multi-class classification : P2P (same team), P2P (different team), P2G<ul>\n<li>output is 6 labels which are used as features of team's 2nd stage</li></ul></li>\n<li>execution time : 2h</li>\n<li>validation : StratifiedGroupKFold(n_splits=3).split(y=\"contact_org\", groups=\"game_id\")<ul>\n<li>same as anonamename part </li></ul></li>\n<li>dataset<ul>\n<li>train data under sampling : Reduce negative sample of P2P contact (same team) by one-third.</li></ul></li>\n<li>feature<ul>\n<li>table feature : 54<ul>\n<li>3 types distance : euclidean, chebyshev, cityblock</li>\n<li>3 types distance rank : among all, same team, different team</li>\n<li>median of helmet width and height</li>\n<li>normalized distance by mean of helmet width and height<ul>\n<li>The mean of helmet width and height are calculated from all players.</li></ul></li>\n<li>total rank from the center coordinates of 2player's helmets</li>\n<li>ratio of helmet detection exist : both players, each player</li>\n<li>cosine similarity : direction, orientation</li>\n<li>predicted euclidean distance</li>\n<li>other simple features : step, is_same_team, ground_flag, etc.</li></ul></li>\n<li>image feature<ul>\n<li>10 images in 2.5D CNN<ul>\n<li>5frames each for Sideline and Endline (n-4, n-2, n, n+2, n+4)</li>\n<li>image_size = (256, 256)</li></ul></li>\n<li>cropping method<ul>\n<li>Change the cropping method depending on whether both players’ helmets exist.<ul>\n<li>both players exist : Make sure both players are visible.</li>\n<li>one player exist : Make sure the player is in the center.</li></ul></li>\n<li>Crop the image with the mean of helmet width and height as a variable.</li>\n<li>Give priority to the downward direction.</li></ul></li>\n<li>mean of image exist : 4<ul>\n<li>each for Sideline and Endline</li></ul></li></ul></li></ul></li>\n<li>TTA<ul>\n<li>flip sensor and image in one of three models inferences<ul>\n<li>sensor : exchange player1,2</li>\n<li>image : HorizontalFlip</li></ul></li></ul></li>\n</ul>\n<h1>2nd stage</h1>\n<hr>\n<ul>\n<li>model : lgbm × 4</li>\n<li>feature : shift features of each models’ predictions and sensor data (-13~+13)</li>\n<li>postprocessing : 4 predictions by lgbm -&gt; simple average -&gt; moving average -&gt; final prediction</li>\n</ul>",
  "messages": [
    {
      "id": "2169464",
      "postDate": "03/05/2023 06:21:22",
      "content": "<p>I would like to thank the organizers for such an interesting competition!   <br>\nWe share the Team JK's solution.  <br>\nTeam Member: <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>, <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>, <a href=\"https://www.kaggle.com/anonamename\" target=\"_blank\">@anonamename</a>, <a href=\"https://www.kaggle.com/yururoi\" target=\"_blank\">@yururoi</a>, <a href=\"https://www.kaggle.com/tomo20180402\" target=\"_blank\">@tomo20180402</a><br>\n<br></p>\n<h1>Overview</h1>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126876%2F4daeeaa70bf25eda3c4bb2519bcac346%2Fjk_solution_image.png?generation=1677991277683737&amp;alt=media\" alt=\"\"></p>\n<h1>1st stage</h1>\n<hr>\n<h2>yuki part</h2>\n<ul>\n<li>(1) of fig.</li>\n<li>See <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392046\" target=\"_blank\">yuki's post</a>.</li>\n</ul>\n<h2>Vladislav part</h2>\n<ul>\n<li>(2) of fig.</li>\n<li>Features are mainly created from sensor data, but helmets bboxes information is also used.</li>\n<li>Trained XGB and LGBM models for P2P and P2G individually.</li>\n<li>P2P and P2G have different features. There are 133 features for pair contact and 119 features for ground contact.</li>\n<li>Here is explanation of some features:<ul>\n<li>Excluded speed, since it correlates to distance.</li>\n<li>Step (or frame_number), it boosts score a lot.</li>\n<li>Player position on field (defense, offense or special)</li>\n<li>Twist feature (direction-orientation)</li>\n<li>Same team feature</li>\n<li>Is home team feature</li>\n<li>Number of players/opponents in (1,3,5 meters) is quite good feature</li>\n<li>Number of players in opposite orientation</li>\n<li>Acceleration of player ratio to mean acceleration of all players per step</li>\n<li>Diff of features of same player (in time domain)</li>\n<li>Time features (just copy of previous and future steps features)</li>\n<li>Difference of features between two players</li>\n<li>Euclidean distance is the main feature and other features based on it as well</li>\n<li>Features from helmets dataframe (bboxes coordinates, bboxes height &amp; width for each view and perimeter)</li>\n<li>IoU helmets features</li></ul></li>\n<li>XGB/LGBM models were trained with common hyperperameters that can be seen on public notebooks. Only added reg_alpha = 0.1 for both models.</li>\n</ul>\n<h2>anonamename part - combined knowledge of team members</h2>\n<ul>\n<li>(3) of fig.</li>\n<li>2-stage model of 2.5D/3D CNN and GBDT (5fold CV:0.778/Public:0.775/Private:0.773)</li>\n<li>2.5D/3D CNN<ul>\n<li>based <a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">public notebook</a>.</li>\n<li>input<ul>\n<li>image<ul>\n<li>15frames (±7frame, skip_frame=1)</li>\n<li>use both view (Endzone and Sideline)</li></ul></li>\n<li>tracking data<ul>\n<li>64 features (created by <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>)</li></ul></li></ul></li>\n<li>model<ul>\n<li>based <a href=\"https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932\" target=\"_blank\">DFL competition 1st solution</a>.</li>\n<li>pipeline : 15frames 2.5D -&gt; Residual3DBlock -&gt; GeM (created by <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>)</li>\n<li>2.5D backborn : tf_mobilenetv3_small_minimal_100.in1k</li>\n<li>multi-label classification (created by <a href=\"https://www.kaggle.com/anonamename\" target=\"_blank\">@anonamename</a>)<ul>\n<li>num_classes=2(Player-Player contact(P2P) and Player-Ground contact(P2G)) + nn.BCEWithLogitsLoss</li></ul></li>\n<li>fold : StratifiedGroupKFold(n_splits=5).split(y=\"contact_org\", groups=\"game_id\") (created by <a href=\"https://www.kaggle.com/tomo20180402\" target=\"_blank\">@tomo20180402</a>)<ul>\n<li>Set different labels for contacts between same team, different teams and ground.</li>\n<li>train data under sampling : positive:negative = 1:5 (change under sampling data for each epoch)</li></ul></li></ul></li>\n<li>optimaizer : AdamW(lr=1e-3-&gt;1e-5 CosineAnnealingLR, weight_decay=1e-5)</li>\n<li>epoch : 15</li>\n<li>augmentation<ul>\n<li>HorizontalFlip, ShiftScaleRotate, MotionBlur, OpticalDistortion, CoarseDropout</li>\n<li>Mixup at the last layer (like a <a href=\"https://arxiv.org/abs/1806.05236\" target=\"_blank\">Manifold mixup</a>. created by <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>)</li></ul></li>\n<li>TTA : HorizontalFlip</li></ul></li>\n<li>GBDT<ul>\n<li>Create xgboost and lightgbm for P2P and P2G individually.</li>\n<li>tracking feature (created by <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>)</li>\n<li>2.5D/3D CNN prob feature<ul>\n<li>groupby([\"game_play\", \"nfl_player_id_1\", \"nfl_player_id_2\"]) : shift(), diff(), mean(), max(), min(), std()</li></ul></li></ul></li>\n</ul>\n<h2>tomo part</h2>\n<ul>\n<li>(4) of fig.</li>\n<li>single-stage NN model (3fold CV:0.771/Public:0.759/Private:0.760)</li>\n<li>multi-class classification : P2P (same team), P2P (different team), P2G<ul>\n<li>output is 6 labels which are used as features of team's 2nd stage</li></ul></li>\n<li>execution time : 2h</li>\n<li>validation : StratifiedGroupKFold(n_splits=3).split(y=\"contact_org\", groups=\"game_id\")<ul>\n<li>same as anonamename part </li></ul></li>\n<li>dataset<ul>\n<li>train data under sampling : Reduce negative sample of P2P contact (same team) by one-third.</li></ul></li>\n<li>feature<ul>\n<li>table feature : 54<ul>\n<li>3 types distance : euclidean, chebyshev, cityblock</li>\n<li>3 types distance rank : among all, same team, different team</li>\n<li>median of helmet width and height</li>\n<li>normalized distance by mean of helmet width and height<ul>\n<li>The mean of helmet width and height are calculated from all players.</li></ul></li>\n<li>total rank from the center coordinates of 2player's helmets</li>\n<li>ratio of helmet detection exist : both players, each player</li>\n<li>cosine similarity : direction, orientation</li>\n<li>predicted euclidean distance</li>\n<li>other simple features : step, is_same_team, ground_flag, etc.</li></ul></li>\n<li>image feature<ul>\n<li>10 images in 2.5D CNN<ul>\n<li>5frames each for Sideline and Endline (n-4, n-2, n, n+2, n+4)</li>\n<li>image_size = (256, 256)</li></ul></li>\n<li>cropping method<ul>\n<li>Change the cropping method depending on whether both players’ helmets exist.<ul>\n<li>both players exist : Make sure both players are visible.</li>\n<li>one player exist : Make sure the player is in the center.</li></ul></li>\n<li>Crop the image with the mean of helmet width and height as a variable.</li>\n<li>Give priority to the downward direction.</li></ul></li>\n<li>mean of image exist : 4<ul>\n<li>each for Sideline and Endline</li></ul></li></ul></li></ul></li>\n<li>TTA<ul>\n<li>flip sensor and image in one of three models inferences<ul>\n<li>sensor : exchange player1,2</li>\n<li>image : HorizontalFlip</li></ul></li></ul></li>\n</ul>\n<h1>2nd stage</h1>\n<hr>\n<ul>\n<li>model : lgbm × 4</li>\n<li>feature : shift features of each models’ predictions and sensor data (-13~+13)</li>\n<li>postprocessing : 4 predictions by lgbm -&gt; simple average -&gt; moving average -&gt; final prediction</li>\n</ul>",
      "rawMarkdown": "I would like to thank the organizers for such an interesting competition!   \nWe share the Team JK's solution.  \nTeam Member: @vostankovich, @tereka, @anonamename, @yururoi, @tomo20180402\n<br>\n\n# Overview\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126876%2F4daeeaa70bf25eda3c4bb2519bcac346%2Fjk_solution_image.png?generation=1677991277683737&alt=media)\n\n# 1st stage\n\n---\n\n## yuki part\n- (1) of fig.\n- See [yuki's post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392046).\n\n## Vladislav part\n- (2) of fig.\n- Features are mainly created from sensor data, but helmets bboxes information is also used.\n- Trained XGB and LGBM models for P2P and P2G individually.\n- P2P and P2G have different features. There are 133 features for pair contact and 119 features for ground contact.\n- Here is explanation of some features:\n    - Excluded speed, since it correlates to distance.\n    - Step (or frame_number), it boosts score a lot.\n    - Player position on field (defense, offense or special)\n    - Twist feature (direction-orientation)\n    - Same team feature\n    - Is home team feature\n    - Number of players/opponents in (1,3,5 meters) is quite good feature\n    - Number of players in opposite orientation\n    - Acceleration of player ratio to mean acceleration of all players per step\n    - Diff of features of same player (in time domain)\n    - Time features (just copy of previous and future steps features)\n    - Difference of features between two players\n    - Euclidean distance is the main feature and other features based on it as well\n    - Features from helmets dataframe (bboxes coordinates, bboxes height & width for each view and perimeter)\n    - IoU helmets features\n- XGB/LGBM models were trained with common hyperperameters that can be seen on public notebooks. Only added reg_alpha = 0.1 for both models.\n\n## anonamename part - combined knowledge of team members\n- (3) of fig.\n- 2-stage model of 2.5D/3D CNN and GBDT (5fold CV:0.778/Public:0.775/Private:0.773)\n- 2.5D/3D CNN\n    - based [public notebook](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference).\n    - input\n        - image\n            - 15frames (±7frame, skip_frame=1)\n            - use both view (Endzone and Sideline)\n        - tracking data\n            - 64 features (created by @vostankovich)\n    - model\n        - based [DFL competition 1st solution](https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932).\n        - pipeline : 15frames 2.5D -> Residual3DBlock -> GeM (created by @tereka)\n        - 2.5D backborn : tf_mobilenetv3_small_minimal_100.in1k\n        - multi-label classification (created by @anonamename)\n            - num_classes=2(Player-Player contact(P2P) and Player-Ground contact(P2G)) + nn.BCEWithLogitsLoss\n        - fold : StratifiedGroupKFold(n_splits=5).split(y=\"contact_org\", groups=\"game_id\") (created by @tomo20180402)\n            - Set different labels for contacts between same team, different teams and ground.\n            - train data under sampling : positive:negative = 1:5 (change under sampling data for each epoch)\n    - optimaizer : AdamW(lr=1e-3->1e-5 CosineAnnealingLR, weight_decay=1e-5)\n    - epoch : 15\n    - augmentation\n        - HorizontalFlip, ShiftScaleRotate, MotionBlur, OpticalDistortion, CoarseDropout\n        - Mixup at the last layer (like a [Manifold mixup](https://arxiv.org/abs/1806.05236). created by @tereka)\n    - TTA : HorizontalFlip\n- GBDT\n    - Create xgboost and lightgbm for P2P and P2G individually.\n    - tracking feature (created by @vostankovich)\n    - 2.5D/3D CNN prob feature\n        - groupby([\"game_play\", \"nfl_player_id_1\", \"nfl_player_id_2\"]) : shift(), diff(), mean(), max(), min(), std()\n\n## tomo part\n- (4) of fig.\n- single-stage NN model (3fold CV:0.771/Public:0.759/Private:0.760)\n- multi-class classification : P2P (same team), P2P (different team), P2G\n    - output is 6 labels which are used as features of team's 2nd stage\n- execution time : 2h\n- validation : StratifiedGroupKFold(n_splits=3).split(y=\"contact_org\", groups=\"game_id\")\n    - same as anonamename part \n- dataset\n    - train data under sampling : Reduce negative sample of P2P contact (same team) by one-third.\n- feature\n    - table feature : 54\n        - 3 types distance : euclidean, chebyshev, cityblock\n        - 3 types distance rank : among all, same team, different team\n        - median of helmet width and height\n        - normalized distance by mean of helmet width and height\n            - The mean of helmet width and height are calculated from all players.\n        - total rank from the center coordinates of 2player's helmets\n        - ratio of helmet detection exist : both players, each player\n        - cosine similarity : direction, orientation\n        - predicted euclidean distance\n        - other simple features : step, is_same_team, ground_flag, etc.\n    - image feature\n        - 10 images in 2.5D CNN\n            - 5frames each for Sideline and Endline (n-4, n-2, n, n+2, n+4)\n            - image_size = (256, 256)\n        - cropping method\n            - Change the cropping method depending on whether both players’ helmets exist.\n                - both players exist : Make sure both players are visible.\n                - one player exist : Make sure the player is in the center.\n            - Crop the image with the mean of helmet width and height as a variable.\n            - Give priority to the downward direction.\n        - mean of image exist : 4\n            - each for Sideline and Endline\n- TTA\n    - flip sensor and image in one of three models inferences\n        - sensor : exchange player1,2\n        - image : HorizontalFlip\n\n# 2nd stage\n\n---\n\n- model : lgbm × 4\n- feature : shift features of each models’ predictions and sensor data (-13~+13)\n- postprocessing : 4 predictions by lgbm -> simple average -> moving average -> final prediction",
      "votes": null
    },
    {
      "id": "2170010",
      "postDate": "03/05/2023 15:56:47",
      "content": "<p>I explain about Residual3DBlock.<br>\nIn this competition, 3d(t, h, w) is very important, so I decided to use any time analysis method.<br>\n15 frames reshape this ({-28, -24, -20}{-16, -12, -8}{-4, 0, 4}{8, 12, 16}{20, 24, 28}) and extract feature using backbone, then hidden output is applied it. here is a sample for using residual3d block(final architecture is very similar)</p>\n<pre><code>class Residual3DBlock(nn.Module):\n    def __init__(self):\n        super(Residual3DBlock, self).__init__()\n\n        self.block = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n            nn.ReLU(512)\n        )\n\n        self.block2 = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n        )\n\n    def forward(self, images):\n        short_cut = images\n        h = self.block(images)\n        h = self.block2(h)\n\n        return F.relu(h + short_cut)\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.backbone = timm.create_model(\"tf_efficientnet_b0_ns\", pretrained=True, num_classes=1, in_chans=3)\n        self.mlp = nn.Sequential(\n            nn.Linear(68, 256),\n            nn.BatchNorm1d(256),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n        n_hidden = 1024\n\n        self.conv_proj = nn.Sequential(\n            nn.Conv2d(1280, 512, 1, stride=1),\n            nn.BatchNorm2d(512),\n            nn.ReLU(),\n        )\n\n        self.neck = nn.Sequential(\n            nn.Linear(1024, 1024),\n            nn.BatchNorm1d(1024),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n\n        self.triple_layer = nn.Sequential(\n            Residual3DBlock(),\n        )\n\n        self.pool = GeM() \n\n        self.fc = nn.Linear(256+1024, 1)\n\n    def forward(self, images,feature,  target=None, mixup_hidden = False,  mixup_alpha = 0.1, layer_mix=None):\n        b, t, h, w = images.shape\n        images = images.view(b * t // 3, 3, h, w)\n        feature_maps = self.conv_proj(self.backbone.forward_features(images))\n        _, c, h, w = feature_maps.size()\n        feature_maps = feature_maps.contiguous().view(b * 2,c,t // 2 // 3, h, w)\n        feature_maps = self.triple_layer(feature_maps)\n        middle_maps = feature_maps[:, :, 2, :, :]\n        #b, h, w= middle_maps.size()\n        #middle_maps = middle_maps.view(b, 1, h, w)\n        nn_feature = self.neck(self.pool(middle_maps).reshape(b, -1))\n        feature = self.mlp(feature)\n        cat_features = torch.cat([nn_feature, feature], dim=1)\n        if target is not None:\n            cat_features, y_a, y_b, lam = mixup_data(cat_features, target, mixup_alpha)\n            y = self.fc(cat_features)\n            return y, y_a, y_b, lam\n        else:\n            y = self.fc(cat_features)\n            return y\n</code></pre>",
      "rawMarkdown": "I explain about Residual3DBlock.\nIn this competition, 3d(t, h, w) is very important, so I decided to use any time analysis method.\n15 frames reshape this ({-28, -24, -20}{-16, -12, -8}{-4, 0, 4}{8, 12, 16}{20, 24, 28}) and extract feature using backbone, then hidden output is applied it. here is a sample for using residual3d block(final architecture is very similar)\n\n```\nclass Residual3DBlock(nn.Module):\n    def __init__(self):\n        super(Residual3DBlock, self).__init__()\n\n        self.block = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n            nn.ReLU(512)\n        )\n\n        self.block2 = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n        )\n\n    def forward(self, images):\n        short_cut = images\n        h = self.block(images)\n        h = self.block2(h)\n\n        return F.relu(h + short_cut)\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.backbone = timm.create_model(\"tf_efficientnet_b0_ns\", pretrained=True, num_classes=1, in_chans=3)\n        self.mlp = nn.Sequential(\n            nn.Linear(68, 256),\n            nn.BatchNorm1d(256),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n        n_hidden = 1024\n\n        self.conv_proj = nn.Sequential(\n            nn.Conv2d(1280, 512, 1, stride=1),\n            nn.BatchNorm2d(512),\n            nn.ReLU(),\n        )\n\n        self.neck = nn.Sequential(\n            nn.Linear(1024, 1024),\n            nn.BatchNorm1d(1024),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n\n        self.triple_layer = nn.Sequential(\n            Residual3DBlock(),\n        )\n\n        self.pool = GeM() \n\n        self.fc = nn.Linear(256+1024, 1)\n\n    def forward(self, images,feature,  target=None, mixup_hidden = False,  mixup_alpha = 0.1, layer_mix=None):\n        b, t, h, w = images.shape\n        images = images.view(b * t // 3, 3, h, w)\n        feature_maps = self.conv_proj(self.backbone.forward_features(images))\n        _, c, h, w = feature_maps.size()\n        feature_maps = feature_maps.contiguous().view(b * 2,c,t // 2 // 3, h, w)\n        feature_maps = self.triple_layer(feature_maps)\n        middle_maps = feature_maps[:, :, 2, :, :]\n        #b, h, w= middle_maps.size()\n        #middle_maps = middle_maps.view(b, 1, h, w)\n        nn_feature = self.neck(self.pool(middle_maps).reshape(b, -1))\n        feature = self.mlp(feature)\n        cat_features = torch.cat([nn_feature, feature], dim=1)\n        if target is not None:\n            cat_features, y_a, y_b, lam = mixup_data(cat_features, target, mixup_alpha)\n            y = self.fc(cat_features)\n            return y, y_a, y_b, lam\n        else:\n            y = self.fc(cat_features)\n            return y\n```",
      "votes": null
    },
    {
      "id": "2172067",
      "postDate": "03/07/2023 09:34:06",
      "content": "<p>I tried to run the code, but it seems to give me an error. What am I doing wrong?</p>\n<pre><code>model = Model()\nim = torch.randn((8, 15, 512, 512))\nfeature = torch.randn((8, 68))\nmodel(im, feature)\n</code></pre>",
      "rawMarkdown": "I tried to run the code, but it seems to give me an error. What am I doing wrong?\n\n```\nmodel = Model()\nim = torch.randn((8, 15, 512, 512))\nfeature = torch.randn((8, 68))\nmodel(im, feature)\n```",
      "votes": null
    },
    {
      "id": "2172076",
      "postDate": "03/07/2023 09:41:38",
      "content": "<p>Oh, do you use the following because of the end and side?</p>\n<pre><code>im = torch.randn((8, 30, 512, 512))\n</code></pre>",
      "rawMarkdown": "Oh, do you use the following because of the end and side?\n```\nim = torch.randn((8, 30, 512, 512))\n```",
      "votes": null
    },
    {
      "id": "2172287",
      "postDate": "03/07/2023 12:29:08",
      "content": "<p>May I ask a question? what is moving average? I have seen that it is mentioned several times</p>",
      "rawMarkdown": "May I ask a question? what is moving average? I have seen that it is mentioned several times",
      "votes": null
    },
    {
      "id": "2172501",
      "postDate": "03/07/2023 15:04:23",
      "content": "<p>yes, it's correct.<br>\nb, t(side + endzone), height, width.</p>\n<blockquote>\n  <p>im = torch.randn((8, 30, 512, 512))</p>\n</blockquote>",
      "rawMarkdown": "yes, it's correct.\nb, t(side + endzone), height, width.\n\n>im = torch.randn((8, 30, 512, 512))",
      "votes": null
    },
    {
      "id": "2172956",
      "postDate": "03/08/2023 00:58:06",
      "content": "<p>Moving average is the average of previous and next data.<br>\nIn short, <code>df['pred'].rolling().mean()</code>.<br>\nBelow is the pseudo code.<br>\n<code>test_df['pred_ma'] = test_df.groupby(['game_play', 'nfl_player_id_1', 'nfl_player_id_2'])['pred'].rolling(3, center=True, min_periods=1).mean().to_frame('pred_ma').reset_index()['pred_ma']</code></p>",
      "rawMarkdown": "Moving average is the average of previous and next data.\nIn short, `df['pred'].rolling().mean()`.\nBelow is the pseudo code.\n`test_df['pred_ma'] = test_df.groupby(['game_play', 'nfl_player_id_1', 'nfl_player_id_2'])['pred'].rolling(3, center=True, min_periods=1).mean().to_frame('pred_ma').reset_index()['pred_ma']`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2170010,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "03/05/2023 15:56:47",
      "content": "<p>I explain about Residual3DBlock.<br>\nIn this competition, 3d(t, h, w) is very important, so I decided to use any time analysis method.<br>\n15 frames reshape this ({-28, -24, -20}{-16, -12, -8}{-4, 0, 4}{8, 12, 16}{20, 24, 28}) and extract feature using backbone, then hidden output is applied it. here is a sample for using residual3d block(final architecture is very similar)</p>\n<pre><code>class Residual3DBlock(nn.Module):\n    def __init__(self):\n        super(Residual3DBlock, self).__init__()\n\n        self.block = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n            nn.ReLU(512)\n        )\n\n        self.block2 = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n        )\n\n    def forward(self, images):\n        short_cut = images\n        h = self.block(images)\n        h = self.block2(h)\n\n        return F.relu(h + short_cut)\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.backbone = timm.create_model(\"tf_efficientnet_b0_ns\", pretrained=True, num_classes=1, in_chans=3)\n        self.mlp = nn.Sequential(\n            nn.Linear(68, 256),\n            nn.BatchNorm1d(256),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n        n_hidden = 1024\n\n        self.conv_proj = nn.Sequential(\n            nn.Conv2d(1280, 512, 1, stride=1),\n            nn.BatchNorm2d(512),\n            nn.ReLU(),\n        )\n\n        self.neck = nn.Sequential(\n            nn.Linear(1024, 1024),\n            nn.BatchNorm1d(1024),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n\n        self.triple_layer = nn.Sequential(\n            Residual3DBlock(),\n        )\n\n        self.pool = GeM() \n\n        self.fc = nn.Linear(256+1024, 1)\n\n    def forward(self, images,feature,  target=None, mixup_hidden = False,  mixup_alpha = 0.1, layer_mix=None):\n        b, t, h, w = images.shape\n        images = images.view(b * t // 3, 3, h, w)\n        feature_maps = self.conv_proj(self.backbone.forward_features(images))\n        _, c, h, w = feature_maps.size()\n        feature_maps = feature_maps.contiguous().view(b * 2,c,t // 2 // 3, h, w)\n        feature_maps = self.triple_layer(feature_maps)\n        middle_maps = feature_maps[:, :, 2, :, :]\n        #b, h, w= middle_maps.size()\n        #middle_maps = middle_maps.view(b, 1, h, w)\n        nn_feature = self.neck(self.pool(middle_maps).reshape(b, -1))\n        feature = self.mlp(feature)\n        cat_features = torch.cat([nn_feature, feature], dim=1)\n        if target is not None:\n            cat_features, y_a, y_b, lam = mixup_data(cat_features, target, mixup_alpha)\n            y = self.fc(cat_features)\n            return y, y_a, y_b, lam\n        else:\n            y = self.fc(cat_features)\n            return y\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2172067,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "03/07/2023 09:34:06",
          "content": "<p>I tried to run the code, but it seems to give me an error. What am I doing wrong?</p>\n<pre><code>model = Model()\nim = torch.randn((8, 15, 512, 512))\nfeature = torch.randn((8, 68))\nmodel(im, feature)\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2172076,
              "author_name": "yujiariyasu",
              "author_url": "",
              "post_date": "03/07/2023 09:41:38",
              "content": "<p>Oh, do you use the following because of the end and side?</p>\n<pre><code>im = torch.randn((8, 30, 512, 512))\n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 2172501,
                  "author_name": "tereka",
                  "author_url": "",
                  "post_date": "03/07/2023 15:04:23",
                  "content": "<p>yes, it's correct.<br>\nb, t(side + endzone), height, width.</p>\n<blockquote>\n  <p>im = torch.randn((8, 30, 512, 512))</p>\n</blockquote>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2172287,
      "author_name": "chg0901",
      "author_url": "",
      "post_date": "03/07/2023 12:29:08",
      "content": "<p>May I ask a question? what is moving average? I have seen that it is mentioned several times</p>",
      "votes": null,
      "replies": [
        {
          "id": 2172956,
          "author_name": "anonamename",
          "author_url": "",
          "post_date": "03/08/2023 00:58:06",
          "content": "<p>Moving average is the average of previous and next data.<br>\nIn short, <code>df['pred'].rolling().mean()</code>.<br>\nBelow is the pseudo code.<br>\n<code>test_df['pred_ma'] = test_df.groupby(['game_play', 'nfl_player_id_1', 'nfl_player_id_2'])['pred'].rolling(3, center=True, min_periods=1).mean().to_frame('pred_ma').reset_index()['pred_ma']</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2169464": "I would like to thank the organizers for such an interesting competition!   \nWe share the Team JK's solution.  \nTeam Member: @vostankovich, @tereka, @anonamename, @yururoi, @tomo20180402\n<br>\n\n# Overview\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2126876%2F4daeeaa70bf25eda3c4bb2519bcac346%2Fjk_solution_image.png?generation=1677991277683737&alt=media)\n\n# 1st stage\n\n---\n\n## yuki part\n- (1) of fig.\n- See [yuki's post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392046).\n\n## Vladislav part\n- (2) of fig.\n- Features are mainly created from sensor data, but helmets bboxes information is also used.\n- Trained XGB and LGBM models for P2P and P2G individually.\n- P2P and P2G have different features. There are 133 features for pair contact and 119 features for ground contact.\n- Here is explanation of some features:\n    - Excluded speed, since it correlates to distance.\n    - Step (or frame_number), it boosts score a lot.\n    - Player position on field (defense, offense or special)\n    - Twist feature (direction-orientation)\n    - Same team feature\n    - Is home team feature\n    - Number of players/opponents in (1,3,5 meters) is quite good feature\n    - Number of players in opposite orientation\n    - Acceleration of player ratio to mean acceleration of all players per step\n    - Diff of features of same player (in time domain)\n    - Time features (just copy of previous and future steps features)\n    - Difference of features between two players\n    - Euclidean distance is the main feature and other features based on it as well\n    - Features from helmets dataframe (bboxes coordinates, bboxes height & width for each view and perimeter)\n    - IoU helmets features\n- XGB/LGBM models were trained with common hyperperameters that can be seen on public notebooks. Only added reg_alpha = 0.1 for both models.\n\n## anonamename part - combined knowledge of team members\n- (3) of fig.\n- 2-stage model of 2.5D/3D CNN and GBDT (5fold CV:0.778/Public:0.775/Private:0.773)\n- 2.5D/3D CNN\n    - based [public notebook](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference).\n    - input\n        - image\n            - 15frames (±7frame, skip_frame=1)\n            - use both view (Endzone and Sideline)\n        - tracking data\n            - 64 features (created by @vostankovich)\n    - model\n        - based [DFL competition 1st solution](https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932).\n        - pipeline : 15frames 2.5D -> Residual3DBlock -> GeM (created by @tereka)\n        - 2.5D backborn : tf_mobilenetv3_small_minimal_100.in1k\n        - multi-label classification (created by @anonamename)\n            - num_classes=2(Player-Player contact(P2P) and Player-Ground contact(P2G)) + nn.BCEWithLogitsLoss\n        - fold : StratifiedGroupKFold(n_splits=5).split(y=\"contact_org\", groups=\"game_id\") (created by @tomo20180402)\n            - Set different labels for contacts between same team, different teams and ground.\n            - train data under sampling : positive:negative = 1:5 (change under sampling data for each epoch)\n    - optimaizer : AdamW(lr=1e-3->1e-5 CosineAnnealingLR, weight_decay=1e-5)\n    - epoch : 15\n    - augmentation\n        - HorizontalFlip, ShiftScaleRotate, MotionBlur, OpticalDistortion, CoarseDropout\n        - Mixup at the last layer (like a [Manifold mixup](https://arxiv.org/abs/1806.05236). created by @tereka)\n    - TTA : HorizontalFlip\n- GBDT\n    - Create xgboost and lightgbm for P2P and P2G individually.\n    - tracking feature (created by @vostankovich)\n    - 2.5D/3D CNN prob feature\n        - groupby([\"game_play\", \"nfl_player_id_1\", \"nfl_player_id_2\"]) : shift(), diff(), mean(), max(), min(), std()\n\n## tomo part\n- (4) of fig.\n- single-stage NN model (3fold CV:0.771/Public:0.759/Private:0.760)\n- multi-class classification : P2P (same team), P2P (different team), P2G\n    - output is 6 labels which are used as features of team's 2nd stage\n- execution time : 2h\n- validation : StratifiedGroupKFold(n_splits=3).split(y=\"contact_org\", groups=\"game_id\")\n    - same as anonamename part \n- dataset\n    - train data under sampling : Reduce negative sample of P2P contact (same team) by one-third.\n- feature\n    - table feature : 54\n        - 3 types distance : euclidean, chebyshev, cityblock\n        - 3 types distance rank : among all, same team, different team\n        - median of helmet width and height\n        - normalized distance by mean of helmet width and height\n            - The mean of helmet width and height are calculated from all players.\n        - total rank from the center coordinates of 2player's helmets\n        - ratio of helmet detection exist : both players, each player\n        - cosine similarity : direction, orientation\n        - predicted euclidean distance\n        - other simple features : step, is_same_team, ground_flag, etc.\n    - image feature\n        - 10 images in 2.5D CNN\n            - 5frames each for Sideline and Endline (n-4, n-2, n, n+2, n+4)\n            - image_size = (256, 256)\n        - cropping method\n            - Change the cropping method depending on whether both players’ helmets exist.\n                - both players exist : Make sure both players are visible.\n                - one player exist : Make sure the player is in the center.\n            - Crop the image with the mean of helmet width and height as a variable.\n            - Give priority to the downward direction.\n        - mean of image exist : 4\n            - each for Sideline and Endline\n- TTA\n    - flip sensor and image in one of three models inferences\n        - sensor : exchange player1,2\n        - image : HorizontalFlip\n\n# 2nd stage\n\n---\n\n- model : lgbm × 4\n- feature : shift features of each models’ predictions and sensor data (-13~+13)\n- postprocessing : 4 predictions by lgbm -> simple average -> moving average -> final prediction",
    "2170010": "I explain about Residual3DBlock.\nIn this competition, 3d(t, h, w) is very important, so I decided to use any time analysis method.\n15 frames reshape this ({-28, -24, -20}{-16, -12, -8}{-4, 0, 4}{8, 12, 16}{20, 24, 28}) and extract feature using backbone, then hidden output is applied it. here is a sample for using residual3d block(final architecture is very similar)\n\n```\nclass Residual3DBlock(nn.Module):\n    def __init__(self):\n        super(Residual3DBlock, self).__init__()\n\n        self.block = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n            nn.ReLU(512)\n        )\n\n        self.block2 = nn.Sequential(\n            nn.Conv3d(512, 512, 3, stride=1, padding=1),\n            nn.BatchNorm3d(512),\n        )\n\n    def forward(self, images):\n        short_cut = images\n        h = self.block(images)\n        h = self.block2(h)\n\n        return F.relu(h + short_cut)\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.backbone = timm.create_model(\"tf_efficientnet_b0_ns\", pretrained=True, num_classes=1, in_chans=3)\n        self.mlp = nn.Sequential(\n            nn.Linear(68, 256),\n            nn.BatchNorm1d(256),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n        n_hidden = 1024\n\n        self.conv_proj = nn.Sequential(\n            nn.Conv2d(1280, 512, 1, stride=1),\n            nn.BatchNorm2d(512),\n            nn.ReLU(),\n        )\n\n        self.neck = nn.Sequential(\n            nn.Linear(1024, 1024),\n            nn.BatchNorm1d(1024),\n            nn.LeakyReLU(),\n            nn.Dropout(0.2),\n        )\n\n        self.triple_layer = nn.Sequential(\n            Residual3DBlock(),\n        )\n\n        self.pool = GeM() \n\n        self.fc = nn.Linear(256+1024, 1)\n\n    def forward(self, images,feature,  target=None, mixup_hidden = False,  mixup_alpha = 0.1, layer_mix=None):\n        b, t, h, w = images.shape\n        images = images.view(b * t // 3, 3, h, w)\n        feature_maps = self.conv_proj(self.backbone.forward_features(images))\n        _, c, h, w = feature_maps.size()\n        feature_maps = feature_maps.contiguous().view(b * 2,c,t // 2 // 3, h, w)\n        feature_maps = self.triple_layer(feature_maps)\n        middle_maps = feature_maps[:, :, 2, :, :]\n        #b, h, w= middle_maps.size()\n        #middle_maps = middle_maps.view(b, 1, h, w)\n        nn_feature = self.neck(self.pool(middle_maps).reshape(b, -1))\n        feature = self.mlp(feature)\n        cat_features = torch.cat([nn_feature, feature], dim=1)\n        if target is not None:\n            cat_features, y_a, y_b, lam = mixup_data(cat_features, target, mixup_alpha)\n            y = self.fc(cat_features)\n            return y, y_a, y_b, lam\n        else:\n            y = self.fc(cat_features)\n            return y\n```",
    "2172067": "I tried to run the code, but it seems to give me an error. What am I doing wrong?\n\n```\nmodel = Model()\nim = torch.randn((8, 15, 512, 512))\nfeature = torch.randn((8, 68))\nmodel(im, feature)\n```",
    "2172076": "Oh, do you use the following because of the end and side?\n```\nim = torch.randn((8, 30, 512, 512))\n```",
    "2172287": "May I ask a question? what is moving average? I have seen that it is mentioned several times",
    "2172501": "yes, it's correct.\nb, t(side + endzone), height, width.\n\n>im = torch.randn((8, 30, 512, 512))",
    "2172956": "Moving average is the average of previous and next data.\nIn short, `df['pred'].rolling().mean()`.\nBelow is the pseudo code.\n`test_df['pred_ma'] = test_df.groupby(['game_play', 'nfl_player_id_1', 'nfl_player_id_2'])['pred'].rolling(3, center=True, min_periods=1).mean().to_frame('pred_ma').reset_index()['pred_ma']`"
  },
  "source": "meta"
}