{
  "id": 392182,
  "title": "3rd place solution, single stage approach",
  "url": "/competitions/nfl-player-contact-detection/writeups/dmytro-poplavskiy-3rd-place-solution-single-stage-",
  "author_name": "",
  "post_date": "2023-03-12T01:33:00.347Z",
  "votes": 37,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'd like to thank the I'd like to thank organisers for a very interesting challenge (especially <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for providing very useful answers and helping teams). It was interesting to participate.</p>\n<h2>Overview</h2>\n<p>The approach is single-stage, trained end-to-end with a single model executed per player and step interval (instead of per pairs or players) and predicting for all input steps range the ground contact for the current player and contact with 7 nearest players. The model has a video encoder part to process input video frames and a transformer decoder to combine tracking features and video activations.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fa93573968664b0e0af1534e3819e9d49%2FKaggle%20model_1.png?generation=1677986296204472&amp;alt=media\" alt=\"\"></p>\n<h3>Video encoders</h3>\n<p>The video encoders used a number of input video frames around requested steps and produced activations at corresponding steps at downsampled resolution, usually for 16 steps with corresponding 96 frames using every second frame for input.</p>\n<p>I used a few different models for video encoders:</p>\n<ul>\n<li>2d imagenet pretrained models + 3d Conv layer (credits to the Team Hydrogen solution of one of previous competitions). 3 input frames around the current step are converted to grayscale and used as an input to 2d model, with the results combined using 3d conv. Usually larger models performed better for me, with the best performing model based on the convnext large backbone. Other Convnext based models or DPN92 also worked ok.</li>\n<li>2d imagenet pretrained models + TSM, with the color inputs for every 2nd or 3rd frame and TSM like activation exchange between frames before every convolution. Worked better with smaller models like convnext pico or resnet 34 (would probably work better with larger models if the TSM converted model were pretrained on video tasks).</li>\n<li>3D/Video models like CLIP-X (X-CLIP-B/16 was the second best performing model) or the Video Swin Transformer (performed okeish but not included in the final submission).</li>\n</ul>\n<p>Video frames were cropped to 224x224 resolution with the current player's helmet placed at the center/top part of the frame and scaled so the average size of helmets in surrounding frames would be scaled to 34 pixels.<br>\nI applied augmentations to randomly shift, scale, rotate images, shift HUV, added blur and noise.</p>\n<p>For video model activations (at the 32x downsampled 7x7 resolution) I added the positional encoding and learnable separate sideline / endzone markers.<br>\nOptionally the video activations may be encoded using transformers per frame in a similar way as done in DETR but I found it has little to no impact on the result.</p>\n<h3>Transformer player features / video activations decoder</h3>\n<p>The idea is to use attention mechanisms to combine the players features with other surrounding players information and to query the relevant parts of the images.</p>\n<p>For particular player and step, I selected the current player features for surrounding -7..+8 steps and for every step I selected up to 7 nearest players within 2.4 yards, so in total 16 steps * (7+1) players inputs.</p>\n<p>For every player/step input I used the following features, added together using per feature linear transformation to match the transformer features dim:</p>\n<ul>\n<li>position encoding for the helmet pos on the sideline and endzone video, if within 128 pixels from the crop.</li>\n<li>is it visible on sideline and endzone frames</li>\n<li>pos encoding for the step number</li>\n<li>is player the current selected player</li>\n<li>is player from the same team as the current player or not</li>\n<li>player position (not xy but the role from the tracking metadata)</li>\n<li>speed over +- 2 frames</li>\n<li>signed acceleration over +- 2 frames</li>\n<li>distance to the current player, both values and one hot encoding over +- 2 frames</li>\n<li>relative orientation, of the player relative to player-player0 and of player0 relative to player, encoded as sin and cos over +- 2 frames</li>\n<li>for visible helmets, I also added the activations from the video at the helmet position directly to player features. The idea was - it's most likely relevant and may help to avoid using the attention heads for the same task, but I found no difference in the final result.</li>\n</ul>\n<p>Player/step features are used as inputs/targets for a few iterations of transformer layers:</p>\n<ul>\n<li>For all step/player input, I applied the transformer decoder layer with the query over video activations from the same step. </li>\n<li>For all step/player inputs I applied the transformer encoder with the self attention over all  players/steps:</li>\n</ul>\n<pre><code>        # video shape is HW*2 x T*B x C\n        # player_features shape is P, T, B, C\n        # where P - players, T - time_steps, B - batch, C - features, HW - video activations dims\n        x = player_features\n        for step in range(self.num_decoder_layers):\n            x = x.reshape(P, T*B, C)  # reshape to move time steps to batch to use attention only over the current step\n            x = self.video_decoders[step](x, video)\n\n            x = x.reshape(P*T, B, C) # attention over all players/steps\n            x = self.player_decoders[step](x)\n</code></pre>\n<p>I tested with the number of iterations between 2 and 8 and the results were comparable, so I used 2 iterations for most of models.</p>\n<h2>Data pre-processing</h2>\n<p>Mostly to smooth the predicted helmets trajectory, smoothed the prediction to find and remove outliers and interpolated/extrapolated.<br>\nDuring the early test the impact on the performance was not very large, so not conclusive.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Feb690ac7195ad8d03205681175d3c979%2Fplayers_trajectory_pp.png?generation=1677916456068367&amp;alt=media\" alt=\"\"></p>\n<h2>Training</h2>\n<p>For training I selected all players and steps with helmet detected on at least one video (so model would have the tracking features for a few steps before or after the player was visible for the first/last time). I have not excluded any samples using other rules.</p>\n<p>I used the AdamW optimiser with quite a small batch size of 1 to 4 and CosineAnnealingWarmRestarts scheduler with the epoch size of 1024-2048 samples, trained for 68 epochs. It takes about 6-10 hours to train a single model on 3090 GPU.<br>\nI evaluated model every time the scheduler reaches the min rate at epochs 14, 36 and 68.</p>\n<p>I used the BCE loss with slight label smoothing of 0.001..0.999 (it was a guess, I have not tuned hyperparameters much).</p>\n<p>I added aux outputs to the video models to predict if the current player has contact with other players or ground and heatmap of other player helmets with contacts, but the impact on the score was not very large.</p>\n<h2>Prediction</h2>\n<p>The prediction is very straightforward, for model with the input interval of 11 or 16 steps I run it with the smaller offset of 5 steps to predict over the overlapped intervals for every player.</p>\n<p>predictions = defaultdict(list)  # key is (game, step, player1, player2)</p>\n<p>Every prediction between the current and another player, it's added to the list at the dictionary key (gameplay, step, min(player0, player), max(player0, player))<br>\nand all predictions are averages. Usually predictions for the pair of players at a certain step would include predictions with each player as the current one and a few step intervals when the current step is closer to the beginning, middle and end of the intervals.</p>\n<p>When ensembles multiple models, their predictions are added to the same predictions dictionary, with better models added 2-3 times to increase their weight.<br>\nIn total, I used 7 models for the best submission.</p>\n<h2>Individual models performance</h2>\n<table>\n<thead>\n<tr>\n<th>Video model type, backbone</th>\n<th>Notes</th>\n<th>Private LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Convnext large, 2D + 3D conv</td>\n<td>16 steps/96 frames, skip 1 frame.</td>\n<td>0.7915</td>\n</tr>\n<tr>\n<td>Convnext base, 2D + 3D conv</td>\n<td>16 steps/96 frames, skip 1 frame.</td>\n<td>0.786</td>\n</tr>\n<tr>\n<td>DPN92, 2D + 3D conv</td>\n<td>16 steps/96 frames, skip 1 frame.</td>\n<td>0.784</td>\n</tr>\n<tr>\n<td>X-CLIP-B/16</td>\n<td>11 steps/64 frames, skip 1 frame.</td>\n<td>0.791</td>\n</tr>\n<tr>\n<td>X-CLIP-B/32</td>\n<td>11 steps/64 frames, skip 1 frame.</td>\n<td>0.784</td>\n</tr>\n<tr>\n<td>Convnext pico, TSM</td>\n<td>63 steps/384 frames, skip 2 frames.</td>\n<td>0.788</td>\n</tr>\n<tr>\n<td>Convnext pico, 2D + 3D conv</td>\n<td>64 steps/384 frames, skip 2 frames.</td>\n<td>Local CV slightly worse than TSM</td>\n</tr>\n<tr>\n<td>2 best models ensemble</td>\n<td>Convnext large and X-CLIP-B/16,</td>\n<td>0.7925</td>\n</tr>\n<tr>\n<td>6 models ensemble</td>\n<td>Without DPN92, re-trained on full data with original helmets</td>\n<td>0.7932</td>\n</tr>\n<tr>\n<td>6 models ensemble</td>\n<td>Without DPN92, re-trained on full data with fixed helmets</td>\n<td>0.7934</td>\n</tr>\n<tr>\n<td>7 models ensemble</td>\n<td>Convnext large  added with weight 3 and X-CLIP-B/16 with weight 2. Models trained on different folds.</td>\n<td>0.7956</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<ul>\n<li>Training Video Encoder model using aux losses before training transformer decoders. Video Encoder overfits.</li>\n<li>Adding much more tracking features to player transformer inputs. When added the history over larger number of steps for each player input, the transformer encoder overfits.</li>\n<li>Larger models with TSM</li>\n<li>Fix players/helmets assignment in the provided baseline helmets prediction. On some folds the impact was negligible, on some the score has improved by ~ 0.005 even without re-training models. On the private LB the score was similar with and without helmets fixed. One submitted model was using the original data pre-processing, another using more complex pipeline with helmets re-assigned.</li>\n</ul>\n<h2>Local CV challenges</h2>\n<p>To check for possible issues with models generalisation, I decided to split to folds using the sorted by game play list of games, with the first 25% of games assigned to fold 0 validation fold and so on.</p>\n<p>I found to have not only the difference between folds in score, but models/ideas performing well on one fold may work much worse on another.<br>\nFor example, I found on the fold 2, the models with the very large receptive field over time/steps (384 steps, over 6 seconds,  convnext pico based models in the submission) performed by about 0.008 better than the best larger models, while the score fo such models was by the similar 0.007 worse on the fold 3.</p>\n<p>All this made the local validation much more challenging and harder to trust. Taking into account the private dataset is even smaller than every fold, I expected to see a significant shakeup.</p>\n<h2>Player helmets re-assignment</h2>\n<p>Since it was not part of the best submission, added as a separate post: <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392392\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392392</a></p>\n<p>Instead of the data pre-processing described above, I used the estimated tracking -&gt; video transformation to interpolate/extrapolate missing helmets information. The best result was when I discarded the first or the last predicted helmet position and extrapolated by 8 steps maintaining the difference with the position predicted from tracking and tracking-&gt;view transformation.</p>\n<p>The submission source is available at <a href=\"https://www.kaggle.com/dmytropoplavskiy/nfl-sub-place3\" target=\"_blank\">https://www.kaggle.com/dmytropoplavskiy/nfl-sub-place3</a></p>",
  "messages": [
    {
      "id": "2168208",
      "postDate": "03/04/2023 02:15:10",
      "content": "<p>I'd like to thank the I'd like to thank organisers for a very interesting challenge (especially <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for providing very useful answers and helping teams). It was interesting to participate.</p>\n<h2>Overview</h2>\n<p>The approach is single-stage, trained end-to-end with a single model executed per player and step interval (instead of per pairs or players) and predicting for all input steps range the ground contact for the current player and contact with 7 nearest players. The model has a video encoder part to process input video frames and a transformer decoder to combine tracking features and video activations.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fa93573968664b0e0af1534e3819e9d49%2FKaggle%20model_1.png?generation=1677986296204472&amp;alt=media\" alt=\"\"></p>\n<h3>Video encoders</h3>\n<p>The video encoders used a number of input video frames around requested steps and produced activations at corresponding steps at downsampled resolution, usually for 16 steps with corresponding 96 frames using every second frame for input.</p>\n<p>I used a few different models for video encoders:</p>\n<ul>\n<li>2d imagenet pretrained models + 3d Conv layer (credits to the Team Hydrogen solution of one of previous competitions). 3 input frames around the current step are converted to grayscale and used as an input to 2d model, with the results combined using 3d conv. Usually larger models performed better for me, with the best performing model based on the convnext large backbone. Other Convnext based models or DPN92 also worked ok.</li>\n<li>2d imagenet pretrained models + TSM, with the color inputs for every 2nd or 3rd frame and TSM like activation exchange between frames before every convolution. Worked better with smaller models like convnext pico or resnet 34 (would probably work better with larger models if the TSM converted model were pretrained on video tasks).</li>\n<li>3D/Video models like CLIP-X (X-CLIP-B/16 was the second best performing model) or the Video Swin Transformer (performed okeish but not included in the final submission).</li>\n</ul>\n<p>Video frames were cropped to 224x224 resolution with the current player's helmet placed at the center/top part of the frame and scaled so the average size of helmets in surrounding frames would be scaled to 34 pixels.<br>\nI applied augmentations to randomly shift, scale, rotate images, shift HUV, added blur and noise.</p>\n<p>For video model activations (at the 32x downsampled 7x7 resolution) I added the positional encoding and learnable separate sideline / endzone markers.<br>\nOptionally the video activations may be encoded using transformers per frame in a similar way as done in DETR but I found it has little to no impact on the result.</p>\n<h3>Transformer player features / video activations decoder</h3>\n<p>The idea is to use attention mechanisms to combine the players features with other surrounding players information and to query the relevant parts of the images.</p>\n<p>For particular player and step, I selected the current player features for surrounding -7..+8 steps and for every step I selected up to 7 nearest players within 2.4 yards, so in total 16 steps * (7+1) players inputs.</p>\n<p>For every player/step input I used the following features, added together using per feature linear transformation to match the transformer features dim:</p>\n<ul>\n<li>position encoding for the helmet pos on the sideline and endzone video, if within 128 pixels from the crop.</li>\n<li>is it visible on sideline and endzone frames</li>\n<li>pos encoding for the step number</li>\n<li>is player the current selected player</li>\n<li>is player from the same team as the current player or not</li>\n<li>player position (not xy but the role from the tracking metadata)</li>\n<li>speed over +- 2 frames</li>\n<li>signed acceleration over +- 2 frames</li>\n<li>distance to the current player, both values and one hot encoding over +- 2 frames</li>\n<li>relative orientation, of the player relative to player-player0 and of player0 relative to player, encoded as sin and cos over +- 2 frames</li>\n<li>for visible helmets, I also added the activations from the video at the helmet position directly to player features. The idea was - it's most likely relevant and may help to avoid using the attention heads for the same task, but I found no difference in the final result.</li>\n</ul>\n<p>Player/step features are used as inputs/targets for a few iterations of transformer layers:</p>\n<ul>\n<li>For all step/player input, I applied the transformer decoder layer with the query over video activations from the same step. </li>\n<li>For all step/player inputs I applied the transformer encoder with the self attention over all  players/steps:</li>\n</ul>\n<pre><code>        # video shape is HW*2 x T*B x C\n        # player_features shape is P, T, B, C\n        # where P - players, T - time_steps, B - batch, C - features, HW - video activations dims\n        x = player_features\n        for step in range(self.num_decoder_layers):\n            x = x.reshape(P, T*B, C)  # reshape to move time steps to batch to use attention only over the current step\n            x = self.video_decoders[step](x, video)\n\n            x = x.reshape(P*T, B, C) # attention over all players/steps\n            x = self.player_decoders[step](x)\n</code></pre>\n<p>I tested with the number of iterations between 2 and 8 and the results were comparable, so I used 2 iterations for most of models.</p>\n<h2>Data pre-processing</h2>\n<p>Mostly to smooth the predicted helmets trajectory, smoothed the prediction to find and remove outliers and interpolated/extrapolated.<br>\nDuring the early test the impact on the performance was not very large, so not conclusive.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Feb690ac7195ad8d03205681175d3c979%2Fplayers_trajectory_pp.png?generation=1677916456068367&amp;alt=media\" alt=\"\"></p>\n<h2>Training</h2>\n<p>For training I selected all players and steps with helmet detected on at least one video (so model would have the tracking features for a few steps before or after the player was visible for the first/last time). I have not excluded any samples using other rules.</p>\n<p>I used the AdamW optimiser with quite a small batch size of 1 to 4 and CosineAnnealingWarmRestarts scheduler with the epoch size of 1024-2048 samples, trained for 68 epochs. It takes about 6-10 hours to train a single model on 3090 GPU.<br>\nI evaluated model every time the scheduler reaches the min rate at epochs 14, 36 and 68.</p>\n<p>I used the BCE loss with slight label smoothing of 0.001..0.999 (it was a guess, I have not tuned hyperparameters much).</p>\n<p>I added aux outputs to the video models to predict if the current player has contact with other players or ground and heatmap of other player helmets with contacts, but the impact on the score was not very large.</p>\n<h2>Prediction</h2>\n<p>The prediction is very straightforward, for model with the input interval of 11 or 16 steps I run it with the smaller offset of 5 steps to predict over the overlapped intervals for every player.</p>\n<p>predictions = defaultdict(list)  # key is (game, step, player1, player2)</p>\n<p>Every prediction between the current and another player, it's added to the list at the dictionary key (gameplay, step, min(player0, player), max(player0, player))<br>\nand all predictions are averages. Usually predictions for the pair of players at a certain step would include predictions with each player as the current one and a few step intervals when the current step is closer to the beginning, middle and end of the intervals.</p>\n<p>When ensembles multiple models, their predictions are added to the same predictions dictionary, with better models added 2-3 times to increase their weight.<br>\nIn total, I used 7 models for the best submission.</p>\n<h2>Individual models performance</h2>\n<table>\n<thead>\n<tr>\n<th>Video model type, backbone</th>\n<th>Notes</th>\n<th>Private LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Convnext large, 2D + 3D conv</td>\n<td>16 steps/96 frames, skip 1 frame.</td>\n<td>0.7915</td>\n</tr>\n<tr>\n<td>Convnext base, 2D + 3D conv</td>\n<td>16 steps/96 frames, skip 1 frame.</td>\n<td>0.786</td>\n</tr>\n<tr>\n<td>DPN92, 2D + 3D conv</td>\n<td>16 steps/96 frames, skip 1 frame.</td>\n<td>0.784</td>\n</tr>\n<tr>\n<td>X-CLIP-B/16</td>\n<td>11 steps/64 frames, skip 1 frame.</td>\n<td>0.791</td>\n</tr>\n<tr>\n<td>X-CLIP-B/32</td>\n<td>11 steps/64 frames, skip 1 frame.</td>\n<td>0.784</td>\n</tr>\n<tr>\n<td>Convnext pico, TSM</td>\n<td>63 steps/384 frames, skip 2 frames.</td>\n<td>0.788</td>\n</tr>\n<tr>\n<td>Convnext pico, 2D + 3D conv</td>\n<td>64 steps/384 frames, skip 2 frames.</td>\n<td>Local CV slightly worse than TSM</td>\n</tr>\n<tr>\n<td>2 best models ensemble</td>\n<td>Convnext large and X-CLIP-B/16,</td>\n<td>0.7925</td>\n</tr>\n<tr>\n<td>6 models ensemble</td>\n<td>Without DPN92, re-trained on full data with original helmets</td>\n<td>0.7932</td>\n</tr>\n<tr>\n<td>6 models ensemble</td>\n<td>Without DPN92, re-trained on full data with fixed helmets</td>\n<td>0.7934</td>\n</tr>\n<tr>\n<td>7 models ensemble</td>\n<td>Convnext large  added with weight 3 and X-CLIP-B/16 with weight 2. Models trained on different folds.</td>\n<td>0.7956</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<ul>\n<li>Training Video Encoder model using aux losses before training transformer decoders. Video Encoder overfits.</li>\n<li>Adding much more tracking features to player transformer inputs. When added the history over larger number of steps for each player input, the transformer encoder overfits.</li>\n<li>Larger models with TSM</li>\n<li>Fix players/helmets assignment in the provided baseline helmets prediction. On some folds the impact was negligible, on some the score has improved by ~ 0.005 even without re-training models. On the private LB the score was similar with and without helmets fixed. One submitted model was using the original data pre-processing, another using more complex pipeline with helmets re-assigned.</li>\n</ul>\n<h2>Local CV challenges</h2>\n<p>To check for possible issues with models generalisation, I decided to split to folds using the sorted by game play list of games, with the first 25% of games assigned to fold 0 validation fold and so on.</p>\n<p>I found to have not only the difference between folds in score, but models/ideas performing well on one fold may work much worse on another.<br>\nFor example, I found on the fold 2, the models with the very large receptive field over time/steps (384 steps, over 6 seconds,  convnext pico based models in the submission) performed by about 0.008 better than the best larger models, while the score fo such models was by the similar 0.007 worse on the fold 3.</p>\n<p>All this made the local validation much more challenging and harder to trust. Taking into account the private dataset is even smaller than every fold, I expected to see a significant shakeup.</p>\n<h2>Player helmets re-assignment</h2>\n<p>Since it was not part of the best submission, added as a separate post: <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392392\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392392</a></p>\n<p>Instead of the data pre-processing described above, I used the estimated tracking -&gt; video transformation to interpolate/extrapolate missing helmets information. The best result was when I discarded the first or the last predicted helmet position and extrapolated by 8 steps maintaining the difference with the position predicted from tracking and tracking-&gt;view transformation.</p>\n<p>The submission source is available at <a href=\"https://www.kaggle.com/dmytropoplavskiy/nfl-sub-place3\" target=\"_blank\">https://www.kaggle.com/dmytropoplavskiy/nfl-sub-place3</a></p>",
      "rawMarkdown": "I'd like to thank the I'd like to thank organisers for a very interesting challenge (especially @robikscube for providing very useful answers and helping teams). It was interesting to participate.\n\n## Overview\n\nThe approach is single-stage, trained end-to-end with a single model executed per player and step interval (instead of per pairs or players) and predicting for all input steps range the ground contact for the current player and contact with 7 nearest players. The model has a video encoder part to process input video frames and a transformer decoder to combine tracking features and video activations.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fa93573968664b0e0af1534e3819e9d49%2FKaggle%20model_1.png?generation=1677986296204472&alt=media)\n\n\n\n### Video encoders\n\nThe video encoders used a number of input video frames around requested steps and produced activations at corresponding steps at downsampled resolution, usually for 16 steps with corresponding 96 frames using every second frame for input.\n\nI used a few different models for video encoders:\n\n- 2d imagenet pretrained models + 3d Conv layer (credits to the Team Hydrogen solution of one of previous competitions). 3 input frames around the current step are converted to grayscale and used as an input to 2d model, with the results combined using 3d conv. Usually larger models performed better for me, with the best performing model based on the convnext large backbone. Other Convnext based models or DPN92 also worked ok.\n- 2d imagenet pretrained models + TSM, with the color inputs for every 2nd or 3rd frame and TSM like activation exchange between frames before every convolution. Worked better with smaller models like convnext pico or resnet 34 (would probably work better with larger models if the TSM converted model were pretrained on video tasks).\n- 3D/Video models like CLIP-X (X-CLIP-B/16 was the second best performing model) or the Video Swin Transformer (performed okeish but not included in the final submission).\n\nVideo frames were cropped to 224x224 resolution with the current player's helmet placed at the center/top part of the frame and scaled so the average size of helmets in surrounding frames would be scaled to 34 pixels.\nI applied augmentations to randomly shift, scale, rotate images, shift HUV, added blur and noise.\n\nFor video model activations (at the 32x downsampled 7x7 resolution) I added the positional encoding and learnable separate sideline / endzone markers.\nOptionally the video activations may be encoded using transformers per frame in a similar way as done in DETR but I found it has little to no impact on the result.\n\n\n### Transformer player features / video activations decoder\n\nThe idea is to use attention mechanisms to combine the players features with other surrounding players information and to query the relevant parts of the images.\n\nFor particular player and step, I selected the current player features for surrounding -7..+8 steps and for every step I selected up to 7 nearest players within 2.4 yards, so in total 16 steps * (7+1) players inputs.\n\nFor every player/step input I used the following features, added together using per feature linear transformation to match the transformer features dim:\n- position encoding for the helmet pos on the sideline and endzone video, if within 128 pixels from the crop.\n- is it visible on sideline and endzone frames\n- pos encoding for the step number\n- is player the current selected player\n- is player from the same team as the current player or not\n- player position (not xy but the role from the tracking metadata)\n- speed over +- 2 frames\n- signed acceleration over +- 2 frames\n- distance to the current player, both values and one hot encoding over +- 2 frames\n- relative orientation, of the player relative to player-player0 and of player0 relative to player, encoded as sin and cos over +- 2 frames\n- for visible helmets, I also added the activations from the video at the helmet position directly to player features. The idea was - it's most likely relevant and may help to avoid using the attention heads for the same task, but I found no difference in the final result.\n\nPlayer/step features are used as inputs/targets for a few iterations of transformer layers:\n- For all step/player input, I applied the transformer decoder layer with the query over video activations from the same step. \n- For all step/player inputs I applied the transformer encoder with the self attention over all  players/steps:\n\n```\n        # video shape is HW*2 x T*B x C\n        # player_features shape is P, T, B, C\n        # where P - players, T - time_steps, B - batch, C - features, HW - video activations dims\n        x = player_features\n        for step in range(self.num_decoder_layers):\n            x = x.reshape(P, T*B, C)  # reshape to move time steps to batch to use attention only over the current step\n            x = self.video_decoders[step](x, video)\n\n            x = x.reshape(P*T, B, C) # attention over all players/steps\n            x = self.player_decoders[step](x)\n```\n\nI tested with the number of iterations between 2 and 8 and the results were comparable, so I used 2 iterations for most of models.\n\n## Data pre-processing\n\nMostly to smooth the predicted helmets trajectory, smoothed the prediction to find and remove outliers and interpolated/extrapolated.\nDuring the early test the impact on the performance was not very large, so not conclusive.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Feb690ac7195ad8d03205681175d3c979%2Fplayers_trajectory_pp.png?generation=1677916456068367&alt=media)\n\n## Training\n\nFor training I selected all players and steps with helmet detected on at least one video (so model would have the tracking features for a few steps before or after the player was visible for the first/last time). I have not excluded any samples using other rules.\n\nI used the AdamW optimiser with quite a small batch size of 1 to 4 and CosineAnnealingWarmRestarts scheduler with the epoch size of 1024-2048 samples, trained for 68 epochs. It takes about 6-10 hours to train a single model on 3090 GPU.\nI evaluated model every time the scheduler reaches the min rate at epochs 14, 36 and 68.\n\nI used the BCE loss with slight label smoothing of 0.001..0.999 (it was a guess, I have not tuned hyperparameters much).\n\nI added aux outputs to the video models to predict if the current player has contact with other players or ground and heatmap of other player helmets with contacts, but the impact on the score was not very large.\n\n## Prediction\n\nThe prediction is very straightforward, for model with the input interval of 11 or 16 steps I run it with the smaller offset of 5 steps to predict over the overlapped intervals for every player.\n\npredictions = defaultdict(list)  # key is (game, step, player1, player2)\n\nEvery prediction between the current and another player, it's added to the list at the dictionary key (gameplay, step, min(player0, player), max(player0, player))\nand all predictions are averages. Usually predictions for the pair of players at a certain step would include predictions with each player as the current one and a few step intervals when the current step is closer to the beginning, middle and end of the intervals.\n\nWhen ensembles multiple models, their predictions are added to the same predictions dictionary, with better models added 2-3 times to increase their weight.\nIn total, I used 7 models for the best submission.\n\n## Individual models performance\n\n| Video model type, backbone | Notes                                      | Private LB score |\n|----------------------------|--------------------------------------------|------------------|\n| Convnext large, 2D + 3D conv| 16 steps/96 frames, skip 1 frame.         | 0.7915           |\n| Convnext base, 2D + 3D conv| 16 steps/96 frames, skip 1 frame.          | 0.786            |\n| DPN92, 2D + 3D conv        | 16 steps/96 frames, skip 1 frame.          | 0.784            |\n| X-CLIP-B/16                | 11 steps/64 frames, skip 1 frame.          | 0.791            |\n| X-CLIP-B/32                | 11 steps/64 frames, skip 1 frame.          | 0.784            |\n| Convnext pico, TSM         | 63 steps/384 frames, skip 2 frames.        | 0.788            |\n| Convnext pico, 2D + 3D conv| 64 steps/384 frames, skip 2 frames.        | Local CV slightly worse than TSM |\n| 2 best models ensemble |  Convnext large and X-CLIP-B/16,   | 0.7925 |\n| 6 models ensemble |  Without DPN92, re-trained on full data with original helmets  | 0.7932 |\n| 6 models ensemble |  Without DPN92, re-trained on full data with fixed helmets      | 0.7934 |\n| 7 models ensemble |  Convnext large  added with weight 3 and X-CLIP-B/16 with weight 2. Models trained on different folds.  | 0.7956 |\n\n## What did not work\n\n- Training Video Encoder model using aux losses before training transformer decoders. Video Encoder overfits.\n- Adding much more tracking features to player transformer inputs. When added the history over larger number of steps for each player input, the transformer encoder overfits.\n- Larger models with TSM\n- Fix players/helmets assignment in the provided baseline helmets prediction. On some folds the impact was negligible, on some the score has improved by ~ 0.005 even without re-training models. On the private LB the score was similar with and without helmets fixed. One submitted model was using the original data pre-processing, another using more complex pipeline with helmets re-assigned.\n\n## Local CV challenges\n\nTo check for possible issues with models generalisation, I decided to split to folds using the sorted by game play list of games, with the first 25% of games assigned to fold 0 validation fold and so on.\n\nI found to have not only the difference between folds in score, but models/ideas performing well on one fold may work much worse on another.\nFor example, I found on the fold 2, the models with the very large receptive field over time/steps (384 steps, over 6 seconds,  convnext pico based models in the submission) performed by about 0.008 better than the best larger models, while the score fo such models was by the similar 0.007 worse on the fold 3.\n\nAll this made the local validation much more challenging and harder to trust. Taking into account the private dataset is even smaller than every fold, I expected to see a significant shakeup.\n\n## Player helmets re-assignment\n\nSince it was not part of the best submission, added as a separate post: https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392392\n\nInstead of the data pre-processing described above, I used the estimated tracking -> video transformation to interpolate/extrapolate missing helmets information. The best result was when I discarded the first or the last predicted helmet position and extrapolated by 8 steps maintaining the difference with the position predicted from tracking and tracking->view transformation.\n\n\nThe submission source is available at https://www.kaggle.com/dmytropoplavskiy/nfl-sub-place3",
      "votes": null
    },
    {
      "id": "2168466",
      "postDate": "03/04/2023 09:14:28",
      "content": "<p>Congrats, what a cool and beautiful architecture!!<br>\nif I understand correctly, you stack 96 frames of 224x224 image and feed them to Convnext large. <br>\nHow much GPU memory is required??</p>",
      "rawMarkdown": "Congrats, what a cool and beautiful architecture!!\nif I understand correctly, you stack 96 frames of 224x224 image and feed them to Convnext large. \nHow much GPU memory is required??",
      "votes": null
    },
    {
      "id": "2168488",
      "postDate": "03/04/2023 09:30:58",
      "content": "<p>Thanks!</p>\n<p>I should probably clarify, 96 frames is the slice length/duration, I used only every second frame (or even 3rd frame for the last 2 models).</p>\n<p>With 2D+3D approach in addition I converted 3 frames to monochrome and used it as an input to 2d CNN, so it was actually 96/(3*2) = 16 combined frames/runs of 224x224 convnext large. With the batch size of 2, it used 19GB of VRAM for ConvNext Large and ~13GB for ConvNext Base during training.</p>",
      "rawMarkdown": "Thanks!\n\nI should probably clarify, 96 frames is the slice length/duration, I used only every second frame (or even 3rd frame for the last 2 models).\n\nWith 2D+3D approach in addition I converted 3 frames to monochrome and used it as an input to 2d CNN, so it was actually 96/(3*2) = 16 combined frames/runs of 224x224 convnext large. With the batch size of 2, it used 19GB of VRAM for ConvNext Large and ~13GB for ConvNext Base during training.",
      "votes": null
    },
    {
      "id": "2168902",
      "postDate": "03/04/2023 16:09:31",
      "content": "<p>Really great solution <a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> - and congrats on the result! I'll probably need to read this again a few times before I can understand exactly how the archetecture works 😄.</p>\n<p>It's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.</p>\n<p>I'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?</p>\n<p>The helmet imputation is a clever idea. Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?</p>\n<p>Thanks for sharing your writeup.</p>",
      "rawMarkdown": "Really great solution @dmytropoplavskiy - and congrats on the result! I'll probably need to read this again a few times before I can understand exactly how the archetecture works 😄.\n\nIt's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.\n\nI'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?\n\nThe helmet imputation is a clever idea. Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?\n\nThanks for sharing your writeup.",
      "votes": null
    },
    {
      "id": "2168904",
      "postDate": "03/04/2023 16:12:03",
      "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> Thanks for the clarification, now I understand it's actually a reasonable number! It seems stacking gray 3 neighbor images on channel dim is very useful technique to reduce memory usage in video deep learning.</p>",
      "rawMarkdown": "dmytropoplavskiy Thanks for the clarification, now I understand it's actually a reasonable number! It seems stacking gray 3 neighbor images on channel dim is very useful technique to reduce memory usage in video deep learning.",
      "votes": null
    },
    {
      "id": "2169260",
      "postDate": "03/05/2023 00:12:16",
      "content": "<p>Hi, thank you again for organizing a very interesting competition, it was a pleasure to participate.</p>\n<blockquote>\n  <p>It's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.</p>\n</blockquote>\n<p>I checked the distribution of Nth nearest player with contact (calculated for both players in the contact pair):</p>\n<table>\n<thead>\n<tr>\n<th>Nearest player num</th>\n<th>number of contacts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>69207</td>\n</tr>\n<tr>\n<td>2</td>\n<td>17584</td>\n</tr>\n<tr>\n<td>3</td>\n<td>5244</td>\n</tr>\n<tr>\n<td>4</td>\n<td>2000</td>\n</tr>\n<tr>\n<td>5</td>\n<td>785</td>\n</tr>\n<tr>\n<td>6</td>\n<td>271</td>\n</tr>\n<tr>\n<td>7</td>\n<td>135</td>\n</tr>\n<tr>\n<td>8</td>\n<td>59</td>\n</tr>\n<tr>\n<td>9</td>\n<td>22</td>\n</tr>\n<tr>\n<td>10</td>\n<td>23</td>\n</tr>\n<tr>\n<td>11</td>\n<td>15</td>\n</tr>\n<tr>\n<td>12</td>\n<td>14</td>\n</tr>\n<tr>\n<td>13</td>\n<td>8</td>\n</tr>\n<tr>\n<td>14</td>\n<td>12</td>\n</tr>\n<tr>\n<td>15</td>\n<td>37</td>\n</tr>\n</tbody>\n</table>\n<p>When I checked contacts only within the distance of 2.4 (edited/fixed):</p>\n<table>\n<thead>\n<tr>\n<th>Nearest player num</th>\n<th>number of contacts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>69177</td>\n</tr>\n<tr>\n<td>2</td>\n<td>17569</td>\n</tr>\n<tr>\n<td>3</td>\n<td>5222</td>\n</tr>\n<tr>\n<td>4</td>\n<td>1992</td>\n</tr>\n<tr>\n<td>5</td>\n<td>768</td>\n</tr>\n<tr>\n<td>6</td>\n<td>263</td>\n</tr>\n<tr>\n<td>7</td>\n<td>119</td>\n</tr>\n<tr>\n<td>8</td>\n<td>48</td>\n</tr>\n<tr>\n<td>9</td>\n<td>15</td>\n</tr>\n<tr>\n<td>10</td>\n<td>12</td>\n</tr>\n<tr>\n<td>11</td>\n<td>10</td>\n</tr>\n<tr>\n<td>12</td>\n<td>4</td>\n</tr>\n<tr>\n<td>13</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>Since the contact prediction is averaged when evaluated from both players of view, some (likely most or even all)<br>\ncontacts would still be checked. For example if player2 is 8th nearest player for player1 in contact, player1 may be the 5th nearest player for player 2, so the contact would still be evaluated from player2 point of view.</p>\n<p>I have not tested the model score with the different number of nearest players, but since the model can accept the variable size input, I tried one of the models on one of folds:</p>\n<table>\n<thead>\n<tr>\n<th>Number of nearest players</th>\n<th>threshold for the best score</th>\n<th>score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>15</td>\n<td>0.5800</td>\n<td>0.7626</td>\n</tr>\n<tr>\n<td>13</td>\n<td>0.5400</td>\n<td>0.7708</td>\n</tr>\n<tr>\n<td>11</td>\n<td>0.5200</td>\n<td>0.7784</td>\n</tr>\n<tr>\n<td>9</td>\n<td>0.4400</td>\n<td>0.7881</td>\n</tr>\n<tr>\n<td>8</td>\n<td>0.4000</td>\n<td>0.7910</td>\n</tr>\n<tr>\n<td>7</td>\n<td>0.3400</td>\n<td>0.7926</td>\n</tr>\n<tr>\n<td>6</td>\n<td>0.3000</td>\n<td>0.7938</td>\n</tr>\n<tr>\n<td>5</td>\n<td>0.2200</td>\n<td>0.7921</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.1800</td>\n<td>0.7900</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.1200</td>\n<td>0.7873</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.1000</td>\n<td>0.7824</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.0600</td>\n<td>0.7499</td>\n</tr>\n</tbody>\n</table>\n<p>So looks like the selected 7 players choice was reasonable, 6 players worked slightly better with the score of 0.7938. Maybe when trained on the 15 players input the model would learn better how such messy cases are annotated.</p>\n<blockquote>\n  <p>I'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?</p>\n</blockquote>\n<p>I added the position encoding (grid of sin/cos values at different frequencies, like used with NLP) to 7x7 grid of video encoders activations (starting from -128, -128 pix to encode positions around the visible area) and I also added similar position encoding for the helmet position on the sideline and endzone views (with different linear projections to allow models to query both views).</p>\n<p>This way the similar position encoding is used for both key and query parts of the transformer decoder attention and allows to associate and query parts of images relevant to the player visible position. I allowed to encode positions within 128pix of the visible area to be able to query players with contact but the helmet not visible in the current step.</p>\n<p>I accidentally introduced a bug in the dataset class and provided the main player video position to all nearest players and this caused the significant degradation of model performance. I also tried to supply one of the activations from 7x7 grid with the player helmet directly to players features, but I have not noticed the significant difference, looks like the model is able to use the supplied position encodings.</p>\n<blockquote>\n  <p>Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?</p>\n</blockquote>\n<p>I only used the position of the helmet on views (if visible). If the helmet is outside of [-128pix..crop+128pix] box, the pos encoding for corresponding view values are set to zero. </p>\n<p>I run prediction for the current player only for steps when the player is visible on at least one view, but since the prediction is done for a number of steps (for example 16 steps, or +-0.8s from the current timestamp, with the current timestamp sampled at 0.5s steps), it's possible the player will be not visible on the previous or next timestamp. But the model would still predict contacts for steps around the visible interval, using the previously visible frames and tracking information (the self attention part of the encoder which uses attention over all players and all time steps).</p>\n<p>If the nearest player is not visible on either view, I think it's still included but model would have access to only tracking information or images of this player from surrounding steps if he was visible (it may be hard to associate players only using the tracking info).</p>",
      "rawMarkdown": "Hi, thank you again for organizing a very interesting competition, it was a pleasure to participate.\n\n\n>It's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.\n\nI checked the distribution of Nth nearest player with contact (calculated for both players in the contact pair):\n\nNearest player num  |  number of contacts\n---------------------|----------------------\n1   |  69207\n2   |  17584\n3   |   5244\n4   |   2000\n5   |    785\n6   |    271\n7   |    135\n8   |     59\n9    |    22\n10  |     23\n11   |    15\n12   |    14\n13   |     8\n14   |    12\n15  |     37\n\nWhen I checked contacts only within the distance of 2.4 (edited/fixed):\n\nNearest player num  |  number of contacts\n---------------------|----------------------\n1   |  69177\n2   |  17569\n3   |   5222\n4   |   1992\n5   |    768\n6   |    263\n7   |    119\n8   |     48\n9   |     15\n10 |      12\n11  |     10\n12  |      4\n13  |      1\n\nSince the contact prediction is averaged when evaluated from both players of view, some (likely most or even all)\ncontacts would still be checked. For example if player2 is 8th nearest player for player1 in contact, player1 may be the 5th nearest player for player 2, so the contact would still be evaluated from player2 point of view.\n\nI have not tested the model score with the different number of nearest players, but since the model can accept the variable size input, I tried one of the models on one of folds:\n\nNumber of nearest players | threshold for the best score | score\n----------------------------|------------------------------|------\n15 | 0.5800 | 0.7626\n13 | 0.5400 | 0.7708\n11 | 0.5200 | 0.7784\n9  | 0.4400 | 0.7881\n8   | 0.4000 | 0.7910\n7   | 0.3400 | 0.7926\n6   | 0.3000 | 0.7938\n5   | 0.2200 | 0.7921\n4   | 0.1800  | 0.7900\n3   | 0.1200  | 0.7873\n2   | 0.1000  | 0.7824\n1   | 0.0600 | 0.7499\n\nSo looks like the selected 7 players choice was reasonable, 6 players worked slightly better with the score of 0.7938. Maybe when trained on the 15 players input the model would learn better how such messy cases are annotated.\n\n> I'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?\n\nI added the position encoding (grid of sin/cos values at different frequencies, like used with NLP) to 7x7 grid of video encoders activations (starting from -128, -128 pix to encode positions around the visible area) and I also added similar position encoding for the helmet position on the sideline and endzone views (with different linear projections to allow models to query both views).\n\nThis way the similar position encoding is used for both key and query parts of the transformer decoder attention and allows to associate and query parts of images relevant to the player visible position. I allowed to encode positions within 128pix of the visible area to be able to query players with contact but the helmet not visible in the current step.\n\nI accidentally introduced a bug in the dataset class and provided the main player video position to all nearest players and this caused the significant degradation of model performance. I also tried to supply one of the activations from 7x7 grid with the player helmet directly to players features, but I have not noticed the significant difference, looks like the model is able to use the supplied position encodings.\n\n\n> Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?\n\nI only used the position of the helmet on views (if visible). If the helmet is outside of [-128pix..crop+128pix] box, the pos encoding for corresponding view values are set to zero. \n\nI run prediction for the current player only for steps when the player is visible on at least one view, but since the prediction is done for a number of steps (for example 16 steps, or +-0.8s from the current timestamp, with the current timestamp sampled at 0.5s steps), it's possible the player will be not visible on the previous or next timestamp. But the model would still predict contacts for steps around the visible interval, using the previously visible frames and tracking information (the self attention part of the encoder which uses attention over all players and all time steps).\n\nIf the nearest player is not visible on either view, I think it's still included but model would have access to only tracking information or images of this player from surrounding steps if he was visible (it may be hard to associate players only using the tracking info).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168466,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "03/04/2023 09:14:28",
      "content": "<p>Congrats, what a cool and beautiful architecture!!<br>\nif I understand correctly, you stack 96 frames of 224x224 image and feed them to Convnext large. <br>\nHow much GPU memory is required??</p>",
      "votes": null,
      "replies": [
        {
          "id": 2168488,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "03/04/2023 09:30:58",
          "content": "<p>Thanks!</p>\n<p>I should probably clarify, 96 frames is the slice length/duration, I used only every second frame (or even 3rd frame for the last 2 models).</p>\n<p>With 2D+3D approach in addition I converted 3 frames to monochrome and used it as an input to 2d CNN, so it was actually 96/(3*2) = 16 combined frames/runs of 224x224 convnext large. With the batch size of 2, it used 19GB of VRAM for ConvNext Large and ~13GB for ConvNext Base during training.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2168904,
              "author_name": "bamps53",
              "author_url": "",
              "post_date": "03/04/2023 16:12:03",
              "content": "<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> Thanks for the clarification, now I understand it's actually a reasonable number! It seems stacking gray 3 neighbor images on channel dim is very useful technique to reduce memory usage in video deep learning.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2168902,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:09:31",
      "content": "<p>Really great solution <a href=\"https://www.kaggle.com/dmytropoplavskiy\" target=\"_blank\">@dmytropoplavskiy</a> - and congrats on the result! I'll probably need to read this again a few times before I can understand exactly how the archetecture works 😄.</p>\n<p>It's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.</p>\n<p>I'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?</p>\n<p>The helmet imputation is a clever idea. Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?</p>\n<p>Thanks for sharing your writeup.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169260,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "03/05/2023 00:12:16",
          "content": "<p>Hi, thank you again for organizing a very interesting competition, it was a pleasure to participate.</p>\n<blockquote>\n  <p>It's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.</p>\n</blockquote>\n<p>I checked the distribution of Nth nearest player with contact (calculated for both players in the contact pair):</p>\n<table>\n<thead>\n<tr>\n<th>Nearest player num</th>\n<th>number of contacts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>69207</td>\n</tr>\n<tr>\n<td>2</td>\n<td>17584</td>\n</tr>\n<tr>\n<td>3</td>\n<td>5244</td>\n</tr>\n<tr>\n<td>4</td>\n<td>2000</td>\n</tr>\n<tr>\n<td>5</td>\n<td>785</td>\n</tr>\n<tr>\n<td>6</td>\n<td>271</td>\n</tr>\n<tr>\n<td>7</td>\n<td>135</td>\n</tr>\n<tr>\n<td>8</td>\n<td>59</td>\n</tr>\n<tr>\n<td>9</td>\n<td>22</td>\n</tr>\n<tr>\n<td>10</td>\n<td>23</td>\n</tr>\n<tr>\n<td>11</td>\n<td>15</td>\n</tr>\n<tr>\n<td>12</td>\n<td>14</td>\n</tr>\n<tr>\n<td>13</td>\n<td>8</td>\n</tr>\n<tr>\n<td>14</td>\n<td>12</td>\n</tr>\n<tr>\n<td>15</td>\n<td>37</td>\n</tr>\n</tbody>\n</table>\n<p>When I checked contacts only within the distance of 2.4 (edited/fixed):</p>\n<table>\n<thead>\n<tr>\n<th>Nearest player num</th>\n<th>number of contacts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>69177</td>\n</tr>\n<tr>\n<td>2</td>\n<td>17569</td>\n</tr>\n<tr>\n<td>3</td>\n<td>5222</td>\n</tr>\n<tr>\n<td>4</td>\n<td>1992</td>\n</tr>\n<tr>\n<td>5</td>\n<td>768</td>\n</tr>\n<tr>\n<td>6</td>\n<td>263</td>\n</tr>\n<tr>\n<td>7</td>\n<td>119</td>\n</tr>\n<tr>\n<td>8</td>\n<td>48</td>\n</tr>\n<tr>\n<td>9</td>\n<td>15</td>\n</tr>\n<tr>\n<td>10</td>\n<td>12</td>\n</tr>\n<tr>\n<td>11</td>\n<td>10</td>\n</tr>\n<tr>\n<td>12</td>\n<td>4</td>\n</tr>\n<tr>\n<td>13</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>Since the contact prediction is averaged when evaluated from both players of view, some (likely most or even all)<br>\ncontacts would still be checked. For example if player2 is 8th nearest player for player1 in contact, player1 may be the 5th nearest player for player 2, so the contact would still be evaluated from player2 point of view.</p>\n<p>I have not tested the model score with the different number of nearest players, but since the model can accept the variable size input, I tried one of the models on one of folds:</p>\n<table>\n<thead>\n<tr>\n<th>Number of nearest players</th>\n<th>threshold for the best score</th>\n<th>score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>15</td>\n<td>0.5800</td>\n<td>0.7626</td>\n</tr>\n<tr>\n<td>13</td>\n<td>0.5400</td>\n<td>0.7708</td>\n</tr>\n<tr>\n<td>11</td>\n<td>0.5200</td>\n<td>0.7784</td>\n</tr>\n<tr>\n<td>9</td>\n<td>0.4400</td>\n<td>0.7881</td>\n</tr>\n<tr>\n<td>8</td>\n<td>0.4000</td>\n<td>0.7910</td>\n</tr>\n<tr>\n<td>7</td>\n<td>0.3400</td>\n<td>0.7926</td>\n</tr>\n<tr>\n<td>6</td>\n<td>0.3000</td>\n<td>0.7938</td>\n</tr>\n<tr>\n<td>5</td>\n<td>0.2200</td>\n<td>0.7921</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.1800</td>\n<td>0.7900</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.1200</td>\n<td>0.7873</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.1000</td>\n<td>0.7824</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.0600</td>\n<td>0.7499</td>\n</tr>\n</tbody>\n</table>\n<p>So looks like the selected 7 players choice was reasonable, 6 players worked slightly better with the score of 0.7938. Maybe when trained on the 15 players input the model would learn better how such messy cases are annotated.</p>\n<blockquote>\n  <p>I'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?</p>\n</blockquote>\n<p>I added the position encoding (grid of sin/cos values at different frequencies, like used with NLP) to 7x7 grid of video encoders activations (starting from -128, -128 pix to encode positions around the visible area) and I also added similar position encoding for the helmet position on the sideline and endzone views (with different linear projections to allow models to query both views).</p>\n<p>This way the similar position encoding is used for both key and query parts of the transformer decoder attention and allows to associate and query parts of images relevant to the player visible position. I allowed to encode positions within 128pix of the visible area to be able to query players with contact but the helmet not visible in the current step.</p>\n<p>I accidentally introduced a bug in the dataset class and provided the main player video position to all nearest players and this caused the significant degradation of model performance. I also tried to supply one of the activations from 7x7 grid with the player helmet directly to players features, but I have not noticed the significant difference, looks like the model is able to use the supplied position encodings.</p>\n<blockquote>\n  <p>Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?</p>\n</blockquote>\n<p>I only used the position of the helmet on views (if visible). If the helmet is outside of [-128pix..crop+128pix] box, the pos encoding for corresponding view values are set to zero. </p>\n<p>I run prediction for the current player only for steps when the player is visible on at least one view, but since the prediction is done for a number of steps (for example 16 steps, or +-0.8s from the current timestamp, with the current timestamp sampled at 0.5s steps), it's possible the player will be not visible on the previous or next timestamp. But the model would still predict contacts for steps around the visible interval, using the previously visible frames and tracking information (the self attention part of the encoder which uses attention over all players and all time steps).</p>\n<p>If the nearest player is not visible on either view, I think it's still included but model would have access to only tracking information or images of this player from surrounding steps if he was visible (it may be hard to associate players only using the tracking info).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2168208": "I'd like to thank the I'd like to thank organisers for a very interesting challenge (especially @robikscube for providing very useful answers and helping teams). It was interesting to participate.\n\n## Overview\n\nThe approach is single-stage, trained end-to-end with a single model executed per player and step interval (instead of per pairs or players) and predicting for all input steps range the ground contact for the current player and contact with 7 nearest players. The model has a video encoder part to process input video frames and a transformer decoder to combine tracking features and video activations.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fa93573968664b0e0af1534e3819e9d49%2FKaggle%20model_1.png?generation=1677986296204472&alt=media)\n\n\n\n### Video encoders\n\nThe video encoders used a number of input video frames around requested steps and produced activations at corresponding steps at downsampled resolution, usually for 16 steps with corresponding 96 frames using every second frame for input.\n\nI used a few different models for video encoders:\n\n- 2d imagenet pretrained models + 3d Conv layer (credits to the Team Hydrogen solution of one of previous competitions). 3 input frames around the current step are converted to grayscale and used as an input to 2d model, with the results combined using 3d conv. Usually larger models performed better for me, with the best performing model based on the convnext large backbone. Other Convnext based models or DPN92 also worked ok.\n- 2d imagenet pretrained models + TSM, with the color inputs for every 2nd or 3rd frame and TSM like activation exchange between frames before every convolution. Worked better with smaller models like convnext pico or resnet 34 (would probably work better with larger models if the TSM converted model were pretrained on video tasks).\n- 3D/Video models like CLIP-X (X-CLIP-B/16 was the second best performing model) or the Video Swin Transformer (performed okeish but not included in the final submission).\n\nVideo frames were cropped to 224x224 resolution with the current player's helmet placed at the center/top part of the frame and scaled so the average size of helmets in surrounding frames would be scaled to 34 pixels.\nI applied augmentations to randomly shift, scale, rotate images, shift HUV, added blur and noise.\n\nFor video model activations (at the 32x downsampled 7x7 resolution) I added the positional encoding and learnable separate sideline / endzone markers.\nOptionally the video activations may be encoded using transformers per frame in a similar way as done in DETR but I found it has little to no impact on the result.\n\n\n### Transformer player features / video activations decoder\n\nThe idea is to use attention mechanisms to combine the players features with other surrounding players information and to query the relevant parts of the images.\n\nFor particular player and step, I selected the current player features for surrounding -7..+8 steps and for every step I selected up to 7 nearest players within 2.4 yards, so in total 16 steps * (7+1) players inputs.\n\nFor every player/step input I used the following features, added together using per feature linear transformation to match the transformer features dim:\n- position encoding for the helmet pos on the sideline and endzone video, if within 128 pixels from the crop.\n- is it visible on sideline and endzone frames\n- pos encoding for the step number\n- is player the current selected player\n- is player from the same team as the current player or not\n- player position (not xy but the role from the tracking metadata)\n- speed over +- 2 frames\n- signed acceleration over +- 2 frames\n- distance to the current player, both values and one hot encoding over +- 2 frames\n- relative orientation, of the player relative to player-player0 and of player0 relative to player, encoded as sin and cos over +- 2 frames\n- for visible helmets, I also added the activations from the video at the helmet position directly to player features. The idea was - it's most likely relevant and may help to avoid using the attention heads for the same task, but I found no difference in the final result.\n\nPlayer/step features are used as inputs/targets for a few iterations of transformer layers:\n- For all step/player input, I applied the transformer decoder layer with the query over video activations from the same step. \n- For all step/player inputs I applied the transformer encoder with the self attention over all  players/steps:\n\n```\n        # video shape is HW*2 x T*B x C\n        # player_features shape is P, T, B, C\n        # where P - players, T - time_steps, B - batch, C - features, HW - video activations dims\n        x = player_features\n        for step in range(self.num_decoder_layers):\n            x = x.reshape(P, T*B, C)  # reshape to move time steps to batch to use attention only over the current step\n            x = self.video_decoders[step](x, video)\n\n            x = x.reshape(P*T, B, C) # attention over all players/steps\n            x = self.player_decoders[step](x)\n```\n\nI tested with the number of iterations between 2 and 8 and the results were comparable, so I used 2 iterations for most of models.\n\n## Data pre-processing\n\nMostly to smooth the predicted helmets trajectory, smoothed the prediction to find and remove outliers and interpolated/extrapolated.\nDuring the early test the impact on the performance was not very large, so not conclusive.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Feb690ac7195ad8d03205681175d3c979%2Fplayers_trajectory_pp.png?generation=1677916456068367&alt=media)\n\n## Training\n\nFor training I selected all players and steps with helmet detected on at least one video (so model would have the tracking features for a few steps before or after the player was visible for the first/last time). I have not excluded any samples using other rules.\n\nI used the AdamW optimiser with quite a small batch size of 1 to 4 and CosineAnnealingWarmRestarts scheduler with the epoch size of 1024-2048 samples, trained for 68 epochs. It takes about 6-10 hours to train a single model on 3090 GPU.\nI evaluated model every time the scheduler reaches the min rate at epochs 14, 36 and 68.\n\nI used the BCE loss with slight label smoothing of 0.001..0.999 (it was a guess, I have not tuned hyperparameters much).\n\nI added aux outputs to the video models to predict if the current player has contact with other players or ground and heatmap of other player helmets with contacts, but the impact on the score was not very large.\n\n## Prediction\n\nThe prediction is very straightforward, for model with the input interval of 11 or 16 steps I run it with the smaller offset of 5 steps to predict over the overlapped intervals for every player.\n\npredictions = defaultdict(list)  # key is (game, step, player1, player2)\n\nEvery prediction between the current and another player, it's added to the list at the dictionary key (gameplay, step, min(player0, player), max(player0, player))\nand all predictions are averages. Usually predictions for the pair of players at a certain step would include predictions with each player as the current one and a few step intervals when the current step is closer to the beginning, middle and end of the intervals.\n\nWhen ensembles multiple models, their predictions are added to the same predictions dictionary, with better models added 2-3 times to increase their weight.\nIn total, I used 7 models for the best submission.\n\n## Individual models performance\n\n| Video model type, backbone | Notes                                      | Private LB score |\n|----------------------------|--------------------------------------------|------------------|\n| Convnext large, 2D + 3D conv| 16 steps/96 frames, skip 1 frame.         | 0.7915           |\n| Convnext base, 2D + 3D conv| 16 steps/96 frames, skip 1 frame.          | 0.786            |\n| DPN92, 2D + 3D conv        | 16 steps/96 frames, skip 1 frame.          | 0.784            |\n| X-CLIP-B/16                | 11 steps/64 frames, skip 1 frame.          | 0.791            |\n| X-CLIP-B/32                | 11 steps/64 frames, skip 1 frame.          | 0.784            |\n| Convnext pico, TSM         | 63 steps/384 frames, skip 2 frames.        | 0.788            |\n| Convnext pico, 2D + 3D conv| 64 steps/384 frames, skip 2 frames.        | Local CV slightly worse than TSM |\n| 2 best models ensemble |  Convnext large and X-CLIP-B/16,   | 0.7925 |\n| 6 models ensemble |  Without DPN92, re-trained on full data with original helmets  | 0.7932 |\n| 6 models ensemble |  Without DPN92, re-trained on full data with fixed helmets      | 0.7934 |\n| 7 models ensemble |  Convnext large  added with weight 3 and X-CLIP-B/16 with weight 2. Models trained on different folds.  | 0.7956 |\n\n## What did not work\n\n- Training Video Encoder model using aux losses before training transformer decoders. Video Encoder overfits.\n- Adding much more tracking features to player transformer inputs. When added the history over larger number of steps for each player input, the transformer encoder overfits.\n- Larger models with TSM\n- Fix players/helmets assignment in the provided baseline helmets prediction. On some folds the impact was negligible, on some the score has improved by ~ 0.005 even without re-training models. On the private LB the score was similar with and without helmets fixed. One submitted model was using the original data pre-processing, another using more complex pipeline with helmets re-assigned.\n\n## Local CV challenges\n\nTo check for possible issues with models generalisation, I decided to split to folds using the sorted by game play list of games, with the first 25% of games assigned to fold 0 validation fold and so on.\n\nI found to have not only the difference between folds in score, but models/ideas performing well on one fold may work much worse on another.\nFor example, I found on the fold 2, the models with the very large receptive field over time/steps (384 steps, over 6 seconds,  convnext pico based models in the submission) performed by about 0.008 better than the best larger models, while the score fo such models was by the similar 0.007 worse on the fold 3.\n\nAll this made the local validation much more challenging and harder to trust. Taking into account the private dataset is even smaller than every fold, I expected to see a significant shakeup.\n\n## Player helmets re-assignment\n\nSince it was not part of the best submission, added as a separate post: https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392392\n\nInstead of the data pre-processing described above, I used the estimated tracking -> video transformation to interpolate/extrapolate missing helmets information. The best result was when I discarded the first or the last predicted helmet position and extrapolated by 8 steps maintaining the difference with the position predicted from tracking and tracking->view transformation.\n\n\nThe submission source is available at https://www.kaggle.com/dmytropoplavskiy/nfl-sub-place3",
    "2168466": "Congrats, what a cool and beautiful architecture!!\nif I understand correctly, you stack 96 frames of 224x224 image and feed them to Convnext large. \nHow much GPU memory is required??",
    "2168488": "Thanks!\n\nI should probably clarify, 96 frames is the slice length/duration, I used only every second frame (or even 3rd frame for the last 2 models).\n\nWith 2D+3D approach in addition I converted 3 frames to monochrome and used it as an input to 2d CNN, so it was actually 96/(3*2) = 16 combined frames/runs of 224x224 convnext large. With the batch size of 2, it used 19GB of VRAM for ConvNext Large and ~13GB for ConvNext Base during training.",
    "2168902": "Really great solution @dmytropoplavskiy - and congrats on the result! I'll probably need to read this again a few times before I can understand exactly how the archetecture works 😄.\n\nIt's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.\n\nI'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?\n\nThe helmet imputation is a clever idea. Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?\n\nThanks for sharing your writeup.",
    "2168904": "dmytropoplavskiy Thanks for the clarification, now I understand it's actually a reasonable number! It seems stacking gray 3 neighbor images on channel dim is very useful technique to reduce memory usage in video deep learning.",
    "2169260": "Hi, thank you again for organizing a very interesting competition, it was a pleasure to participate.\n\n\n>It's really interesting how your model predicted per player instead of per pair. Did you decide that using up to the 7th closest player was sufficient to capture any contact? Thats honestly slightly more than I'd expect.\n\nI checked the distribution of Nth nearest player with contact (calculated for both players in the contact pair):\n\nNearest player num  |  number of contacts\n---------------------|----------------------\n1   |  69207\n2   |  17584\n3   |   5244\n4   |   2000\n5   |    785\n6   |    271\n7   |    135\n8   |     59\n9    |    22\n10  |     23\n11   |    15\n12   |    14\n13   |     8\n14   |    12\n15  |     37\n\nWhen I checked contacts only within the distance of 2.4 (edited/fixed):\n\nNearest player num  |  number of contacts\n---------------------|----------------------\n1   |  69177\n2   |  17569\n3   |   5222\n4   |   1992\n5   |    768\n6   |    263\n7   |    119\n8   |     48\n9   |     15\n10 |      12\n11  |     10\n12  |      4\n13  |      1\n\nSince the contact prediction is averaged when evaluated from both players of view, some (likely most or even all)\ncontacts would still be checked. For example if player2 is 8th nearest player for player1 in contact, player1 may be the 5th nearest player for player 2, so the contact would still be evaluated from player2 point of view.\n\nI have not tested the model score with the different number of nearest players, but since the model can accept the variable size input, I tried one of the models on one of folds:\n\nNumber of nearest players | threshold for the best score | score\n----------------------------|------------------------------|------\n15 | 0.5800 | 0.7626\n13 | 0.5400 | 0.7708\n11 | 0.5200 | 0.7784\n9  | 0.4400 | 0.7881\n8   | 0.4000 | 0.7910\n7   | 0.3400 | 0.7926\n6   | 0.3000 | 0.7938\n5   | 0.2200 | 0.7921\n4   | 0.1800  | 0.7900\n3   | 0.1200  | 0.7873\n2   | 0.1000  | 0.7824\n1   | 0.0600 | 0.7499\n\nSo looks like the selected 7 players choice was reasonable, 6 players worked slightly better with the score of 0.7938. Maybe when trained on the 15 players input the model would learn better how such messy cases are annotated.\n\n> I'm not clear on how the model was able to identify which of surrounding players in the video were associated with the player tracking (NGS) features that you provided the decoder. Did you add any additional masking to the images or did the model learn these relationships on it's own?\n\nI added the position encoding (grid of sin/cos values at different frequencies, like used with NLP) to 7x7 grid of video encoders activations (starting from -128, -128 pix to encode positions around the visible area) and I also added similar position encoding for the helmet position on the sideline and endzone views (with different linear projections to allow models to query both views).\n\nThis way the similar position encoding is used for both key and query parts of the transformer decoder attention and allows to associate and query parts of images relevant to the player visible position. I allowed to encode positions within 128pix of the visible area to be able to query players with contact but the helmet not visible in the current step.\n\nI accidentally introduced a bug in the dataset class and provided the main player video position to all nearest players and this caused the significant degradation of model performance. I also tried to supply one of the activations from 7x7 grid with the player helmet directly to players features, but I have not noticed the significant difference, looks like the model is able to use the supplied position encodings.\n\n\n> Did you use any of the helmet bounding box data in the model itself other than identifying the player's helmet to predict for. Also, how did you handle when helmets were not visible in either camera?\n\nI only used the position of the helmet on views (if visible). If the helmet is outside of [-128pix..crop+128pix] box, the pos encoding for corresponding view values are set to zero. \n\nI run prediction for the current player only for steps when the player is visible on at least one view, but since the prediction is done for a number of steps (for example 16 steps, or +-0.8s from the current timestamp, with the current timestamp sampled at 0.5s steps), it's possible the player will be not visible on the previous or next timestamp. But the model would still predict contacts for steps around the visible interval, using the previously visible frames and tracking information (the self attention part of the encoder which uses attention over all players and all time steps).\n\nIf the nearest player is not visible on either view, I think it's still included but model would have access to only tracking information or images of this player from surrounding steps if he was visible (it may be hard to associate players only using the tracking info)."
  },
  "source": "meta"
}