{
  "id": 392162,
  "title": "18th place solution : 2d-cnn / 1d-cnn / XGB / 1d-cnn",
  "url": "/competitions/nfl-player-contact-detection/writeups/thomas-zh-chen-18th-place-solution-2d-cnn-1d-cnn-x",
  "author_name": "",
  "post_date": "2023-03-03T22:22:28.470Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, I want to thank the hosts of this competition and my team: <a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">@chenlin1999</a> and <a href=\"https://www.kaggle.com/hanzhou0315\" target=\"_blank\">@hanzhou0315</a></p>\n<h1>Summary</h1>\n<p>Our solution is made of 4 stages :</p>\n<ol>\n<li><strong>2d-cnn</strong>  : The model predicts for each player if the player is in contact as well as if the player is on the ground</li>\n<li><strong>1d-cnn</strong> : This stage is intended to smooth the prediction of 1. using the temporality</li>\n<li><strong>XGB</strong>     : It is at this moment that we associate the contacts between players</li>\n<li><strong>1d-cnn</strong> : This stage is intended to smooth the prediction of 3. using the temporality</li>\n</ol>\n<h1>Validation methodology</h1>\n<p>We opt for a Stratified group 5-Fold cross-validation by <code>game_play</code> : this strategy seemed to be the most correlated with LB and the most obvious. We have a final solution that reaches a CV score : 0.77174 for a LB : 0.77004 and a public LB : 0.76219</p>\n<h1>Stage1 : 2d-cnn</h1>\n<p>At this stage it is very easy to overfit on the data so we only trained for 2 epochs. We used the timm models: efficientnetv2_rw_s and convnext_base_in22k for the final submission.</p>\n<p>The input is composed of 2 RGB images for the Endzone and the Sideline, then we concatenate the features to make a prediction. To add supervision to this model we used features created from the tabular data. Tt's a bit similar to <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015\" target=\"_blank\">this</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6012866%2Fccac7fba20baef4c47053060a748ff75%2Fimage_2023-03-03_164244573.png?generation=1677879764555718&amp;alt=media\" alt=\"\"></p>\n<h1>Stage2 : 1d-cnn</h1>\n<p>It is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by <code>[\"game_play\",\"nfl_player_id\",\"step\"]</code></p>\n<h1>Stage3 : XGB</h1>\n<p>It's XGB like <a href=\"https://www.kaggle.com/code/columbia2131/nfl-player-contact-detection-simple-xgb-baseline\" target=\"_blank\">this</a> by <a href=\"https://www.kaggle.com/columbia2131\" target=\"_blank\">@columbia2131</a> and we added the features from the previous stage</p>\n<h1>Stage4 : 1d-cnn</h1>\n<p>It is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by <code>[\"game_play\",\"nfl_player_id_1\",\"nfl_player_id_2\",\"step\"]</code></p>\n<h1>Final results</h1>\n<table>\n<thead>\n<tr>\n<th>Stage1</th>\n<th>Stage2</th>\n<th>Stage3</th>\n<th>Stage4</th>\n<th>CV</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td></td>\n<td>✓</td>\n<td></td>\n<td>0.65</td>\n<td>0.645</td>\n<td>0.645</td>\n</tr>\n<tr>\n<td>✓</td>\n<td></td>\n<td>✓</td>\n<td></td>\n<td>0.718</td>\n<td>0.715</td>\n<td>0.716</td>\n</tr>\n<tr>\n<td>✓</td>\n<td></td>\n<td>✓</td>\n<td>✓</td>\n<td>0.731</td>\n<td>0.729</td>\n<td>0.725</td>\n</tr>\n<tr>\n<td>✓</td>\n<td>✓</td>\n<td>✓</td>\n<td>✓</td>\n<td>0.771</td>\n<td>0.762</td>\n<td>0.770</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2168109",
      "postDate": "03/03/2023 22:14:48",
      "content": "<p>First of all, I want to thank the hosts of this competition and my team: <a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">@chenlin1999</a> and <a href=\"https://www.kaggle.com/hanzhou0315\" target=\"_blank\">@hanzhou0315</a></p>\n<h1>Summary</h1>\n<p>Our solution is made of 4 stages :</p>\n<ol>\n<li><strong>2d-cnn</strong>  : The model predicts for each player if the player is in contact as well as if the player is on the ground</li>\n<li><strong>1d-cnn</strong> : This stage is intended to smooth the prediction of 1. using the temporality</li>\n<li><strong>XGB</strong>     : It is at this moment that we associate the contacts between players</li>\n<li><strong>1d-cnn</strong> : This stage is intended to smooth the prediction of 3. using the temporality</li>\n</ol>\n<h1>Validation methodology</h1>\n<p>We opt for a Stratified group 5-Fold cross-validation by <code>game_play</code> : this strategy seemed to be the most correlated with LB and the most obvious. We have a final solution that reaches a CV score : 0.77174 for a LB : 0.77004 and a public LB : 0.76219</p>\n<h1>Stage1 : 2d-cnn</h1>\n<p>At this stage it is very easy to overfit on the data so we only trained for 2 epochs. We used the timm models: efficientnetv2_rw_s and convnext_base_in22k for the final submission.</p>\n<p>The input is composed of 2 RGB images for the Endzone and the Sideline, then we concatenate the features to make a prediction. To add supervision to this model we used features created from the tabular data. Tt's a bit similar to <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015\" target=\"_blank\">this</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6012866%2Fccac7fba20baef4c47053060a748ff75%2Fimage_2023-03-03_164244573.png?generation=1677879764555718&amp;alt=media\" alt=\"\"></p>\n<h1>Stage2 : 1d-cnn</h1>\n<p>It is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by <code>[\"game_play\",\"nfl_player_id\",\"step\"]</code></p>\n<h1>Stage3 : XGB</h1>\n<p>It's XGB like <a href=\"https://www.kaggle.com/code/columbia2131/nfl-player-contact-detection-simple-xgb-baseline\" target=\"_blank\">this</a> by <a href=\"https://www.kaggle.com/columbia2131\" target=\"_blank\">@columbia2131</a> and we added the features from the previous stage</p>\n<h1>Stage4 : 1d-cnn</h1>\n<p>It is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by <code>[\"game_play\",\"nfl_player_id_1\",\"nfl_player_id_2\",\"step\"]</code></p>\n<h1>Final results</h1>\n<table>\n<thead>\n<tr>\n<th>Stage1</th>\n<th>Stage2</th>\n<th>Stage3</th>\n<th>Stage4</th>\n<th>CV</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td></td>\n<td>✓</td>\n<td></td>\n<td>0.65</td>\n<td>0.645</td>\n<td>0.645</td>\n</tr>\n<tr>\n<td>✓</td>\n<td></td>\n<td>✓</td>\n<td></td>\n<td>0.718</td>\n<td>0.715</td>\n<td>0.716</td>\n</tr>\n<tr>\n<td>✓</td>\n<td></td>\n<td>✓</td>\n<td>✓</td>\n<td>0.731</td>\n<td>0.729</td>\n<td>0.725</td>\n</tr>\n<tr>\n<td>✓</td>\n<td>✓</td>\n<td>✓</td>\n<td>✓</td>\n<td>0.771</td>\n<td>0.762</td>\n<td>0.770</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "First of all, I want to thank the hosts of this competition and my team: @chenlin1999 and @hanzhou0315\n\n# Summary\n\nOur solution is made of 4 stages :\n1. **2d-cnn**  : The model predicts for each player if the player is in contact as well as if the player is on the ground\n2. **1d-cnn** : This stage is intended to smooth the prediction of 1. using the temporality\n3. **XGB**     : It is at this moment that we associate the contacts between players\n4. **1d-cnn** : This stage is intended to smooth the prediction of 3. using the temporality\n\n# Validation methodology\n\nWe opt for a Stratified group 5-Fold cross-validation by `game_play ` : this strategy seemed to be the most correlated with LB and the most obvious. We have a final solution that reaches a CV score : 0.77174 for a LB : 0.77004 and a public LB : 0.76219\n\n# Stage1 : 2d-cnn\nAt this stage it is very easy to overfit on the data so we only trained for 2 epochs. We used the timm models: efficientnetv2_rw_s and convnext_base_in22k for the final submission.\n\nThe input is composed of 2 RGB images for the Endzone and the Sideline, then we concatenate the features to make a prediction. To add supervision to this model we used features created from the tabular data. Tt's a bit similar to [this]( https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6012866%2Fccac7fba20baef4c47053060a748ff75%2Fimage_2023-03-03_164244573.png?generation=1677879764555718&alt=media)\n\n# Stage2 : 1d-cnn\n\nIt is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by `[\"game_play\",\"nfl_player_id\",\"step\"]`\n\n# Stage3 : XGB\n\nIt's XGB like [this](https://www.kaggle.com/code/columbia2131/nfl-player-contact-detection-simple-xgb-baseline) by @columbia2131 and we added the features from the previous stage\n\n# Stage4 : 1d-cnn\n\nIt is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by `[\"game_play\",\"nfl_player_id_1\",\"nfl_player_id_2\",\"step\"]`\n\n# Final results\n| Stage1 | Stage2 | Stage3 | Stage4 | CV | public LB | private LB |\n|--|--|--|--|--|\n|   |   | ✓ |   | 0.65 | 0.645 | 0.645 |\n| ✓ |   | ✓ |   | 0.718 | 0.715 | 0.716 |\n| ✓ |   | ✓ | ✓ | 0.731 | 0.729 | 0.725 |\n| ✓ | ✓ | ✓ | ✓ | 0.771 | 0.762| 0.770 |",
      "votes": null
    },
    {
      "id": "2168147",
      "postDate": "03/03/2023 23:14:57",
      "content": "<p>Many thanks to <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for organising this interesting competition and to my two teammates <a href=\"https://www.kaggle.com/thomasdubail\" target=\"_blank\">@thomasdubail</a> and <a href=\"https://www.kaggle.com/hanzhou0315\" target=\"_blank\">@hanzhou0315</a> </p>\n<p>This is my first time participating in a competition using computer vision and tabular data together, and I have seen a very wide variety of completely different solutions in the top solution, each using images and tabular data in a different order and in different ways, as well as being able to achieve very close results in the end.</p>\n<p>I noticed that the first half of our solution was very similar to the stage1 of the solution shared by <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723\" target=\"_blank\">here</a>, but in the second stage we only used the tracking data and some manual features created to train XGB, and did not use the CNN again to classify the images of the player pairs, perhaps this was the last step we needed to take to get to the gold zone.</p>\n<p>Besides, we saw from the top solution in the previous NFL competition that the 2.5d window size was typically 9 frames, but this window size did not lead to a boost in our stage1. Due to time constraints in the end, we decided to abandon the 2.5d which consumes a lot of computing power and did not explore more window size, and now it seems that others got a boost on the larger windows.</p>",
      "rawMarkdown": "Many thanks to @robikscube for organising this interesting competition and to my two teammates @thomasdubail and @hanzhou0315 \n\nThis is my first time participating in a competition using computer vision and tabular data together, and I have seen a very wide variety of completely different solutions in the top solution, each using images and tabular data in a different order and in different ways, as well as being able to achieve very close results in the end.\n\nI noticed that the first half of our solution was very similar to the stage1 of the solution shared by @haqishen [here](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723), but in the second stage we only used the tracking data and some manual features created to train XGB, and did not use the CNN again to classify the images of the player pairs, perhaps this was the last step we needed to take to get to the gold zone.\n\nBesides, we saw from the top solution in the previous NFL competition that the 2.5d window size was typically 9 frames, but this window size did not lead to a boost in our stage1. Due to time constraints in the end, we decided to abandon the 2.5d which consumes a lot of computing power and did not explore more window size, and now it seems that others got a boost on the larger windows.",
      "votes": null
    },
    {
      "id": "2174318",
      "postDate": "03/09/2023 03:43:03",
      "content": "<p>Thank you for sharing solution.<br>\nIt's interesting that the result shows 1D-CNN is more effective in stage2 than stage4.<br>\nHow many window size do you use for 1D-CNN in stage 2 and 4?</p>",
      "rawMarkdown": "Thank you for sharing solution.\nIt's interesting that the result shows 1D-CNN is more effective in stage2 than stage4.\nHow many window size do you use for 1D-CNN in stage 2 and 4?",
      "votes": null
    },
    {
      "id": "2174887",
      "postDate": "03/09/2023 13:39:22",
      "content": "<p>Thank you for the feedback: indeed for me stage1 was already very effective in highlighting the people in contact, stage2 just had the effect of smoothing its values. However in stage3 the association of the players was much more noizy so the gain after this stage was much greater using time. For training I used 64 steps windows (for batch optimization) and then in inference mode I used the whole game as input</p>",
      "rawMarkdown": "Thank you for the feedback: indeed for me stage1 was already very effective in highlighting the people in contact, stage2 just had the effect of smoothing its values. However in stage3 the association of the players was much more noizy so the gain after this stage was much greater using time. For training I used 64 steps windows (for batch optimization) and then in inference mode I used the whole game as input",
      "votes": null
    },
    {
      "id": "2174966",
      "postDate": "03/09/2023 14:41:52",
      "content": "<p>Thanks. One more question: Why stage 2 only focus on single player?<br>\nDo you train the model to the contact event for a single player to any other players? (like <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723\" target=\"_blank\">Qishen's solution</a>?)<br>\n[\"game_play\",\"nfl_player_id\",\"step\"]</p>\n<p>Do you think modeling with contact event for single player is better than that of two players?</p>",
      "rawMarkdown": "Thanks. One more question: Why stage 2 only focus on single player?\nDo you train the model to the contact event for a single player to any other players? (like [Qishen's solution](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723)?)\n[\"game_play\",\"nfl_player_id\",\"step\"]\n\nDo you think modeling with contact event for single player is better than that of two players?",
      "votes": null
    },
    {
      "id": "2175348",
      "postDate": "03/09/2023 19:44:37",
      "content": "<p><code>Why stage 2 only focus on single player?</code><br>\nThis is because the predictions of stage1 are based on the contact of a single player.<br>\n<code>Do you train the model to the contact event for a single player to any other players? (like Qishen's solution?)</code><br>\nYes.<br>\n<code>Do you think modeling with contact event for a single player is better than that of two players?</code><br>\nI have seen many other solutions using the contact situation of two players as the label of the CNN that won the gold medal. Qishen only used the single players CNN in the first stage, but then used the two players CNN in the second stage. Our solution had no part of the two players' CNN at all, which is perhaps what we needed to get the gold medal. My initial thought was that the single player's image would be cropped out with less noise and the network would learn how to extract features more efficiently.</p>",
      "rawMarkdown": "`Why stage 2 only focus on single player?`\nThis is because the predictions of stage1 are based on the contact of a single player.\n`Do you train the model to the contact event for a single player to any other players? (like Qishen's solution?)`\nYes.\n`Do you think modeling with contact event for a single player is better than that of two players?`\nI have seen many other solutions using the contact situation of two players as the label of the CNN that won the gold medal. Qishen only used the single players CNN in the first stage, but then used the two players CNN in the second stage. Our solution had no part of the two players' CNN at all, which is perhaps what we needed to get the gold medal. My initial thought was that the single player's image would be cropped out with less noise and the network would learn how to extract features more efficiently.",
      "votes": null
    },
    {
      "id": "2175470",
      "postDate": "03/09/2023 21:28:55",
      "content": "<p><a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">@chenlin1999</a> Thanks for the explanation.</p>\n<p>Let me add my thought. Since the contact events in train dataset is few (although it seems plenty of them, most of them are very similar to each other), modeling mainly on player to player contact event easy to overfit. So I believe modeling with single player setting is more robust to overfitting to specific contact events than modeling with player to player contacts. My pipeline is mainly on XGBs, and the overfit was serious. CNNs seems more robust to overfitting as some participants says.</p>",
      "rawMarkdown": "chenlin1999 Thanks for the explanation.\n\nLet me add my thought. Since the contact events in train dataset is few (although it seems plenty of them, most of them are very similar to each other), modeling mainly on player to player contact event easy to overfit. So I believe modeling with single player setting is more robust to overfitting to specific contact events than modeling with player to player contacts. My pipeline is mainly on XGBs, and the overfit was serious. CNNs seems more robust to overfitting as some participants says.",
      "votes": null
    },
    {
      "id": "2175481",
      "postDate": "03/09/2023 21:37:10",
      "content": "<p>Yes I think you are right : I have the same thing in mind</p>",
      "rawMarkdown": "Yes I think you are right : I have the same thing in mind",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168147,
      "author_name": "chenlin1999",
      "author_url": "",
      "post_date": "03/03/2023 23:14:57",
      "content": "<p>Many thanks to <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for organising this interesting competition and to my two teammates <a href=\"https://www.kaggle.com/thomasdubail\" target=\"_blank\">@thomasdubail</a> and <a href=\"https://www.kaggle.com/hanzhou0315\" target=\"_blank\">@hanzhou0315</a> </p>\n<p>This is my first time participating in a competition using computer vision and tabular data together, and I have seen a very wide variety of completely different solutions in the top solution, each using images and tabular data in a different order and in different ways, as well as being able to achieve very close results in the end.</p>\n<p>I noticed that the first half of our solution was very similar to the stage1 of the solution shared by <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723\" target=\"_blank\">here</a>, but in the second stage we only used the tracking data and some manual features created to train XGB, and did not use the CNN again to classify the images of the player pairs, perhaps this was the last step we needed to take to get to the gold zone.</p>\n<p>Besides, we saw from the top solution in the previous NFL competition that the 2.5d window size was typically 9 frames, but this window size did not lead to a boost in our stage1. Due to time constraints in the end, we decided to abandon the 2.5d which consumes a lot of computing power and did not explore more window size, and now it seems that others got a boost on the larger windows.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2174318,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "03/09/2023 03:43:03",
      "content": "<p>Thank you for sharing solution.<br>\nIt's interesting that the result shows 1D-CNN is more effective in stage2 than stage4.<br>\nHow many window size do you use for 1D-CNN in stage 2 and 4?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174887,
          "author_name": "thomasdubail",
          "author_url": "",
          "post_date": "03/09/2023 13:39:22",
          "content": "<p>Thank you for the feedback: indeed for me stage1 was already very effective in highlighting the people in contact, stage2 just had the effect of smoothing its values. However in stage3 the association of the players was much more noizy so the gain after this stage was much greater using time. For training I used 64 steps windows (for batch optimization) and then in inference mode I used the whole game as input</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174966,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "03/09/2023 14:41:52",
              "content": "<p>Thanks. One more question: Why stage 2 only focus on single player?<br>\nDo you train the model to the contact event for a single player to any other players? (like <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723\" target=\"_blank\">Qishen's solution</a>?)<br>\n[\"game_play\",\"nfl_player_id\",\"step\"]</p>\n<p>Do you think modeling with contact event for single player is better than that of two players?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2175348,
                  "author_name": "chenlin1999",
                  "author_url": "",
                  "post_date": "03/09/2023 19:44:37",
                  "content": "<p><code>Why stage 2 only focus on single player?</code><br>\nThis is because the predictions of stage1 are based on the contact of a single player.<br>\n<code>Do you train the model to the contact event for a single player to any other players? (like Qishen's solution?)</code><br>\nYes.<br>\n<code>Do you think modeling with contact event for a single player is better than that of two players?</code><br>\nI have seen many other solutions using the contact situation of two players as the label of the CNN that won the gold medal. Qishen only used the single players CNN in the first stage, but then used the two players CNN in the second stage. Our solution had no part of the two players' CNN at all, which is perhaps what we needed to get the gold medal. My initial thought was that the single player's image would be cropped out with less noise and the network would learn how to extract features more efficiently.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2175470,
                      "author_name": "tatamikenn",
                      "author_url": "",
                      "post_date": "03/09/2023 21:28:55",
                      "content": "<p><a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">@chenlin1999</a> Thanks for the explanation.</p>\n<p>Let me add my thought. Since the contact events in train dataset is few (although it seems plenty of them, most of them are very similar to each other), modeling mainly on player to player contact event easy to overfit. So I believe modeling with single player setting is more robust to overfitting to specific contact events than modeling with player to player contacts. My pipeline is mainly on XGBs, and the overfit was serious. CNNs seems more robust to overfitting as some participants says.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2175481,
                          "author_name": "thomasdubail",
                          "author_url": "",
                          "post_date": "03/09/2023 21:37:10",
                          "content": "<p>Yes I think you are right : I have the same thing in mind</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2168109": "First of all, I want to thank the hosts of this competition and my team: @chenlin1999 and @hanzhou0315\n\n# Summary\n\nOur solution is made of 4 stages :\n1. **2d-cnn**  : The model predicts for each player if the player is in contact as well as if the player is on the ground\n2. **1d-cnn** : This stage is intended to smooth the prediction of 1. using the temporality\n3. **XGB**     : It is at this moment that we associate the contacts between players\n4. **1d-cnn** : This stage is intended to smooth the prediction of 3. using the temporality\n\n# Validation methodology\n\nWe opt for a Stratified group 5-Fold cross-validation by `game_play ` : this strategy seemed to be the most correlated with LB and the most obvious. We have a final solution that reaches a CV score : 0.77174 for a LB : 0.77004 and a public LB : 0.76219\n\n# Stage1 : 2d-cnn\nAt this stage it is very easy to overfit on the data so we only trained for 2 epochs. We used the timm models: efficientnetv2_rw_s and convnext_base_in22k for the final submission.\n\nThe input is composed of 2 RGB images for the Endzone and the Sideline, then we concatenate the features to make a prediction. To add supervision to this model we used features created from the tabular data. Tt's a bit similar to [this]( https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6012866%2Fccac7fba20baef4c47053060a748ff75%2Fimage_2023-03-03_164244573.png?generation=1677879764555718&alt=media)\n\n# Stage2 : 1d-cnn\n\nIt is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by `[\"game_play\",\"nfl_player_id\",\"step\"]`\n\n# Stage3 : XGB\n\nIt's XGB like [this](https://www.kaggle.com/code/columbia2131/nfl-player-contact-detection-simple-xgb-baseline) by @columbia2131 and we added the features from the previous stage\n\n# Stage4 : 1d-cnn\n\nIt is a simple CNN with 5 layers and kernels of 3. To ensure temporal consistency we sorted by `[\"game_play\",\"nfl_player_id_1\",\"nfl_player_id_2\",\"step\"]`\n\n# Final results\n| Stage1 | Stage2 | Stage3 | Stage4 | CV | public LB | private LB |\n|--|--|--|--|--|\n|   |   | ✓ |   | 0.65 | 0.645 | 0.645 |\n| ✓ |   | ✓ |   | 0.718 | 0.715 | 0.716 |\n| ✓ |   | ✓ | ✓ | 0.731 | 0.729 | 0.725 |\n| ✓ | ✓ | ✓ | ✓ | 0.771 | 0.762| 0.770 |",
    "2168147": "Many thanks to @robikscube for organising this interesting competition and to my two teammates @thomasdubail and @hanzhou0315 \n\nThis is my first time participating in a competition using computer vision and tabular data together, and I have seen a very wide variety of completely different solutions in the top solution, each using images and tabular data in a different order and in different ways, as well as being able to achieve very close results in the end.\n\nI noticed that the first half of our solution was very similar to the stage1 of the solution shared by @haqishen [here](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723), but in the second stage we only used the tracking data and some manual features created to train XGB, and did not use the CNN again to classify the images of the player pairs, perhaps this was the last step we needed to take to get to the gold zone.\n\nBesides, we saw from the top solution in the previous NFL competition that the 2.5d window size was typically 9 frames, but this window size did not lead to a boost in our stage1. Due to time constraints in the end, we decided to abandon the 2.5d which consumes a lot of computing power and did not explore more window size, and now it seems that others got a boost on the larger windows.",
    "2174318": "Thank you for sharing solution.\nIt's interesting that the result shows 1D-CNN is more effective in stage2 than stage4.\nHow many window size do you use for 1D-CNN in stage 2 and 4?",
    "2174887": "Thank you for the feedback: indeed for me stage1 was already very effective in highlighting the people in contact, stage2 just had the effect of smoothing its values. However in stage3 the association of the players was much more noizy so the gain after this stage was much greater using time. For training I used 64 steps windows (for batch optimization) and then in inference mode I used the whole game as input",
    "2174966": "Thanks. One more question: Why stage 2 only focus on single player?\nDo you train the model to the contact event for a single player to any other players? (like [Qishen's solution](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391723)?)\n[\"game_play\",\"nfl_player_id\",\"step\"]\n\nDo you think modeling with contact event for single player is better than that of two players?",
    "2175348": "`Why stage 2 only focus on single player?`\nThis is because the predictions of stage1 are based on the contact of a single player.\n`Do you train the model to the contact event for a single player to any other players? (like Qishen's solution?)`\nYes.\n`Do you think modeling with contact event for a single player is better than that of two players?`\nI have seen many other solutions using the contact situation of two players as the label of the CNN that won the gold medal. Qishen only used the single players CNN in the first stage, but then used the two players CNN in the second stage. Our solution had no part of the two players' CNN at all, which is perhaps what we needed to get the gold medal. My initial thought was that the single player's image would be cropped out with less noise and the network would learn how to extract features more efficiently.",
    "2175470": "chenlin1999 Thanks for the explanation.\n\nLet me add my thought. Since the contact events in train dataset is few (although it seems plenty of them, most of them are very similar to each other), modeling mainly on player to player contact event easy to overfit. So I believe modeling with single player setting is more robust to overfitting to specific contact events than modeling with player to player contacts. My pipeline is mainly on XGBs, and the overfit was serious. CNNs seems more robust to overfitting as some participants says.",
    "2175481": "Yes I think you are right : I have the same thing in mind"
  },
  "source": "meta"
}