{
  "id": 391609,
  "title": "14th Place Solution",
  "url": "/competitions/nfl-player-contact-detection/writeups/kurupical-14th-place-solution",
  "author_name": "",
  "post_date": "2023-03-07T11:04:04.257Z",
  "votes": 39,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, I would like to thank our hosts for organizing the competition.<br>\nIt was a task I've never solved before, and it was both educational and a lot of fun trying different approaches!</p>\n<h1>Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fe323a35a9255e8e50131807aa19b56ef%2Fsolution.png?generation=1677715798541475&amp;alt=media\" alt=\"\"></p>\n<h1>Model Detail</h1>\n<h2>3D-CNN (cv: 0.770)</h2>\n<ul>\n<li>backbone: r3d_18 (from torchvision: <a href=\"https://pytorch.org/vision/stable/models/generated/torchvision.models.video.r3d_18.html#torchvision.models.video.R3D_18_Weights\" target=\"_blank\">https://pytorch.org/vision/stable/models/generated/torchvision.models.video.r3d_18.html#torchvision.models.video.R3D_18_Weights</a>)</li>\n<li>use 63 frames(20fps)</li>\n<li>predict 19 steps</li>\n<li>train every 9 steps</li>\n<li>StepLR Scheduler(~2epochs: lr=1e-3/1e-4)</li>\n</ul>\n<h2>2.5D3D-CNN (cv: 0.768)</h2>\n<ul>\n<li>Almost same as DFL's 1st solution by Team Hydrogen (<a href=\"https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932\" target=\"_blank\">https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932</a>)</li>\n<li>backbone: legacy_seresnet34</li>\n<li>use 123 frames(20fps)</li>\n<li>predict 3 frames</li>\n<li>down sampling (g: 10%, contact: 30%)</li>\n<li>label smoothing (0.1-0.9)</li>\n</ul>\n<h2>Both 3D, 2.5D3D</h2>\n<ul>\n<li>Linear layers for g and contact</li>\n</ul>\n<pre><code>\nx_contact = model_contact(x)  \nx_g = model_g(x)  \nnot_is_g = (is_g == )\nx = x_contact * not_is_g + x_g * is_g  \n</code></pre>\n<ul>\n<li>output 3 prediction and calculate loss: only sideline, only endzone, concat sideline-endzone feature.</li>\n</ul>\n<pre><code>\n ():\n  x_sideline = cnn(x_sideline_image) \n  x_endzone = cnn(x_endzone_image)\n   fc(torch.cat([x_sideline, x_endzone])), fc_sideline(x_sideline), fc_endzone(x_endzone)\n</code></pre>\n<h2>LGBM (cv: 0.740)</h2>\n<ul>\n<li>about 1100 features</li>\n<li>feautres<ul>\n<li>player's distance (tracking, helmet)</li>\n<li>lag, diff</li>\n<li>top_n nearest player's distance (n: parameters)</li>\n<li>number of people within distance n (n: parameters)</li></ul></li>\n<li>groupby<ul>\n<li>game_play</li>\n<li>is_g</li>\n<li>is_same_team</li>\n<li>number of people within distance n </li></ul></li>\n</ul>\n<h2>ensemble</h2>\n<p>Weighted ensemble, G and contact respectively.</p>\n<h2>What worked for me</h2>\n<ul>\n<li>image preprocessing<ul>\n<li>draw bbox -&gt; draw bbox and paint out</li>\n<li>use 2 colors(g, contact) -&gt; use 3 colors(g, same team contact, different team contact)</li>\n<li>crop the image with keeping the aspect ratio</li></ul></li>\n</ul>\n<pre><code>  bbox_left_ratio = \n  bbox_right_ratio = \n  bbox_top_ratio = \n  bbox_down_ratio = \n   col  [, , , ]:\n      df[col] = df[[, ]].mean(axis=)\n  df[] = df[[, ]].mean(axis=)\n  df[] = df.groupby([, , ])[].transform()\n\n  series = df.iloc[]  \n  left = (series[] - series[] * bbox_left_ratio)\n  right = (series[] + series[] * bbox_right_ratio)\n  top = (series[] + series[] * bbox_top_ratio)\n  down = (series[] - series[] * bbox_down_ratio)\n  img = img[down:top, left:right]\n  img = cv2.resize(img, (, ))\n</code></pre>\n<ul>\n<li>StepLR with warmup scheduler</li>\n<li>label smoothing (worked for 2.5D3D, but not worked for 3D)</li>\n</ul>\n<h2>What not worked for me</h2>\n<ul>\n<li>Transformers<ul>\n<li>use top 100~400 features of lgbm feature importances</li>\n<li>tuned hard but got cv 0.02 lower than lgbm.</li></ul></li>\n<li>2D-&gt;1D CNN<ul>\n<li>contact score is same as 2.5D3D, 3D but very poor G score in my work.</li></ul></li>\n<li>interpolate bbox</li>\n</ul>\n<h2>Other</h2>\n<ul>\n<li>tools: I make tools to investigate wrong inference and make a hypothesize to improve score.<br>\n<a href=\"https://github.com/kurupical/nfl_contact_detection/blob/master/58218_003210_contact_0.506591796875_score0.0_H23_V10.gif\" target=\"_blank\">https://github.com/kurupical/nfl_contact_detection/blob/master/58218_003210_contact_0.506591796875_score0.0_H23_V10.gif</a></li>\n</ul>",
  "messages": [
    {
      "id": "2165032",
      "postDate": "03/02/2023 00:22:07",
      "content": "<p>First of all, I would like to thank our hosts for organizing the competition.<br>\nIt was a task I've never solved before, and it was both educational and a lot of fun trying different approaches!</p>\n<h1>Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fe323a35a9255e8e50131807aa19b56ef%2Fsolution.png?generation=1677715798541475&amp;alt=media\" alt=\"\"></p>\n<h1>Model Detail</h1>\n<h2>3D-CNN (cv: 0.770)</h2>\n<ul>\n<li>backbone: r3d_18 (from torchvision: <a href=\"https://pytorch.org/vision/stable/models/generated/torchvision.models.video.r3d_18.html#torchvision.models.video.R3D_18_Weights\" target=\"_blank\">https://pytorch.org/vision/stable/models/generated/torchvision.models.video.r3d_18.html#torchvision.models.video.R3D_18_Weights</a>)</li>\n<li>use 63 frames(20fps)</li>\n<li>predict 19 steps</li>\n<li>train every 9 steps</li>\n<li>StepLR Scheduler(~2epochs: lr=1e-3/1e-4)</li>\n</ul>\n<h2>2.5D3D-CNN (cv: 0.768)</h2>\n<ul>\n<li>Almost same as DFL's 1st solution by Team Hydrogen (<a href=\"https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932\" target=\"_blank\">https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932</a>)</li>\n<li>backbone: legacy_seresnet34</li>\n<li>use 123 frames(20fps)</li>\n<li>predict 3 frames</li>\n<li>down sampling (g: 10%, contact: 30%)</li>\n<li>label smoothing (0.1-0.9)</li>\n</ul>\n<h2>Both 3D, 2.5D3D</h2>\n<ul>\n<li>Linear layers for g and contact</li>\n</ul>\n<pre><code>\nx_contact = model_contact(x)  \nx_g = model_g(x)  \nnot_is_g = (is_g == )\nx = x_contact * not_is_g + x_g * is_g  \n</code></pre>\n<ul>\n<li>output 3 prediction and calculate loss: only sideline, only endzone, concat sideline-endzone feature.</li>\n</ul>\n<pre><code>\n ():\n  x_sideline = cnn(x_sideline_image) \n  x_endzone = cnn(x_endzone_image)\n   fc(torch.cat([x_sideline, x_endzone])), fc_sideline(x_sideline), fc_endzone(x_endzone)\n</code></pre>\n<h2>LGBM (cv: 0.740)</h2>\n<ul>\n<li>about 1100 features</li>\n<li>feautres<ul>\n<li>player's distance (tracking, helmet)</li>\n<li>lag, diff</li>\n<li>top_n nearest player's distance (n: parameters)</li>\n<li>number of people within distance n (n: parameters)</li></ul></li>\n<li>groupby<ul>\n<li>game_play</li>\n<li>is_g</li>\n<li>is_same_team</li>\n<li>number of people within distance n </li></ul></li>\n</ul>\n<h2>ensemble</h2>\n<p>Weighted ensemble, G and contact respectively.</p>\n<h2>What worked for me</h2>\n<ul>\n<li>image preprocessing<ul>\n<li>draw bbox -&gt; draw bbox and paint out</li>\n<li>use 2 colors(g, contact) -&gt; use 3 colors(g, same team contact, different team contact)</li>\n<li>crop the image with keeping the aspect ratio</li></ul></li>\n</ul>\n<pre><code>  bbox_left_ratio = \n  bbox_right_ratio = \n  bbox_top_ratio = \n  bbox_down_ratio = \n   col  [, , , ]:\n      df[col] = df[[, ]].mean(axis=)\n  df[] = df[[, ]].mean(axis=)\n  df[] = df.groupby([, , ])[].transform()\n\n  series = df.iloc[]  \n  left = (series[] - series[] * bbox_left_ratio)\n  right = (series[] + series[] * bbox_right_ratio)\n  top = (series[] + series[] * bbox_top_ratio)\n  down = (series[] - series[] * bbox_down_ratio)\n  img = img[down:top, left:right]\n  img = cv2.resize(img, (, ))\n</code></pre>\n<ul>\n<li>StepLR with warmup scheduler</li>\n<li>label smoothing (worked for 2.5D3D, but not worked for 3D)</li>\n</ul>\n<h2>What not worked for me</h2>\n<ul>\n<li>Transformers<ul>\n<li>use top 100~400 features of lgbm feature importances</li>\n<li>tuned hard but got cv 0.02 lower than lgbm.</li></ul></li>\n<li>2D-&gt;1D CNN<ul>\n<li>contact score is same as 2.5D3D, 3D but very poor G score in my work.</li></ul></li>\n<li>interpolate bbox</li>\n</ul>\n<h2>Other</h2>\n<ul>\n<li>tools: I make tools to investigate wrong inference and make a hypothesize to improve score.<br>\n<a href=\"https://github.com/kurupical/nfl_contact_detection/blob/master/58218_003210_contact_0.506591796875_score0.0_H23_V10.gif\" target=\"_blank\">https://github.com/kurupical/nfl_contact_detection/blob/master/58218_003210_contact_0.506591796875_score0.0_H23_V10.gif</a></li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank our hosts for organizing the competition.\nIt was a task I've never solved before, and it was both educational and a lot of fun trying different approaches!\n\n# Summary\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fe323a35a9255e8e50131807aa19b56ef%2Fsolution.png?generation=1677715798541475&alt=media)\n\n# Model Detail\n## 3D-CNN (cv: 0.770)\n- backbone: r3d_18 (from torchvision: https://pytorch.org/vision/stable/models/generated/torchvision.models.video.r3d_18.html#torchvision.models.video.R3D_18_Weights)\n- use 63 frames(20fps)\n- predict 19 steps\n- train every 9 steps\n- StepLR Scheduler(~2epochs: lr=1e-3/1e-4)\n\n## 2.5D3D-CNN (cv: 0.768)\n- Almost same as DFL's 1st solution by Team Hydrogen (https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932)\n- backbone: legacy_seresnet34\n- use 123 frames(20fps)\n- predict 3 frames\n- down sampling (g: 10%, contact: 30%)\n- label smoothing (0.1-0.9)\n\n## Both 3D, 2.5D3D\n- Linear layers for g and contact\n```python\n# x: (bs, cnn_features)\nx_contact = model_contact(x)  # (bs, n_predict_frames)\nx_g = model_g(x)  # (bs, n_predict_frames)\nnot_is_g = (is_g == 0)\nx = x_contact * not_is_g + x_g * is_g  # (bs, n_predict_frames)\n```\n- output 3 prediction and calculate loss: only sideline, only endzone, concat sideline-endzone feature.\n```python\n# pseudo code\ndef forward(self, x_sideline_image, x_endzone_image):\n  x_sideline = cnn(x_sideline_image) \n  x_endzone = cnn(x_endzone_image)\n  return fc(torch.cat([x_sideline, x_endzone])), fc_sideline(x_sideline), fc_endzone(x_endzone)\n```\n\n## LGBM (cv: 0.740)\n- about 1100 features\n- feautres\n    - player's distance (tracking, helmet)\n    - lag, diff\n    - top_n nearest player's distance (n: parameters)\n    - number of people within distance n (n: parameters)\n- groupby\n    - game_play\n    - is_g\n    - is_same_team\n    - number of people within distance n \n\n## ensemble\nWeighted ensemble, G and contact respectively.\n\n## What worked for me\n- image preprocessing\n  - draw bbox -> draw bbox and paint out\n  - use 2 colors(g, contact) -> use 3 colors(g, same team contact, different team contact)\n  - crop the image with keeping the aspect ratio\n  ```python\n  bbox_left_ratio = 4.5\n  bbox_right_ratio = 4.5\n  bbox_top_ratio = 4.5\n  bbox_down_ratio = 2.25\n  for col in [\"x\", \"y\", \"width\", \"height\"]:\n      df[col] = df[[f\"{col}_1\", f\"{col}_2\"]].mean(axis=1)\n  df[\"bbox_size\"] = df[[\"width\", \"height\"]].mean(axis=1)\n  df[\"bbox_size\"] = df.groupby([\"view\", \"step\", \"game_play\"])[\"bbox_size\"].transform(\"mean\")\n\n  series = df.iloc[0]  # sample\n  left = int(series[\"x\"] - series[\"bbox_size\"] * bbox_left_ratio)\n  right = int(series[\"x\"] + series[\"bbox_size\"] * bbox_right_ratio)\n  top = int(series[\"y\"] + series[\"bbox_size\"] * bbox_top_ratio)\n  down = int(series[\"y\"] - series[\"bbox_size\"] * bbox_down_ratio)\n  img = img[down:top, left:right]\n  img = cv2.resize(img, (128, 96))\n  ```\n  \n- StepLR with warmup scheduler\n- label smoothing (worked for 2.5D3D, but not worked for 3D)\n\n## What not worked for me\n- Transformers\n  - use top 100~400 features of lgbm feature importances\n  - tuned hard but got cv 0.02 lower than lgbm.\n- 2D->1D CNN\n  - contact score is same as 2.5D3D, 3D but very poor G score in my work.\n- interpolate bbox\n\n\n## Other\n- tools: I make tools to investigate wrong inference and make a hypothesize to improve score.\nhttps://github.com/kurupical/nfl_contact_detection/blob/master/58218_003210_contact_0.506591796875_score0.0_H23_V10.gif",
      "votes": null
    },
    {
      "id": "2165136",
      "postDate": "03/02/2023 02:29:05",
      "content": "<p>First of all congratulations, thanks for sharing this great solution. I have some questions that in 3D CNN you predict 19 steps, that is, given that input sequence you will predict which step it will belong to? (classification)</p>",
      "rawMarkdown": "First of all congratulations, thanks for sharing this great solution. I have some questions that in 3D CNN you predict 19 steps, that is, given that input sequence you will predict which step it will belong to? (classification)",
      "votes": null
    },
    {
      "id": "2168963",
      "postDate": "03/04/2023 16:59:24",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kurupical\" target=\"_blank\">@kurupical</a> and thanks for sharing this solution write up.</p>",
      "rawMarkdown": "Great work @kurupical and thanks for sharing this solution write up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2165136,
      "author_name": "nhtnguyntrng",
      "author_url": "",
      "post_date": "03/02/2023 02:29:05",
      "content": "<p>First of all congratulations, thanks for sharing this great solution. I have some questions that in 3D CNN you predict 19 steps, that is, given that input sequence you will predict which step it will belong to? (classification)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2168963,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:59:24",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kurupical\" target=\"_blank\">@kurupical</a> and thanks for sharing this solution write up.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2165032": "First of all, I would like to thank our hosts for organizing the competition.\nIt was a task I've never solved before, and it was both educational and a lot of fun trying different approaches!\n\n# Summary\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fe323a35a9255e8e50131807aa19b56ef%2Fsolution.png?generation=1677715798541475&alt=media)\n\n# Model Detail\n## 3D-CNN (cv: 0.770)\n- backbone: r3d_18 (from torchvision: https://pytorch.org/vision/stable/models/generated/torchvision.models.video.r3d_18.html#torchvision.models.video.R3D_18_Weights)\n- use 63 frames(20fps)\n- predict 19 steps\n- train every 9 steps\n- StepLR Scheduler(~2epochs: lr=1e-3/1e-4)\n\n## 2.5D3D-CNN (cv: 0.768)\n- Almost same as DFL's 1st solution by Team Hydrogen (https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout/discussion/359932)\n- backbone: legacy_seresnet34\n- use 123 frames(20fps)\n- predict 3 frames\n- down sampling (g: 10%, contact: 30%)\n- label smoothing (0.1-0.9)\n\n## Both 3D, 2.5D3D\n- Linear layers for g and contact\n```python\n# x: (bs, cnn_features)\nx_contact = model_contact(x)  # (bs, n_predict_frames)\nx_g = model_g(x)  # (bs, n_predict_frames)\nnot_is_g = (is_g == 0)\nx = x_contact * not_is_g + x_g * is_g  # (bs, n_predict_frames)\n```\n- output 3 prediction and calculate loss: only sideline, only endzone, concat sideline-endzone feature.\n```python\n# pseudo code\ndef forward(self, x_sideline_image, x_endzone_image):\n  x_sideline = cnn(x_sideline_image) \n  x_endzone = cnn(x_endzone_image)\n  return fc(torch.cat([x_sideline, x_endzone])), fc_sideline(x_sideline), fc_endzone(x_endzone)\n```\n\n## LGBM (cv: 0.740)\n- about 1100 features\n- feautres\n    - player's distance (tracking, helmet)\n    - lag, diff\n    - top_n nearest player's distance (n: parameters)\n    - number of people within distance n (n: parameters)\n- groupby\n    - game_play\n    - is_g\n    - is_same_team\n    - number of people within distance n \n\n## ensemble\nWeighted ensemble, G and contact respectively.\n\n## What worked for me\n- image preprocessing\n  - draw bbox -> draw bbox and paint out\n  - use 2 colors(g, contact) -> use 3 colors(g, same team contact, different team contact)\n  - crop the image with keeping the aspect ratio\n  ```python\n  bbox_left_ratio = 4.5\n  bbox_right_ratio = 4.5\n  bbox_top_ratio = 4.5\n  bbox_down_ratio = 2.25\n  for col in [\"x\", \"y\", \"width\", \"height\"]:\n      df[col] = df[[f\"{col}_1\", f\"{col}_2\"]].mean(axis=1)\n  df[\"bbox_size\"] = df[[\"width\", \"height\"]].mean(axis=1)\n  df[\"bbox_size\"] = df.groupby([\"view\", \"step\", \"game_play\"])[\"bbox_size\"].transform(\"mean\")\n\n  series = df.iloc[0]  # sample\n  left = int(series[\"x\"] - series[\"bbox_size\"] * bbox_left_ratio)\n  right = int(series[\"x\"] + series[\"bbox_size\"] * bbox_right_ratio)\n  top = int(series[\"y\"] + series[\"bbox_size\"] * bbox_top_ratio)\n  down = int(series[\"y\"] - series[\"bbox_size\"] * bbox_down_ratio)\n  img = img[down:top, left:right]\n  img = cv2.resize(img, (128, 96))\n  ```\n  \n- StepLR with warmup scheduler\n- label smoothing (worked for 2.5D3D, but not worked for 3D)\n\n## What not worked for me\n- Transformers\n  - use top 100~400 features of lgbm feature importances\n  - tuned hard but got cv 0.02 lower than lgbm.\n- 2D->1D CNN\n  - contact score is same as 2.5D3D, 3D but very poor G score in my work.\n- interpolate bbox\n\n\n## Other\n- tools: I make tools to investigate wrong inference and make a hypothesize to improve score.\nhttps://github.com/kurupical/nfl_contact_detection/blob/master/58218_003210_contact_0.506591796875_score0.0_H23_V10.gif",
    "2165136": "First of all congratulations, thanks for sharing this great solution. I have some questions that in 3D CNN you predict 19 steps, that is, given that input sequence you will predict which step it will belong to? (classification)",
    "2168963": "Great work @kurupical and thanks for sharing this solution write up."
  },
  "source": "meta"
}