{
  "id": 391719,
  "title": "CNN can Visualize Contact (4th place K_mat part)",
  "url": "/competitions/nfl-player-contact-detection/discussion/391719",
  "author_name": "",
  "post_date": "2023-03-02T12:56:51.231527600Z",
  "votes": 49,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks to NFL and Kaggle for hosting such interesting challenge every year. Thanks <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> , <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>  and <a href=\"https://www.kaggle.com/hattan0523\" target=\"_blank\">@hattan0523</a> for teaming up with me. I really enjoyed and learned a lot from you guys! Congrats to the winners and everyone who enjoyed this competition!</p>\n<p>My main contributions are as follows:<br>\n<strong>1. Contact prediction by 2D CNN</strong><br>\n<strong>2. Feature engineering with bbox and tracking data such as the registration error.</strong><br>\nHere, I show the former.</p>\n<h1>Model Architecture</h1>\n<p>I applied 2D CNN(U-Net) to each player's image and predict some masks. Then the masks of each pair's area of intersection are multiplied and the maximum value of it is outputted as the contact prediction between two players. I expected the model to predict the contact event from the intersected area of each pair of players.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F117b7f73480671f35bd9c582e6f65fb7%2Ffig1.jpg?generation=1677759122353197&amp;alt=media\"></p>\n<p>Additionally, I created a slightly different model based on the same concept for ensemble. I input whole image into UNet and crop players in the feature space.</p>\n<h1>Characteristics</h1>\n<p>This model can learn and predict the contact from the intersection of two players’ area. It is beneficial in many aspects. </p>\n<ul>\n<li>The model can be trained efficiently.</li>\n<li>Quick inference since this model runs for the times of the number of players. (not number of pairs)</li>\n<li>Though this model isn’t trained to learn mask directly, it automatically becomes to predict the area of contact as shown in the figure below. Unfortunately, This didn’t help so much during competition but I believe it will help NFL people in their analysis.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fd41a5f5947368bef89a5266750eb865e%2Ffig2_sampleout.jpg?generation=1677759416178394&amp;alt=media\"></p>\n<p>The output of my model is used in the second stage. <br>\nOur team solution is <br>\n<a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391761\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391761</a></p>",
  "messages": [
    {
      "id": "2165776",
      "postDate": "03/02/2023 12:56:51",
      "content": "<p>Thanks to NFL and Kaggle for hosting such interesting challenge every year. Thanks <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> , <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>  and <a href=\"https://www.kaggle.com/hattan0523\" target=\"_blank\">@hattan0523</a> for teaming up with me. I really enjoyed and learned a lot from you guys! Congrats to the winners and everyone who enjoyed this competition!</p>\n<p>My main contributions are as follows:<br>\n<strong>1. Contact prediction by 2D CNN</strong><br>\n<strong>2. Feature engineering with bbox and tracking data such as the registration error.</strong><br>\nHere, I show the former.</p>\n<h1>Model Architecture</h1>\n<p>I applied 2D CNN(U-Net) to each player's image and predict some masks. Then the masks of each pair's area of intersection are multiplied and the maximum value of it is outputted as the contact prediction between two players. I expected the model to predict the contact event from the intersected area of each pair of players.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F117b7f73480671f35bd9c582e6f65fb7%2Ffig1.jpg?generation=1677759122353197&amp;alt=media\"></p>\n<p>Additionally, I created a slightly different model based on the same concept for ensemble. I input whole image into UNet and crop players in the feature space.</p>\n<h1>Characteristics</h1>\n<p>This model can learn and predict the contact from the intersection of two players’ area. It is beneficial in many aspects. </p>\n<ul>\n<li>The model can be trained efficiently.</li>\n<li>Quick inference since this model runs for the times of the number of players. (not number of pairs)</li>\n<li>Though this model isn’t trained to learn mask directly, it automatically becomes to predict the area of contact as shown in the figure below. Unfortunately, This didn’t help so much during competition but I believe it will help NFL people in their analysis.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fd41a5f5947368bef89a5266750eb865e%2Ffig2_sampleout.jpg?generation=1677759416178394&amp;alt=media\"></p>\n<p>The output of my model is used in the second stage. <br>\nOur team solution is <br>\n<a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391761\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391761</a></p>",
      "rawMarkdown": "Thanks to NFL and Kaggle for hosting such interesting challenge every year. Thanks @nyanpn , @bamps53  and @hattan0523 for teaming up with me. I really enjoyed and learned a lot from you guys! Congrats to the winners and everyone who enjoyed this competition!\n\nMy main contributions are as follows:\n**1. Contact prediction by 2D CNN**\n**2. Feature engineering with bbox and tracking data such as the registration error.**\nHere, I show the former.\n\n# Model Architecture\nI applied 2D CNN(U-Net) to each player's image and predict some masks. Then the masks of each pair's area of intersection are multiplied and the maximum value of it is outputted as the contact prediction between two players. I expected the model to predict the contact event from the intersected area of each pair of players.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F117b7f73480671f35bd9c582e6f65fb7%2Ffig1.jpg?generation=1677759122353197&alt=media\" width=\"1024\">\n\nAdditionally, I created a slightly different model based on the same concept for ensemble. I input whole image into UNet and crop players in the feature space.\n\n# Characteristics\nThis model can learn and predict the contact from the intersection of two players’ area. It is beneficial in many aspects. \n- The model can be trained efficiently.\n- Quick inference since this model runs for the times of the number of players. (not number of pairs)\n- Though this model isn’t trained to learn mask directly, it automatically becomes to predict the area of contact as shown in the figure below. Unfortunately, This didn’t help so much during competition but I believe it will help NFL people in their analysis.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fd41a5f5947368bef89a5266750eb865e%2Ffig2_sampleout.jpg?generation=1677759416178394&alt=media\" width=\"512\">\n\nThe output of my model is used in the second stage. \nOur team solution is \nhttps://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391761",
      "votes": null
    },
    {
      "id": "2167066",
      "postDate": "03/03/2023 08:41:37",
      "content": "<p><a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> congratulation for the 4th place.</p>\n<p>Does this model is trained end-to-end?</p>\n<p>Actually, I'm not clear about how model can learn where is the contact region is only from the feed back of log loss of contact or not. <br>\nThe dimension of the output of U-Net is NxN, on the contrary, the contact label is only 1 dimension.<br>\nHow to robustly feedback (back-props) the contact label information to the high dimension output channels (e.g. 0_a, 0_b, 0_c)?</p>",
      "rawMarkdown": "kmat2019 congratulation for the 4th place.\n\nDoes this model is trained end-to-end?\n\nActually, I'm not clear about how model can learn where is the contact region is only from the feed back of log loss of contact or not. \nThe dimension of the output of U-Net is NxN, on the contrary, the contact label is only 1 dimension.\nHow to robustly feedback (back-props) the contact label information to the high dimension output channels (e.g. 0_a, 0_b, 0_c)?",
      "votes": null
    },
    {
      "id": "2167530",
      "postDate": "03/03/2023 14:45:48",
      "content": "<p>Thank you.<br>\nYes, it's end to end.</p>\n<p>Let's start from simple exmaple. The code below is a much simpler model than mine.</p>\n<pre><code>input = image_two_players\nmask = UNet(input)\noutput = sigmoid(mask).max()\n</code></pre>\n<p>But I guess this model can predict where the contact is a little because the most of the contact information between two players exist there. If the model is a classifier to predict cat or dog, the mask area would be scattered in various places in the image. </p>\n<p>I designed my model to predict the contact area more explicitly than the model above. Model is forced to predict contact or not in the intersection area between two players.<br>\nNow, assuming there are three players in a image</p>\n<pre><code>input_1 = image_player_1\ninput_2 = image_player_2\ninput_3 = image_player_2\nmask_1 = UNet(input_1)\nmask_2 = UNet(input_2)\nmask_3 = UNet(input_3)\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_23 = intersection_multiply(mask_2, mask_3)\ncontact_mask_31 = intersection_multiply(mask_3, mask_1)\ncontact_12 = sigmoid(contact_mask_12).max()\ncontact_23 = sigmoid(contact_mask_23).max()\ncontact_31 = sigmoid(contact_mask_31).max()\n</code></pre>\n<p>If there is a contact between player_1 and player_2, contact_12 should be positive while the other contact_23 and contact_31 should be negative like the image below. As the model trained by many samples, it will be able to predict the contact area. Of cource, as you are concerning, this is not perfect. It sometimes doesn't work well.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F06dca7f4d64310faa6bfe72635202532%2Fcontact_logloss.jpg?generation=1677853666410519&amp;alt=media\" alt=\"\"></p>\n<p>I think it is important to control how the model learns. Once you build the architecture, just beleieve it.</p>",
      "rawMarkdown": "Thank you.\nYes, it's end to end.\n\nLet's start from simple exmaple. The code below is a much simpler model than mine.\n\n```\ninput = image_two_players\nmask = UNet(input)\noutput = sigmoid(mask).max()\n```\nBut I guess this model can predict where the contact is a little because the most of the contact information between two players exist there. If the model is a classifier to predict cat or dog, the mask area would be scattered in various places in the image. \n\nI designed my model to predict the contact area more explicitly than the model above. Model is forced to predict contact or not in the intersection area between two players.\nNow, assuming there are three players in a image\n```\ninput_1 = image_player_1\ninput_2 = image_player_2\ninput_3 = image_player_2\nmask_1 = UNet(input_1)\nmask_2 = UNet(input_2)\nmask_3 = UNet(input_3)\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_23 = intersection_multiply(mask_2, mask_3)\ncontact_mask_31 = intersection_multiply(mask_3, mask_1)\ncontact_12 = sigmoid(contact_mask_12).max()\ncontact_23 = sigmoid(contact_mask_23).max()\ncontact_31 = sigmoid(contact_mask_31).max()\n```\nIf there is a contact between player_1 and player_2, contact_12 should be positive while the other contact_23 and contact_31 should be negative like the image below. As the model trained by many samples, it will be able to predict the contact area. Of cource, as you are concerning, this is not perfect. It sometimes doesn't work well.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F06dca7f4d64310faa6bfe72635202532%2Fcontact_logloss.jpg?generation=1677853666410519&alt=media)\n\nI think it is important to control how the model learns. Once you build the architecture, just beleieve it.",
      "votes": null
    },
    {
      "id": "2167726",
      "postDate": "03/03/2023 17:12:34",
      "content": "<p>Please point out if I am wrong. In my understanding, when it comes to back-propagate phase, max pooling propagate gradient for only one pixel that has maximum pixel value. It seems to me training process could be very slow if you feed the contact mask in the original resolution. (However, since you seemed to successfully train the contact region in this setting, it should be O.K. somehow.)</p>\n<p>Well, this might be too niche to discuss.</p>\n<p>I roughly understand the main point of your approach is feeding the hint of contact region to the model by giving guidance of the intersection area of bounding box of the two players.</p>\n<p>Thank you for the detailed explanation.</p>",
      "rawMarkdown": "Please point out if I am wrong. In my understanding, when it comes to back-propagate phase, max pooling propagate gradient for only one pixel that has maximum pixel value. It seems to me training process could be very slow if you feed the contact mask in the original resolution. (However, since you seemed to successfully train the contact region in this setting, it should be O.K. somehow.)\n\nWell, this might be too niche to discuss.\n\nI roughly understand the main point of your approach is feeding the hint of contact region to the model by giving guidance of the intersection area of bounding box of the two players.\n\nThank you for the detailed explanation.",
      "votes": null
    },
    {
      "id": "2167732",
      "postDate": "03/03/2023 17:17:20",
      "content": "<p>I'm not sure I'm saying correct, but maybe lower resolution mask might be sufficient if the model can learn the roughly highlighted area of contact region.</p>",
      "rawMarkdown": "I'm not sure I'm saying correct, but maybe lower resolution mask might be sufficient if the model can learn the roughly highlighted area of contact region.",
      "votes": null
    },
    {
      "id": "2168212",
      "postDate": "03/04/2023 02:28:50",
      "content": "<blockquote>\n  <p>max pooling propagate gradient for only one pixel that has maximum pixel value</p>\n</blockquote>\n<p>You're right. But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification. </p>\n<p>I'm not sure I'm correct, but I think many convolutions propagate the loss widely even when the gradient occurs in only small section of output. Additionally, since filters of CNN is same in anywhere in the image, it doesn't matter so much where the loss occurs.</p>",
      "rawMarkdown": ">max pooling propagate gradient for only one pixel that has maximum pixel value\n\nYou're right. But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification. \n\nI'm not sure I'm correct, but I think many convolutions propagate the loss widely even when the gradient occurs in only small section of output. Additionally, since filters of CNN is same in anywhere in the image, it doesn't matter so much where the loss occurs.",
      "votes": null
    },
    {
      "id": "2168216",
      "postDate": "03/04/2023 02:32:13",
      "content": "<p>If you use lower resolution mask, you cannot predict the contact precisely. There would be many false positives and false negatives when multiplying every pair of players' low resolution mask.</p>",
      "rawMarkdown": "If you use lower resolution mask, you cannot predict the contact precisely. There would be many false positives and false negatives when multiplying every pair of players' low resolution mask.",
      "votes": null
    },
    {
      "id": "2168271",
      "postDate": "03/04/2023 04:16:38",
      "content": "<blockquote>\n  <p>But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification.</p>\n</blockquote>\n<p>You are right. I found a paper[1] which uses global max-pooling to detect object location in weakly supervised manner. Since this situation (contact location label doesn't provided) is similar to this competition's task, it makes sense to me.<br>\nThank you for pointing this out.</p>\n<p>[1] <a href=\"https://www.cv-foundation.org/openaccess/content_cvpr_2015/app/1A_075.pdf\" target=\"_blank\">Is object localization for free? –\nWeakly-supervised learning with convolutional neural networks</a></p>",
      "rawMarkdown": "> But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification.\n\nYou are right. I found a paper[1] which uses global max-pooling to detect object location in weakly supervised manner. Since this situation (contact location label doesn't provided) is similar to this competition's task, it makes sense to me.\nThank you for pointing this out.\n\n[1] [Is object localization for free? –\nWeakly-supervised learning with convolutional neural networks](https://www.cv-foundation.org/openaccess/content_cvpr_2015/app/1A_075.pdf)",
      "votes": null
    },
    {
      "id": "2168287",
      "postDate": "03/04/2023 04:46:22",
      "content": "<p>By the way, since this task is kind of weakly supervised training, I think it takes some amount of time to learn decent prediction. How much epoch does it take to train this task?</p>",
      "rawMarkdown": "By the way, since this task is kind of weakly supervised training, I think it takes some amount of time to learn decent prediction. How much epoch does it take to train this task?",
      "votes": null
    },
    {
      "id": "2168911",
      "postDate": "03/04/2023 16:20:49",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> and congrats on yet another strong finish!</p>\n<p>The model's ability to predict the specific region is really cool! This could have a lot of use cases for us. </p>\n<p>Am I understanding what you did correctly?</p>\n<ol>\n<li>You preprocessed the training data to extract a the bounding box of each player</li>\n<li>Then trained a model to predict the areas of intercetion.</li>\n<li>This helped the model identify the region of contact.</li>\n</ol>\n<p>Do you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.</p>\n<p>Thanks for sharing.</p>",
      "rawMarkdown": "Nice work @kmat2019 and congrats on yet another strong finish!\n\nThe model's ability to predict the specific region is really cool! This could have a lot of use cases for us. \n\nAm I understanding what you did correctly?\n1. You preprocessed the training data to extract a the bounding box of each player\n2. Then trained a model to predict the areas of intercetion.\n3. This helped the model identify the region of contact.\n\nDo you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.\n\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "2169405",
      "postDate": "03/05/2023 04:51:05",
      "content": "<p>Thank you! It's my pleasure to hear that as I focused on usability and simplicity this time.</p>\n<blockquote>\n  <p>1.You preprocessed the training data to extract a the bounding box of each player</p>\n</blockquote>\n<p>Exactly. I enlarged the bbox enough to contain a player.</p>\n<blockquote>\n  <p>2.Then trained a model to predict the areas of intercetion.</p>\n</blockquote>\n<p>Strictly speaking, my model predict there's contact or not for each players' pair. Please see the pseudo code below which is a simpler version of my model. The intermediate output just before taking the maximum shows the region of contact.</p>\n<blockquote>\n  <p>Do you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.</p>\n</blockquote>\n<p>Yes, that sounds interesting. Labels of the area are more informative than the binary label in this competition. So it would improve the accuracy of the contact prediction as well as the contact region.</p>\n<p>One concern is that labeling contact area is not easy because there are multiple contact regions in a pair and moreover, they sometimes ocluded by the another player. So I think identifying the body parts directry would be better than predicting the area in the image for the evaluation metric if it were a competition.</p>\n<hr>\n<p>pseudo code</p>\n<pre><code># preprocess \nimage_players = crop_resize_around_helmets_bbox(entire_image, helmets_bbox) \n\n# per player prediction. Three players as an example.\n# this mask means \"contact with other players\"\nmask_0 = UNet(image_players[0])\nmask_1 = UNet(image_players[1])\nmask_2 = UNet(image_players[2])\n\n# multiply each pair of masks considering the relative position in the frame\n# this mask means \"contact with a specific player\"\ncontact_mask_01 = intersection_multiply(mask_0, mask_1) # I visualized these outputs\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_20 = intersection_multiply(mask_2, mask_0)\n\n# take maximum at each intersection.\ncontact_01 = contact_mask_01.max()\ncontact_12 = contact_mask_12.max()\ncontact_20 = contact_mask_20.max()\n\nloss_01 = logloss(contact_01, ground_truth_01)\nloss_12 = logloss(contact_12, ground_truth_12)\nloss_20 = logloss(contact_20, ground_truth_20)\n</code></pre>\n<p>Additionally, I created a slightly different model based on the same concept for ensemble. <br>\nI input entire image into UNet and crop players in the feature space.</p>\n<pre><code># extract feature from whole image\nfeatures = UNet(entire_image)\nfeatures_players = crop_resize_around_helmets_bbox(features, helmets_bbox) # RoI\n\n# per player prediction. Three players as an example.\nmask_0 = shallow_CNN(features_players[0])\nmask_1 = shallow_CNN(features_players[1])\nmask_2 = shallow_CNN(features_players[2])\n</code></pre>",
      "rawMarkdown": "Thank you! It's my pleasure to hear that as I focused on usability and simplicity this time.\n\n> 1.You preprocessed the training data to extract a the bounding box of each player\n\nExactly. I enlarged the bbox enough to contain a player.\n\n> 2.Then trained a model to predict the areas of intercetion.\n\nStrictly speaking, my model predict there's contact or not for each players' pair. Please see the pseudo code below which is a simpler version of my model. The intermediate output just before taking the maximum shows the region of contact.\n\n> Do you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.\n\nYes, that sounds interesting. Labels of the area are more informative than the binary label in this competition. So it would improve the accuracy of the contact prediction as well as the contact region.\n\nOne concern is that labeling contact area is not easy because there are multiple contact regions in a pair and moreover, they sometimes ocluded by the another player. So I think identifying the body parts directry would be better than predicting the area in the image for the evaluation metric if it were a competition.\n\n---\npseudo code\n\n```\n# preprocess \nimage_players = crop_resize_around_helmets_bbox(entire_image, helmets_bbox) \n\n# per player prediction. Three players as an example.\n# this mask means \"contact with other players\"\nmask_0 = UNet(image_players[0])\nmask_1 = UNet(image_players[1])\nmask_2 = UNet(image_players[2])\n\n# multiply each pair of masks considering the relative position in the frame\n# this mask means \"contact with a specific player\"\ncontact_mask_01 = intersection_multiply(mask_0, mask_1) # I visualized these outputs\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_20 = intersection_multiply(mask_2, mask_0)\n\n# take maximum at each intersection.\ncontact_01 = contact_mask_01.max()\ncontact_12 = contact_mask_12.max()\ncontact_20 = contact_mask_20.max()\n\nloss_01 = logloss(contact_01, ground_truth_01)\nloss_12 = logloss(contact_12, ground_truth_12)\nloss_20 = logloss(contact_20, ground_truth_20)\n```\n\nAdditionally, I created a slightly different model based on the same concept for ensemble. \nI input entire image into UNet and crop players in the feature space.\n\n```\n# extract feature from whole image\nfeatures = UNet(entire_image)\nfeatures_players = crop_resize_around_helmets_bbox(features, helmets_bbox) # RoI\n\n# per player prediction. Three players as an example.\nmask_0 = shallow_CNN(features_players[0])\nmask_1 = shallow_CNN(features_players[1])\nmask_2 = shallow_CNN(features_players[2])\n```",
      "votes": null
    },
    {
      "id": "2173253",
      "postDate": "03/08/2023 08:23:39",
      "content": "<p>It was rather faster than the simple classification in my experiment. That's because I applied the prior knowledge of people that the contact occurs in the intersection. Models are trained for 20 epoch at the final submission but 10 to 15 is enough.</p>",
      "rawMarkdown": "It was rather faster than the simple classification in my experiment. That's because I applied the prior knowledge of people that the contact occurs in the intersection. Models are trained for 20 epoch at the final submission but 10 to 15 is enough.",
      "votes": null
    },
    {
      "id": "2173322",
      "postDate": "03/08/2023 10:10:05",
      "content": "<p>Thanks, it is interesting.</p>",
      "rawMarkdown": "Thanks, it is interesting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2167066,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "03/03/2023 08:41:37",
      "content": "<p><a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> congratulation for the 4th place.</p>\n<p>Does this model is trained end-to-end?</p>\n<p>Actually, I'm not clear about how model can learn where is the contact region is only from the feed back of log loss of contact or not. <br>\nThe dimension of the output of U-Net is NxN, on the contrary, the contact label is only 1 dimension.<br>\nHow to robustly feedback (back-props) the contact label information to the high dimension output channels (e.g. 0_a, 0_b, 0_c)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2167530,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "03/03/2023 14:45:48",
          "content": "<p>Thank you.<br>\nYes, it's end to end.</p>\n<p>Let's start from simple exmaple. The code below is a much simpler model than mine.</p>\n<pre><code>input = image_two_players\nmask = UNet(input)\noutput = sigmoid(mask).max()\n</code></pre>\n<p>But I guess this model can predict where the contact is a little because the most of the contact information between two players exist there. If the model is a classifier to predict cat or dog, the mask area would be scattered in various places in the image. </p>\n<p>I designed my model to predict the contact area more explicitly than the model above. Model is forced to predict contact or not in the intersection area between two players.<br>\nNow, assuming there are three players in a image</p>\n<pre><code>input_1 = image_player_1\ninput_2 = image_player_2\ninput_3 = image_player_2\nmask_1 = UNet(input_1)\nmask_2 = UNet(input_2)\nmask_3 = UNet(input_3)\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_23 = intersection_multiply(mask_2, mask_3)\ncontact_mask_31 = intersection_multiply(mask_3, mask_1)\ncontact_12 = sigmoid(contact_mask_12).max()\ncontact_23 = sigmoid(contact_mask_23).max()\ncontact_31 = sigmoid(contact_mask_31).max()\n</code></pre>\n<p>If there is a contact between player_1 and player_2, contact_12 should be positive while the other contact_23 and contact_31 should be negative like the image below. As the model trained by many samples, it will be able to predict the contact area. Of cource, as you are concerning, this is not perfect. It sometimes doesn't work well.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F06dca7f4d64310faa6bfe72635202532%2Fcontact_logloss.jpg?generation=1677853666410519&amp;alt=media\" alt=\"\"></p>\n<p>I think it is important to control how the model learns. Once you build the architecture, just beleieve it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2167726,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "03/03/2023 17:12:34",
              "content": "<p>Please point out if I am wrong. In my understanding, when it comes to back-propagate phase, max pooling propagate gradient for only one pixel that has maximum pixel value. It seems to me training process could be very slow if you feed the contact mask in the original resolution. (However, since you seemed to successfully train the contact region in this setting, it should be O.K. somehow.)</p>\n<p>Well, this might be too niche to discuss.</p>\n<p>I roughly understand the main point of your approach is feeding the hint of contact region to the model by giving guidance of the intersection area of bounding box of the two players.</p>\n<p>Thank you for the detailed explanation.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2167732,
                  "author_name": "tatamikenn",
                  "author_url": "",
                  "post_date": "03/03/2023 17:17:20",
                  "content": "<p>I'm not sure I'm saying correct, but maybe lower resolution mask might be sufficient if the model can learn the roughly highlighted area of contact region.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2168212,
                  "author_name": "kmat2019",
                  "author_url": "",
                  "post_date": "03/04/2023 02:28:50",
                  "content": "<blockquote>\n  <p>max pooling propagate gradient for only one pixel that has maximum pixel value</p>\n</blockquote>\n<p>You're right. But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification. </p>\n<p>I'm not sure I'm correct, but I think many convolutions propagate the loss widely even when the gradient occurs in only small section of output. Additionally, since filters of CNN is same in anywhere in the image, it doesn't matter so much where the loss occurs.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2168216,
                      "author_name": "kmat2019",
                      "author_url": "",
                      "post_date": "03/04/2023 02:32:13",
                      "content": "<p>If you use lower resolution mask, you cannot predict the contact precisely. There would be many false positives and false negatives when multiplying every pair of players' low resolution mask.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2168271,
                      "author_name": "tatamikenn",
                      "author_url": "",
                      "post_date": "03/04/2023 04:16:38",
                      "content": "<blockquote>\n  <p>But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification.</p>\n</blockquote>\n<p>You are right. I found a paper[1] which uses global max-pooling to detect object location in weakly supervised manner. Since this situation (contact location label doesn't provided) is similar to this competition's task, it makes sense to me.<br>\nThank you for pointing this out.</p>\n<p>[1] <a href=\"https://www.cv-foundation.org/openaccess/content_cvpr_2015/app/1A_075.pdf\" target=\"_blank\">Is object localization for free? –\nWeakly-supervised learning with convolutional neural networks</a></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2168287,
                          "author_name": "tatamikenn",
                          "author_url": "",
                          "post_date": "03/04/2023 04:46:22",
                          "content": "<p>By the way, since this task is kind of weakly supervised training, I think it takes some amount of time to learn decent prediction. How much epoch does it take to train this task?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2173253,
                              "author_name": "kmat2019",
                              "author_url": "",
                              "post_date": "03/08/2023 08:23:39",
                              "content": "<p>It was rather faster than the simple classification in my experiment. That's because I applied the prior knowledge of people that the contact occurs in the intersection. Models are trained for 20 epoch at the final submission but 10 to 15 is enough.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2173322,
                                  "author_name": "tatamikenn",
                                  "author_url": "",
                                  "post_date": "03/08/2023 10:10:05",
                                  "content": "<p>Thanks, it is interesting.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2168911,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/04/2023 16:20:49",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> and congrats on yet another strong finish!</p>\n<p>The model's ability to predict the specific region is really cool! This could have a lot of use cases for us. </p>\n<p>Am I understanding what you did correctly?</p>\n<ol>\n<li>You preprocessed the training data to extract a the bounding box of each player</li>\n<li>Then trained a model to predict the areas of intercetion.</li>\n<li>This helped the model identify the region of contact.</li>\n</ol>\n<p>Do you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.</p>\n<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169405,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "03/05/2023 04:51:05",
          "content": "<p>Thank you! It's my pleasure to hear that as I focused on usability and simplicity this time.</p>\n<blockquote>\n  <p>1.You preprocessed the training data to extract a the bounding box of each player</p>\n</blockquote>\n<p>Exactly. I enlarged the bbox enough to contain a player.</p>\n<blockquote>\n  <p>2.Then trained a model to predict the areas of intercetion.</p>\n</blockquote>\n<p>Strictly speaking, my model predict there's contact or not for each players' pair. Please see the pseudo code below which is a simpler version of my model. The intermediate output just before taking the maximum shows the region of contact.</p>\n<blockquote>\n  <p>Do you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.</p>\n</blockquote>\n<p>Yes, that sounds interesting. Labels of the area are more informative than the binary label in this competition. So it would improve the accuracy of the contact prediction as well as the contact region.</p>\n<p>One concern is that labeling contact area is not easy because there are multiple contact regions in a pair and moreover, they sometimes ocluded by the another player. So I think identifying the body parts directry would be better than predicting the area in the image for the evaluation metric if it were a competition.</p>\n<hr>\n<p>pseudo code</p>\n<pre><code># preprocess \nimage_players = crop_resize_around_helmets_bbox(entire_image, helmets_bbox) \n\n# per player prediction. Three players as an example.\n# this mask means \"contact with other players\"\nmask_0 = UNet(image_players[0])\nmask_1 = UNet(image_players[1])\nmask_2 = UNet(image_players[2])\n\n# multiply each pair of masks considering the relative position in the frame\n# this mask means \"contact with a specific player\"\ncontact_mask_01 = intersection_multiply(mask_0, mask_1) # I visualized these outputs\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_20 = intersection_multiply(mask_2, mask_0)\n\n# take maximum at each intersection.\ncontact_01 = contact_mask_01.max()\ncontact_12 = contact_mask_12.max()\ncontact_20 = contact_mask_20.max()\n\nloss_01 = logloss(contact_01, ground_truth_01)\nloss_12 = logloss(contact_12, ground_truth_12)\nloss_20 = logloss(contact_20, ground_truth_20)\n</code></pre>\n<p>Additionally, I created a slightly different model based on the same concept for ensemble. <br>\nI input entire image into UNet and crop players in the feature space.</p>\n<pre><code># extract feature from whole image\nfeatures = UNet(entire_image)\nfeatures_players = crop_resize_around_helmets_bbox(features, helmets_bbox) # RoI\n\n# per player prediction. Three players as an example.\nmask_0 = shallow_CNN(features_players[0])\nmask_1 = shallow_CNN(features_players[1])\nmask_2 = shallow_CNN(features_players[2])\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2165776": "Thanks to NFL and Kaggle for hosting such interesting challenge every year. Thanks @nyanpn , @bamps53  and @hattan0523 for teaming up with me. I really enjoyed and learned a lot from you guys! Congrats to the winners and everyone who enjoyed this competition!\n\nMy main contributions are as follows:\n**1. Contact prediction by 2D CNN**\n**2. Feature engineering with bbox and tracking data such as the registration error.**\nHere, I show the former.\n\n# Model Architecture\nI applied 2D CNN(U-Net) to each player's image and predict some masks. Then the masks of each pair's area of intersection are multiplied and the maximum value of it is outputted as the contact prediction between two players. I expected the model to predict the contact event from the intersected area of each pair of players.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F117b7f73480671f35bd9c582e6f65fb7%2Ffig1.jpg?generation=1677759122353197&alt=media\" width=\"1024\">\n\nAdditionally, I created a slightly different model based on the same concept for ensemble. I input whole image into UNet and crop players in the feature space.\n\n# Characteristics\nThis model can learn and predict the contact from the intersection of two players’ area. It is beneficial in many aspects. \n- The model can be trained efficiently.\n- Quick inference since this model runs for the times of the number of players. (not number of pairs)\n- Though this model isn’t trained to learn mask directly, it automatically becomes to predict the area of contact as shown in the figure below. Unfortunately, This didn’t help so much during competition but I believe it will help NFL people in their analysis.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fd41a5f5947368bef89a5266750eb865e%2Ffig2_sampleout.jpg?generation=1677759416178394&alt=media\" width=\"512\">\n\nThe output of my model is used in the second stage. \nOur team solution is \nhttps://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391761",
    "2167066": "kmat2019 congratulation for the 4th place.\n\nDoes this model is trained end-to-end?\n\nActually, I'm not clear about how model can learn where is the contact region is only from the feed back of log loss of contact or not. \nThe dimension of the output of U-Net is NxN, on the contrary, the contact label is only 1 dimension.\nHow to robustly feedback (back-props) the contact label information to the high dimension output channels (e.g. 0_a, 0_b, 0_c)?",
    "2167530": "Thank you.\nYes, it's end to end.\n\nLet's start from simple exmaple. The code below is a much simpler model than mine.\n\n```\ninput = image_two_players\nmask = UNet(input)\noutput = sigmoid(mask).max()\n```\nBut I guess this model can predict where the contact is a little because the most of the contact information between two players exist there. If the model is a classifier to predict cat or dog, the mask area would be scattered in various places in the image. \n\nI designed my model to predict the contact area more explicitly than the model above. Model is forced to predict contact or not in the intersection area between two players.\nNow, assuming there are three players in a image\n```\ninput_1 = image_player_1\ninput_2 = image_player_2\ninput_3 = image_player_2\nmask_1 = UNet(input_1)\nmask_2 = UNet(input_2)\nmask_3 = UNet(input_3)\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_23 = intersection_multiply(mask_2, mask_3)\ncontact_mask_31 = intersection_multiply(mask_3, mask_1)\ncontact_12 = sigmoid(contact_mask_12).max()\ncontact_23 = sigmoid(contact_mask_23).max()\ncontact_31 = sigmoid(contact_mask_31).max()\n```\nIf there is a contact between player_1 and player_2, contact_12 should be positive while the other contact_23 and contact_31 should be negative like the image below. As the model trained by many samples, it will be able to predict the contact area. Of cource, as you are concerning, this is not perfect. It sometimes doesn't work well.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F06dca7f4d64310faa6bfe72635202532%2Fcontact_logloss.jpg?generation=1677853666410519&alt=media)\n\nI think it is important to control how the model learns. Once you build the architecture, just beleieve it.",
    "2167726": "Please point out if I am wrong. In my understanding, when it comes to back-propagate phase, max pooling propagate gradient for only one pixel that has maximum pixel value. It seems to me training process could be very slow if you feed the contact mask in the original resolution. (However, since you seemed to successfully train the contact region in this setting, it should be O.K. somehow.)\n\nWell, this might be too niche to discuss.\n\nI roughly understand the main point of your approach is feeding the hint of contact region to the model by giving guidance of the intersection area of bounding box of the two players.\n\nThank you for the detailed explanation.",
    "2167732": "I'm not sure I'm saying correct, but maybe lower resolution mask might be sufficient if the model can learn the roughly highlighted area of contact region.",
    "2168212": ">max pooling propagate gradient for only one pixel that has maximum pixel value\n\nYou're right. But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification. \n\nI'm not sure I'm correct, but I think many convolutions propagate the loss widely even when the gradient occurs in only small section of output. Additionally, since filters of CNN is same in anywhere in the image, it doesn't matter so much where the loss occurs.",
    "2168216": "If you use lower resolution mask, you cannot predict the contact precisely. There would be many false positives and false negatives when multiplying every pair of players' low resolution mask.",
    "2168271": "> But there are many implementation using maxpooling in CNN and they are working well such as a globalmaxpooling in classification.\n\nYou are right. I found a paper[1] which uses global max-pooling to detect object location in weakly supervised manner. Since this situation (contact location label doesn't provided) is similar to this competition's task, it makes sense to me.\nThank you for pointing this out.\n\n[1] [Is object localization for free? –\nWeakly-supervised learning with convolutional neural networks](https://www.cv-foundation.org/openaccess/content_cvpr_2015/app/1A_075.pdf)",
    "2168287": "By the way, since this task is kind of weakly supervised training, I think it takes some amount of time to learn decent prediction. How much epoch does it take to train this task?",
    "2168911": "Nice work @kmat2019 and congrats on yet another strong finish!\n\nThe model's ability to predict the specific region is really cool! This could have a lot of use cases for us. \n\nAm I understanding what you did correctly?\n1. You preprocessed the training data to extract a the bounding box of each player\n2. Then trained a model to predict the areas of intercetion.\n3. This helped the model identify the region of contact.\n\nDo you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.\n\nThanks for sharing.",
    "2169405": "Thank you! It's my pleasure to hear that as I focused on usability and simplicity this time.\n\n> 1.You preprocessed the training data to extract a the bounding box of each player\n\nExactly. I enlarged the bbox enough to contain a player.\n\n> 2.Then trained a model to predict the areas of intercetion.\n\nStrictly speaking, my model predict there's contact or not for each players' pair. Please see the pseudo code below which is a simpler version of my model. The intermediate output just before taking the maximum shows the region of contact.\n\n> Do you think this could be improved if we had a labeled dataset with the areas of contact annotated by a human? I could see this as a potential idea for future competitions if we could accurately identify the body parts involved in each contact.\n\nYes, that sounds interesting. Labels of the area are more informative than the binary label in this competition. So it would improve the accuracy of the contact prediction as well as the contact region.\n\nOne concern is that labeling contact area is not easy because there are multiple contact regions in a pair and moreover, they sometimes ocluded by the another player. So I think identifying the body parts directry would be better than predicting the area in the image for the evaluation metric if it were a competition.\n\n---\npseudo code\n\n```\n# preprocess \nimage_players = crop_resize_around_helmets_bbox(entire_image, helmets_bbox) \n\n# per player prediction. Three players as an example.\n# this mask means \"contact with other players\"\nmask_0 = UNet(image_players[0])\nmask_1 = UNet(image_players[1])\nmask_2 = UNet(image_players[2])\n\n# multiply each pair of masks considering the relative position in the frame\n# this mask means \"contact with a specific player\"\ncontact_mask_01 = intersection_multiply(mask_0, mask_1) # I visualized these outputs\ncontact_mask_12 = intersection_multiply(mask_1, mask_2)\ncontact_mask_20 = intersection_multiply(mask_2, mask_0)\n\n# take maximum at each intersection.\ncontact_01 = contact_mask_01.max()\ncontact_12 = contact_mask_12.max()\ncontact_20 = contact_mask_20.max()\n\nloss_01 = logloss(contact_01, ground_truth_01)\nloss_12 = logloss(contact_12, ground_truth_12)\nloss_20 = logloss(contact_20, ground_truth_20)\n```\n\nAdditionally, I created a slightly different model based on the same concept for ensemble. \nI input entire image into UNet and crop players in the feature space.\n\n```\n# extract feature from whole image\nfeatures = UNet(entire_image)\nfeatures_players = crop_resize_around_helmets_bbox(features, helmets_bbox) # RoI\n\n# per player prediction. Three players as an example.\nmask_0 = shallow_CNN(features_players[0])\nmask_1 = shallow_CNN(features_players[1])\nmask_2 = shallow_CNN(features_players[2])\n```",
    "2173253": "It was rather faster than the simple classification in my experiment. That's because I applied the prior knowledge of people that the contact occurs in the intersection. Models are trained for 20 epoch at the final submission but 10 to 15 is enough.",
    "2173322": "Thanks, it is interesting."
  },
  "source": "meta"
}