{
  "id": 392392,
  "title": "How to fix predicted helmets assignment",
  "url": "/competitions/nfl-player-contact-detection/discussion/392392",
  "author_name": "",
  "post_date": "2023-03-05T04:34:13.586591900Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I spend a significant part of time working on this competition to try to improve the helmets/players associations.<br>\nI found this was an even more interesting task, I regret I have not participated in the previous competition and given the close score of the top solution, I hoped it had the potential to boost the score. It improved the score on some folds by around 0.005 and had no impact on other folds (I guess depending on the number of the fixed helmet predictions). Fixing helmets did not improve the private leaderboard score, so I'll keep the approach in a separate post.</p>\n<p>The initial idea was to fit the perspective or affine transformation between the tracking and video positions using something like RANSAC regressor which fits only subset of points. It worked for individual wrong player assignments but failed for frames with the small number of players and/or many mistakes. It would create completely unrealistic transformations.</p>\n<p>One way I tried to improve it is to enforce some kind of continuity between step transformations by adding frame corners and maintaining the constant speed for the frame corners (transformation between corners are added to RANSAC regression). After every transformation fit with RANSAC I removed associations where the predicted position (as estrimated_transform * tracking_pos) does not match the helmet pos on the video.<br>\nAfter I did a recursive search to find a better association between unassociated yet players and helmets (minimising the total square distance between predicted and video pos).</p>\n<p>It helped a bit but was not sufficient for frames with many mistakes (usually the second part of the game with a few player running and many wrongly detected helmets outside of the field).</p>\n<p>Next approach was to fit the camera transformation parameters instead of the unrestricted perspective or affine transformation.<br>\nI'd optimize for the camera position, point in the field camera points to and the current focal length.<br>\nI optimised for MAE of predicted and measured helmet position with an extra cost for the second order of the camera parameter changes between frames to enforce the smooth movements (using Pytorh, Adam optimiser over the whole game).<br>\nI did the initial camera parameters, a few rounds of optimisations, re-assigned helmets as above for the obviously wrong assignment and continued optimisation.</p>\n<p>This helped a bit but a few frames of completely wrong baseline associations pushed camera to the wrong direction for a few video files.<br>\nThe issue was for such games the predicted transformation movement did not match the visible optical flow movement, so I decided to try predicting the optical flow between frames.</p>\n<p>The model receives 2 frames as inputs and predict the ground movement optical field not taking into account players movement.<br>\nThe model is based on 2 resnet34 with TSM like links between.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F423d85e9ccedc596b92364d7af76be93%2Fflow_model.png?generation=1677990657178509&amp;alt=media\" alt=\"\"></p>\n<p>It's trained on either random transformation between the selected frame or two frames with transformation estimated by the previous camera parameters based model.<br>\nIt's important to train not only on syntetic pairs of images but also on actual different frames so model would learn to ignore players movement and align only the field markings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Ffb128e72af041fa40c06a9fbdbd6cd18%2Fflow_prediction_alignment.png?generation=1677990686290014&amp;alt=media\" alt=\"\"></p>\n<p>With the information about the optical flow between frames, the players re-association solution worked much more reliably and after visual inspection I found it fixed almost all player association errors.</p>\n<p>The solution was:</p>\n<p>1) For the initial frame with the good set of predicted players, build the initial transformation using RANSAC. Re-assign outliers using approach described above.<br>\n2) For the next frame, apply RANSAC not only on the tracking pos -&gt; video position pairs but also a seet of points from the last frame with the flow applied and projected to the ground (flow corrected view points should point to the same tracking space pos). This allows to restrict movement between frames to be close to predicted flow but to also correct the accumulated flow errors to better match helments. Re-assign outliers using approach described above.<br>\n3) Run the full game optimisation to fit the re-associated helmets perspective transform, this time not taking flow into account but minimizing the second order changes between transformation parameters.</p>\n<p>Example of the fixed frame from the game (after stage 2), the original player / helment association and the fixed one (points and points_pred are the flow predicted points to match):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F6d6ebdb11624424f82e4c401b11e2490%2Ffixed_predictions.png?generation=1677990703517978&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2169398",
      "postDate": "03/05/2023 04:34:13",
      "content": "<p>I spend a significant part of time working on this competition to try to improve the helmets/players associations.<br>\nI found this was an even more interesting task, I regret I have not participated in the previous competition and given the close score of the top solution, I hoped it had the potential to boost the score. It improved the score on some folds by around 0.005 and had no impact on other folds (I guess depending on the number of the fixed helmet predictions). Fixing helmets did not improve the private leaderboard score, so I'll keep the approach in a separate post.</p>\n<p>The initial idea was to fit the perspective or affine transformation between the tracking and video positions using something like RANSAC regressor which fits only subset of points. It worked for individual wrong player assignments but failed for frames with the small number of players and/or many mistakes. It would create completely unrealistic transformations.</p>\n<p>One way I tried to improve it is to enforce some kind of continuity between step transformations by adding frame corners and maintaining the constant speed for the frame corners (transformation between corners are added to RANSAC regression). After every transformation fit with RANSAC I removed associations where the predicted position (as estrimated_transform * tracking_pos) does not match the helmet pos on the video.<br>\nAfter I did a recursive search to find a better association between unassociated yet players and helmets (minimising the total square distance between predicted and video pos).</p>\n<p>It helped a bit but was not sufficient for frames with many mistakes (usually the second part of the game with a few player running and many wrongly detected helmets outside of the field).</p>\n<p>Next approach was to fit the camera transformation parameters instead of the unrestricted perspective or affine transformation.<br>\nI'd optimize for the camera position, point in the field camera points to and the current focal length.<br>\nI optimised for MAE of predicted and measured helmet position with an extra cost for the second order of the camera parameter changes between frames to enforce the smooth movements (using Pytorh, Adam optimiser over the whole game).<br>\nI did the initial camera parameters, a few rounds of optimisations, re-assigned helmets as above for the obviously wrong assignment and continued optimisation.</p>\n<p>This helped a bit but a few frames of completely wrong baseline associations pushed camera to the wrong direction for a few video files.<br>\nThe issue was for such games the predicted transformation movement did not match the visible optical flow movement, so I decided to try predicting the optical flow between frames.</p>\n<p>The model receives 2 frames as inputs and predict the ground movement optical field not taking into account players movement.<br>\nThe model is based on 2 resnet34 with TSM like links between.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F423d85e9ccedc596b92364d7af76be93%2Fflow_model.png?generation=1677990657178509&amp;alt=media\" alt=\"\"></p>\n<p>It's trained on either random transformation between the selected frame or two frames with transformation estimated by the previous camera parameters based model.<br>\nIt's important to train not only on syntetic pairs of images but also on actual different frames so model would learn to ignore players movement and align only the field markings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Ffb128e72af041fa40c06a9fbdbd6cd18%2Fflow_prediction_alignment.png?generation=1677990686290014&amp;alt=media\" alt=\"\"></p>\n<p>With the information about the optical flow between frames, the players re-association solution worked much more reliably and after visual inspection I found it fixed almost all player association errors.</p>\n<p>The solution was:</p>\n<p>1) For the initial frame with the good set of predicted players, build the initial transformation using RANSAC. Re-assign outliers using approach described above.<br>\n2) For the next frame, apply RANSAC not only on the tracking pos -&gt; video position pairs but also a seet of points from the last frame with the flow applied and projected to the ground (flow corrected view points should point to the same tracking space pos). This allows to restrict movement between frames to be close to predicted flow but to also correct the accumulated flow errors to better match helments. Re-assign outliers using approach described above.<br>\n3) Run the full game optimisation to fit the re-associated helmets perspective transform, this time not taking flow into account but minimizing the second order changes between transformation parameters.</p>\n<p>Example of the fixed frame from the game (after stage 2), the original player / helment association and the fixed one (points and points_pred are the flow predicted points to match):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F6d6ebdb11624424f82e4c401b11e2490%2Ffixed_predictions.png?generation=1677990703517978&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I spend a significant part of time working on this competition to try to improve the helmets/players associations.\nI found this was an even more interesting task, I regret I have not participated in the previous competition and given the close score of the top solution, I hoped it had the potential to boost the score. It improved the score on some folds by around 0.005 and had no impact on other folds (I guess depending on the number of the fixed helmet predictions). Fixing helmets did not improve the private leaderboard score, so I'll keep the approach in a separate post.\n\nThe initial idea was to fit the perspective or affine transformation between the tracking and video positions using something like RANSAC regressor which fits only subset of points. It worked for individual wrong player assignments but failed for frames with the small number of players and/or many mistakes. It would create completely unrealistic transformations.\n\nOne way I tried to improve it is to enforce some kind of continuity between step transformations by adding frame corners and maintaining the constant speed for the frame corners (transformation between corners are added to RANSAC regression). After every transformation fit with RANSAC I removed associations where the predicted position (as estrimated_transform * tracking_pos) does not match the helmet pos on the video.\nAfter I did a recursive search to find a better association between unassociated yet players and helmets (minimising the total square distance between predicted and video pos).\n \nIt helped a bit but was not sufficient for frames with many mistakes (usually the second part of the game with a few player running and many wrongly detected helmets outside of the field).\n\nNext approach was to fit the camera transformation parameters instead of the unrestricted perspective or affine transformation.\nI'd optimize for the camera position, point in the field camera points to and the current focal length.\nI optimised for MAE of predicted and measured helmet position with an extra cost for the second order of the camera parameter changes between frames to enforce the smooth movements (using Pytorh, Adam optimiser over the whole game).\nI did the initial camera parameters, a few rounds of optimisations, re-assigned helmets as above for the obviously wrong assignment and continued optimisation.\n\nThis helped a bit but a few frames of completely wrong baseline associations pushed camera to the wrong direction for a few video files.\nThe issue was for such games the predicted transformation movement did not match the visible optical flow movement, so I decided to try predicting the optical flow between frames.\n\nThe model receives 2 frames as inputs and predict the ground movement optical field not taking into account players movement.\nThe model is based on 2 resnet34 with TSM like links between.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F423d85e9ccedc596b92364d7af76be93%2Fflow_model.png?generation=1677990657178509&alt=media)\n\nIt's trained on either random transformation between the selected frame or two frames with transformation estimated by the previous camera parameters based model.\nIt's important to train not only on syntetic pairs of images but also on actual different frames so model would learn to ignore players movement and align only the field markings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Ffb128e72af041fa40c06a9fbdbd6cd18%2Fflow_prediction_alignment.png?generation=1677990686290014&alt=media)\n\nWith the information about the optical flow between frames, the players re-association solution worked much more reliably and after visual inspection I found it fixed almost all player association errors.\n\nThe solution was:\n\n1) For the initial frame with the good set of predicted players, build the initial transformation using RANSAC. Re-assign outliers using approach described above.\n2) For the next frame, apply RANSAC not only on the tracking pos -> video position pairs but also a seet of points from the last frame with the flow applied and projected to the ground (flow corrected view points should point to the same tracking space pos). This allows to restrict movement between frames to be close to predicted flow but to also correct the accumulated flow errors to better match helments. Re-assign outliers using approach described above.\n3) Run the full game optimisation to fit the re-associated helmets perspective transform, this time not taking flow into account but minimizing the second order changes between transformation parameters.\n\nExample of the fixed frame from the game (after stage 2), the original player / helment association and the fixed one (points and points_pred are the flow predicted points to match):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F6d6ebdb11624424f82e4c401b11e2490%2Ffixed_predictions.png?generation=1677990703517978&alt=media)",
      "votes": null
    },
    {
      "id": "2169428",
      "postDate": "03/05/2023 05:20:50",
      "content": "<p>very helpful</p>",
      "rawMarkdown": "very helpful",
      "votes": null
    },
    {
      "id": "2170153",
      "postDate": "03/05/2023 18:11:50",
      "content": "<p>Which models were trained?</p>",
      "rawMarkdown": "Which models were trained?",
      "votes": null
    },
    {
      "id": "2170256",
      "postDate": "03/05/2023 19:33:24",
      "content": "<p>To assign helmets - only the resnet34+TSM based flow estimation model.</p>",
      "rawMarkdown": "To assign helmets - only the resnet34+TSM based flow estimation model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2169428,
      "author_name": "altaibaatarenkhbat",
      "author_url": "",
      "post_date": "03/05/2023 05:20:50",
      "content": "<p>very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2170153,
      "author_name": "marufjawad",
      "author_url": "",
      "post_date": "03/05/2023 18:11:50",
      "content": "<p>Which models were trained?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2170256,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "03/05/2023 19:33:24",
          "content": "<p>To assign helmets - only the resnet34+TSM based flow estimation model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2169398": "I spend a significant part of time working on this competition to try to improve the helmets/players associations.\nI found this was an even more interesting task, I regret I have not participated in the previous competition and given the close score of the top solution, I hoped it had the potential to boost the score. It improved the score on some folds by around 0.005 and had no impact on other folds (I guess depending on the number of the fixed helmet predictions). Fixing helmets did not improve the private leaderboard score, so I'll keep the approach in a separate post.\n\nThe initial idea was to fit the perspective or affine transformation between the tracking and video positions using something like RANSAC regressor which fits only subset of points. It worked for individual wrong player assignments but failed for frames with the small number of players and/or many mistakes. It would create completely unrealistic transformations.\n\nOne way I tried to improve it is to enforce some kind of continuity between step transformations by adding frame corners and maintaining the constant speed for the frame corners (transformation between corners are added to RANSAC regression). After every transformation fit with RANSAC I removed associations where the predicted position (as estrimated_transform * tracking_pos) does not match the helmet pos on the video.\nAfter I did a recursive search to find a better association between unassociated yet players and helmets (minimising the total square distance between predicted and video pos).\n \nIt helped a bit but was not sufficient for frames with many mistakes (usually the second part of the game with a few player running and many wrongly detected helmets outside of the field).\n\nNext approach was to fit the camera transformation parameters instead of the unrestricted perspective or affine transformation.\nI'd optimize for the camera position, point in the field camera points to and the current focal length.\nI optimised for MAE of predicted and measured helmet position with an extra cost for the second order of the camera parameter changes between frames to enforce the smooth movements (using Pytorh, Adam optimiser over the whole game).\nI did the initial camera parameters, a few rounds of optimisations, re-assigned helmets as above for the obviously wrong assignment and continued optimisation.\n\nThis helped a bit but a few frames of completely wrong baseline associations pushed camera to the wrong direction for a few video files.\nThe issue was for such games the predicted transformation movement did not match the visible optical flow movement, so I decided to try predicting the optical flow between frames.\n\nThe model receives 2 frames as inputs and predict the ground movement optical field not taking into account players movement.\nThe model is based on 2 resnet34 with TSM like links between.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F423d85e9ccedc596b92364d7af76be93%2Fflow_model.png?generation=1677990657178509&alt=media)\n\nIt's trained on either random transformation between the selected frame or two frames with transformation estimated by the previous camera parameters based model.\nIt's important to train not only on syntetic pairs of images but also on actual different frames so model would learn to ignore players movement and align only the field markings.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Ffb128e72af041fa40c06a9fbdbd6cd18%2Fflow_prediction_alignment.png?generation=1677990686290014&alt=media)\n\nWith the information about the optical flow between frames, the players re-association solution worked much more reliably and after visual inspection I found it fixed almost all player association errors.\n\nThe solution was:\n\n1) For the initial frame with the good set of predicted players, build the initial transformation using RANSAC. Re-assign outliers using approach described above.\n2) For the next frame, apply RANSAC not only on the tracking pos -> video position pairs but also a seet of points from the last frame with the flow applied and projected to the ground (flow corrected view points should point to the same tracking space pos). This allows to restrict movement between frames to be close to predicted flow but to also correct the accumulated flow errors to better match helments. Re-assign outliers using approach described above.\n3) Run the full game optimisation to fit the re-associated helmets perspective transform, this time not taking flow into account but minimizing the second order changes between transformation parameters.\n\nExample of the fixed frame from the game (after stage 2), the original player / helment association and the fixed one (points and points_pred are the flow predicted points to match):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2F6d6ebdb11624424f82e4c401b11e2490%2Ffixed_predictions.png?generation=1677990703517978&alt=media)",
    "2169428": "very helpful",
    "2170153": "Which models were trained?",
    "2170256": "To assign helmets - only the resnet34+TSM based flow estimation model."
  },
  "source": "meta"
}