{
  "id": 576536,
  "title": "Fine-tuned MViT-V2-S from Torchvision on 24GB GPU （(0.7 Kaggle Score)",
  "url": "/competitions/nexar-collision-prediction/discussion/576536",
  "author_name": "",
  "post_date": "2025-05-05T15:35:36.152815200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<h2>🔑 Key Highlights</h2>\n<p>Pre-trained Transformer Backbone:<br>\nUtilizes MViT-V2-S, originally trained on Kinetics-400, adapted for binary classification by replacing the final head. pretrain in MViT-V2-S</p>\n<h4>Efficient Video Loading &amp; Preprocessing:</h4>\n<ul>\n<li><p>Reads videos frame-by-frame with OpenCV.</p></li>\n<li><p>Uniformly or randomly samples 16 frames per video to ensure input consistency.</p></li>\n</ul>\n<p>Applies basic preprocessing:</p>\n<ul>\n<li><p>Resize to 224x224,</p></li>\n<li><p>Convert to Tensor,</p></li>\n<li><p>Normalize.</p></li>\n</ul>\n<h4>Optimized Training:</h4>\n<ul>\n<li><p>Uses AdamW optimizer and CosineAnnealingLR.</p></li>\n<li><p>Employs mixed precision (AMP) to accelerate training.</p></li>\n<li><p>Monitors validation mAP (Mean Average Precision) for early stopping and saving the best model.</p></li>\n<li><p>Robust Inference with TTA (Test Time Augmentation):</p></li>\n<li><p>Performs multiple inferences with random frame sampling.</p></li>\n<li><p>Averages results for stable final predictions.</p></li>\n</ul>\n<h4>Submission Generation:</h4>\n<p>Outputs prediction results as CSV, mapping video IDs to target probabilities for easy submission.</p>",
  "messages": [
    {
      "id": "3194265",
      "postDate": "05/05/2025 15:35:36",
      "content": "<h2>🔑 Key Highlights</h2>\n<p>Pre-trained Transformer Backbone:<br>\nUtilizes MViT-V2-S, originally trained on Kinetics-400, adapted for binary classification by replacing the final head. pretrain in MViT-V2-S</p>\n<h4>Efficient Video Loading &amp; Preprocessing:</h4>\n<ul>\n<li><p>Reads videos frame-by-frame with OpenCV.</p></li>\n<li><p>Uniformly or randomly samples 16 frames per video to ensure input consistency.</p></li>\n</ul>\n<p>Applies basic preprocessing:</p>\n<ul>\n<li><p>Resize to 224x224,</p></li>\n<li><p>Convert to Tensor,</p></li>\n<li><p>Normalize.</p></li>\n</ul>\n<h4>Optimized Training:</h4>\n<ul>\n<li><p>Uses AdamW optimizer and CosineAnnealingLR.</p></li>\n<li><p>Employs mixed precision (AMP) to accelerate training.</p></li>\n<li><p>Monitors validation mAP (Mean Average Precision) for early stopping and saving the best model.</p></li>\n<li><p>Robust Inference with TTA (Test Time Augmentation):</p></li>\n<li><p>Performs multiple inferences with random frame sampling.</p></li>\n<li><p>Averages results for stable final predictions.</p></li>\n</ul>\n<h4>Submission Generation:</h4>\n<p>Outputs prediction results as CSV, mapping video IDs to target probabilities for easy submission.</p>",
      "rawMarkdown": "## 🔑 Key Highlights\nPre-trained Transformer Backbone:\nUtilizes MViT-V2-S, originally trained on Kinetics-400, adapted for binary classification by replacing the final head. pretrain in MViT-V2-S\n\n#### Efficient Video Loading & Preprocessing:\n\n- Reads videos frame-by-frame with OpenCV.\n\n- Uniformly or randomly samples 16 frames per video to ensure input consistency.\n\nApplies basic preprocessing:\n\n- Resize to 224x224,\n\n- Convert to Tensor,\n\n- Normalize.\n\n\n#### Optimized Training:\n- Uses AdamW optimizer and CosineAnnealingLR.\n\n- Employs mixed precision (AMP) to accelerate training.\n\n- Monitors validation mAP (Mean Average Precision) for early stopping and saving the best model.\n\n- Robust Inference with TTA (Test Time Augmentation):\n\n- Performs multiple inferences with random frame sampling.\n\n- Averages results for stable final predictions.\n\n#### Submission Generation:\n\nOutputs prediction results as CSV, mapping video IDs to target probabilities for easy submission.",
      "votes": null
    },
    {
      "id": "3194279",
      "postDate": "05/05/2025 16:07:46",
      "content": "<p>Interesting, I also used mvit2_v2_s, with almost same training parameters (and 224x224), started at 0.71 and ended up at 0.898 by deleting and weighting samples. Will post everything on github soon.</p>",
      "rawMarkdown": "Interesting, I also used mvit2_v2_s, with almost same training parameters (and 224x224), started at 0.71 and ended up at 0.898 by deleting and weighting samples. Will post everything on github soon.",
      "votes": null
    },
    {
      "id": "3194467",
      "postDate": "05/05/2025 23:33:10",
      "content": "<p>Thank you， I want to know how to deleting and weighting samples. I’m really looking forward to your sharing solution.</p>",
      "rawMarkdown": "Thank you， I want to know how to deleting and weighting samples. I’m really looking forward to your sharing solution.",
      "votes": null
    },
    {
      "id": "3198799",
      "postDate": "05/10/2025 02:16:10",
      "content": "<p><a href=\"https://www.kaggle.com/peacelu\" target=\"_blank\">@peacelu</a> see my post!</p>",
      "rawMarkdown": "peacelu see my post!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3194279,
      "author_name": "paulendresen76",
      "author_url": "",
      "post_date": "05/05/2025 16:07:46",
      "content": "<p>Interesting, I also used mvit2_v2_s, with almost same training parameters (and 224x224), started at 0.71 and ended up at 0.898 by deleting and weighting samples. Will post everything on github soon.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3194467,
          "author_name": "peacelu",
          "author_url": "",
          "post_date": "05/05/2025 23:33:10",
          "content": "<p>Thank you， I want to know how to deleting and weighting samples. I’m really looking forward to your sharing solution.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3198799,
              "author_name": "paulendresen76",
              "author_url": "",
              "post_date": "05/10/2025 02:16:10",
              "content": "<p><a href=\"https://www.kaggle.com/peacelu\" target=\"_blank\">@peacelu</a> see my post!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3194265": "## 🔑 Key Highlights\nPre-trained Transformer Backbone:\nUtilizes MViT-V2-S, originally trained on Kinetics-400, adapted for binary classification by replacing the final head. pretrain in MViT-V2-S\n\n#### Efficient Video Loading & Preprocessing:\n\n- Reads videos frame-by-frame with OpenCV.\n\n- Uniformly or randomly samples 16 frames per video to ensure input consistency.\n\nApplies basic preprocessing:\n\n- Resize to 224x224,\n\n- Convert to Tensor,\n\n- Normalize.\n\n\n#### Optimized Training:\n- Uses AdamW optimizer and CosineAnnealingLR.\n\n- Employs mixed precision (AMP) to accelerate training.\n\n- Monitors validation mAP (Mean Average Precision) for early stopping and saving the best model.\n\n- Robust Inference with TTA (Test Time Augmentation):\n\n- Performs multiple inferences with random frame sampling.\n\n- Averages results for stable final predictions.\n\n#### Submission Generation:\n\nOutputs prediction results as CSV, mapping video IDs to target probabilities for easy submission.",
    "3194279": "Interesting, I also used mvit2_v2_s, with almost same training parameters (and 224x224), started at 0.71 and ended up at 0.898 by deleting and weighting samples. Will post everything on github soon.",
    "3194467": "Thank you， I want to know how to deleting and weighting samples. I’m really looking forward to your sharing solution.",
    "3198799": "peacelu see my post!"
  },
  "source": "meta"
}