{
  "id": 466592,
  "title": "How to focus the model on selected areas of the image?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/466592",
  "author_name": "",
  "post_date": "2024-01-09T08:34:54.272388800Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Reading through the solution description of the winner: <a href=\"https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/discussion/22906\" target=\"_blank\">https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/discussion/22906</a></p>\n<p>I read the following:<br>\n<em>To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver’s right hand cross is highly correlated with c5. The original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3.</em></p>\n<p>How does it actually work this selection and multiple inputs to the model? Did the winner processed every image to give the model an input like the following one?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16289333%2F42b42f703e6b5752e8205316a623c69e%2Fch.jpg?generation=1704789247757506&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2593476",
      "postDate": "01/09/2024 08:34:54",
      "content": "<p>Reading through the solution description of the winner: <a href=\"https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/discussion/22906\" target=\"_blank\">https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/discussion/22906</a></p>\n<p>I read the following:<br>\n<em>To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver’s right hand cross is highly correlated with c5. The original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3.</em></p>\n<p>How does it actually work this selection and multiple inputs to the model? Did the winner processed every image to give the model an input like the following one?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16289333%2F42b42f703e6b5752e8205316a623c69e%2Fch.jpg?generation=1704789247757506&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Reading through the solution description of the winner: https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/discussion/22906\n\nI read the following:\n*To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver’s right hand cross is highly correlated with c5. The original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3.*\n\nHow does it actually work this selection and multiple inputs to the model? Did the winner processed every image to give the model an input like the following one?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16289333%2F42b42f703e6b5752e8205316a623c69e%2Fch.jpg?generation=1704789247757506&alt=media)",
      "votes": null
    },
    {
      "id": "2974245",
      "postDate": "08/30/2024 13:24:54",
      "content": "<p>This is probably a little bit late, but you could use DeepFace and mediapipe to identify faces and hands and the crop the subsets of the image to train. Currently working on that…</p>",
      "rawMarkdown": "This is probably a little bit late, but you could use DeepFace and mediapipe to identify faces and hands and the crop the subsets of the image to train. Currently working on that...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2974245,
      "author_name": "simonschwerd",
      "author_url": "",
      "post_date": "08/30/2024 13:24:54",
      "content": "<p>This is probably a little bit late, but you could use DeepFace and mediapipe to identify faces and hands and the crop the subsets of the image to train. Currently working on that…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2593476": "Reading through the solution description of the winner: https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/discussion/22906\n\nI read the following:\n*To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver’s right hand cross is highly correlated with c5. The original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3.*\n\nHow does it actually work this selection and multiple inputs to the model? Did the winner processed every image to give the model an input like the following one?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16289333%2F42b42f703e6b5752e8205316a623c69e%2Fch.jpg?generation=1704789247757506&alt=media)",
    "2974245": "This is probably a little bit late, but you could use DeepFace and mediapipe to identify faces and hands and the crop the subsets of the image to train. Currently working on that..."
  },
  "source": "meta"
}