{
  "id": 302411,
  "title": "Do we need sliding window ???",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/302411",
  "author_name": "",
  "post_date": "2022-01-22T12:22:45.214691400Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I tried to overcome the problem of small objects in the training dataset and the problem of large GPU needs, so I divided each image to 36 images with 214*120 pixels  and correct labels of these images.<br>\nthen trained my YOLOV5l on this new data with image size 640 for 100 epochs.</p>\n<p>the validation results were great, I got 0.92 F2 score.</p>\n<p>then I used  sliding window algorithm that divide testing images into small images and run the  model inference through these small images individually and then collect these predictions together.</p>\n<p>but I got very bad results nearly 0.185</p>\n<p>I wondering if I made I mistake somewhere?? </p>",
  "messages": [
    {
      "id": "1660113",
      "postDate": "01/22/2022 12:22:45",
      "content": "<p>I tried to overcome the problem of small objects in the training dataset and the problem of large GPU needs, so I divided each image to 36 images with 214*120 pixels  and correct labels of these images.<br>\nthen trained my YOLOV5l on this new data with image size 640 for 100 epochs.</p>\n<p>the validation results were great, I got 0.92 F2 score.</p>\n<p>then I used  sliding window algorithm that divide testing images into small images and run the  model inference through these small images individually and then collect these predictions together.</p>\n<p>but I got very bad results nearly 0.185</p>\n<p>I wondering if I made I mistake somewhere?? </p>",
      "rawMarkdown": "I tried to overcome the problem of small objects in the training dataset and the problem of large GPU needs, so I divided each image to 36 images with 214*120 pixels  and correct labels of these images.\nthen trained my YOLOV5l on this new data with image size 640 for 100 epochs.\n\nthe validation results were great, I got 0.92 F2 score.\n\nthen I used  sliding window algorithm that divide testing images into small images and run the  model inference through these small images individually and then collect these predictions together.\n\nbut I got very bad results nearly 0.185\n\nI wondering if I made I mistake somewhere??",
      "votes": null
    },
    {
      "id": "1660502",
      "postDate": "01/22/2022 17:52:35",
      "content": "<p>It depends, I think 214x120 is too small and I have tried other levels with <a href=\"https://github.com/obss/sahi\" target=\"_blank\">sahi</a>. I recommend maybe doing around 768. I got some acceptable results using a modified detector class:</p>\n<pre><code>class Yolov5DetectionModelHiRez(Yolov5DetectionModel):\n    def perform_inference(self, image: np.ndarray, image_size: int = None):\n        \"\"\"\n        Prediction is performed using self.model and the prediction result is set to self._original_predictions.\n        Args:\n            image: np.ndarray\n                A numpy array that contains the image to be predicted. 3 channel image should be in RGB order.\n            image_size: int\n                Inference input size.\n        \"\"\"\n        try:\n            import yolov5\n        except ImportError:\n            raise ImportError('Please run \"pip install -U yolov5\" ' \"to install YOLOv5 first for YOLOv5 inference.\")\n\n        # Confirm model is loaded\n        assert self.model is not None, \"Model is not loaded, load it by calling .load_model()\"\n\n        prediction_result = self.model(image, size=1280, augment=False) # Inference at a bit higher rez than original for each slize\n\n        self._original_predictions = prediction_result\n</code></pre>\n<p>But results are still not great, it may be worth looking at doing a sahi with high overlap instead of augmentation. Even my results from splitting images smaller than 768 did not do better than high-resolution training. Hopefully, my tests help someone make something better.</p>",
      "rawMarkdown": "It depends, I think 214x120 is too small and I have tried other levels with [sahi](https://github.com/obss/sahi). I recommend maybe doing around 768. I got some acceptable results using a modified detector class:\n\n```\nclass Yolov5DetectionModelHiRez(Yolov5DetectionModel):\n    def perform_inference(self, image: np.ndarray, image_size: int = None):\n        \"\"\"\n        Prediction is performed using self.model and the prediction result is set to self._original_predictions.\n        Args:\n            image: np.ndarray\n                A numpy array that contains the image to be predicted. 3 channel image should be in RGB order.\n            image_size: int\n                Inference input size.\n        \"\"\"\n        try:\n            import yolov5\n        except ImportError:\n            raise ImportError('Please run \"pip install -U yolov5\" ' \"to install YOLOv5 first for YOLOv5 inference.\")\n\n        # Confirm model is loaded\n        assert self.model is not None, \"Model is not loaded, load it by calling .load_model()\"\n        \n        prediction_result = self.model(image, size=1280, augment=False) # Inference at a bit higher rez than original for each slize\n\n        self._original_predictions = prediction_result\n```\n\nBut results are still not great, it may be worth looking at doing a sahi with high overlap instead of augmentation. Even my results from splitting images smaller than 768 did not do better than high-resolution training. Hopefully, my tests help someone make something better.",
      "votes": null
    },
    {
      "id": "1660507",
      "postDate": "01/22/2022 17:54:02",
      "content": "<p>I will publish a notebook for it soon if I get higher results than the baseline yolov5, it depends if figure it out before the cutoff for posting notebooks since this competition will be ending in a few weeks.</p>",
      "rawMarkdown": "I will publish a notebook for it soon if I get higher results than the baseline yolov5, it depends if figure it out before the cutoff for posting notebooks since this competition will be ending in a few weeks.",
      "votes": null
    },
    {
      "id": "1660675",
      "postDate": "01/22/2022 21:05:26",
      "content": "<p><a href=\"https://www.kaggle.com/outwrest\" target=\"_blank\">@outwrest</a>  but why I got 0.92 F2 score while getting these bad results on LB</p>",
      "rawMarkdown": "outwrest  but why I got 0.92 F2 score while getting these bad results on LB",
      "votes": null
    },
    {
      "id": "1660687",
      "postDate": "01/22/2022 21:23:41",
      "content": "<p>I am not sure, it could be a lot of things that can result in you getting a very good F2 score on validation data. Make sure your data is split correctly and that you have a lot of empty images in validation to get an accurate result.</p>",
      "rawMarkdown": "I am not sure, it could be a lot of things that can result in you getting a very good F2 score on validation data. Make sure your data is split correctly and that you have a lot of empty images in validation to get an accurate result.",
      "votes": null
    },
    {
      "id": "1663538",
      "postDate": "01/25/2022 07:38:02",
      "content": "<p>0.92F2? Maybe your validation set is not split correctly.</p>",
      "rawMarkdown": "0.92F2? Maybe your validation set is not split correctly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1660502,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "01/22/2022 17:52:35",
      "content": "<p>It depends, I think 214x120 is too small and I have tried other levels with <a href=\"https://github.com/obss/sahi\" target=\"_blank\">sahi</a>. I recommend maybe doing around 768. I got some acceptable results using a modified detector class:</p>\n<pre><code>class Yolov5DetectionModelHiRez(Yolov5DetectionModel):\n    def perform_inference(self, image: np.ndarray, image_size: int = None):\n        \"\"\"\n        Prediction is performed using self.model and the prediction result is set to self._original_predictions.\n        Args:\n            image: np.ndarray\n                A numpy array that contains the image to be predicted. 3 channel image should be in RGB order.\n            image_size: int\n                Inference input size.\n        \"\"\"\n        try:\n            import yolov5\n        except ImportError:\n            raise ImportError('Please run \"pip install -U yolov5\" ' \"to install YOLOv5 first for YOLOv5 inference.\")\n\n        # Confirm model is loaded\n        assert self.model is not None, \"Model is not loaded, load it by calling .load_model()\"\n\n        prediction_result = self.model(image, size=1280, augment=False) # Inference at a bit higher rez than original for each slize\n\n        self._original_predictions = prediction_result\n</code></pre>\n<p>But results are still not great, it may be worth looking at doing a sahi with high overlap instead of augmentation. Even my results from splitting images smaller than 768 did not do better than high-resolution training. Hopefully, my tests help someone make something better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1660507,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "01/22/2022 17:54:02",
          "content": "<p>I will publish a notebook for it soon if I get higher results than the baseline yolov5, it depends if figure it out before the cutoff for posting notebooks since this competition will be ending in a few weeks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1660675,
      "author_name": "mohameddonia2222",
      "author_url": "",
      "post_date": "01/22/2022 21:05:26",
      "content": "<p><a href=\"https://www.kaggle.com/outwrest\" target=\"_blank\">@outwrest</a>  but why I got 0.92 F2 score while getting these bad results on LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 1660687,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "01/22/2022 21:23:41",
          "content": "<p>I am not sure, it could be a lot of things that can result in you getting a very good F2 score on validation data. Make sure your data is split correctly and that you have a lot of empty images in validation to get an accurate result.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1663538,
      "author_name": "klawensliu",
      "author_url": "",
      "post_date": "01/25/2022 07:38:02",
      "content": "<p>0.92F2? Maybe your validation set is not split correctly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1660113": "I tried to overcome the problem of small objects in the training dataset and the problem of large GPU needs, so I divided each image to 36 images with 214*120 pixels  and correct labels of these images.\nthen trained my YOLOV5l on this new data with image size 640 for 100 epochs.\n\nthe validation results were great, I got 0.92 F2 score.\n\nthen I used  sliding window algorithm that divide testing images into small images and run the  model inference through these small images individually and then collect these predictions together.\n\nbut I got very bad results nearly 0.185\n\nI wondering if I made I mistake somewhere??",
    "1660502": "It depends, I think 214x120 is too small and I have tried other levels with [sahi](https://github.com/obss/sahi). I recommend maybe doing around 768. I got some acceptable results using a modified detector class:\n\n```\nclass Yolov5DetectionModelHiRez(Yolov5DetectionModel):\n    def perform_inference(self, image: np.ndarray, image_size: int = None):\n        \"\"\"\n        Prediction is performed using self.model and the prediction result is set to self._original_predictions.\n        Args:\n            image: np.ndarray\n                A numpy array that contains the image to be predicted. 3 channel image should be in RGB order.\n            image_size: int\n                Inference input size.\n        \"\"\"\n        try:\n            import yolov5\n        except ImportError:\n            raise ImportError('Please run \"pip install -U yolov5\" ' \"to install YOLOv5 first for YOLOv5 inference.\")\n\n        # Confirm model is loaded\n        assert self.model is not None, \"Model is not loaded, load it by calling .load_model()\"\n        \n        prediction_result = self.model(image, size=1280, augment=False) # Inference at a bit higher rez than original for each slize\n\n        self._original_predictions = prediction_result\n```\n\nBut results are still not great, it may be worth looking at doing a sahi with high overlap instead of augmentation. Even my results from splitting images smaller than 768 did not do better than high-resolution training. Hopefully, my tests help someone make something better.",
    "1660507": "I will publish a notebook for it soon if I get higher results than the baseline yolov5, it depends if figure it out before the cutoff for posting notebooks since this competition will be ending in a few weeks.",
    "1660675": "outwrest  but why I got 0.92 F2 score while getting these bad results on LB",
    "1660687": "I am not sure, it could be a lot of things that can result in you getting a very good F2 score on validation data. Make sure your data is split correctly and that you have a lot of empty images in validation to get an accurate result.",
    "1663538": "0.92F2? Maybe your validation set is not split correctly."
  },
  "source": "meta"
}