{
  "id": 305801,
  "title": "Imperfect annotation - the primary reason why some ideas (augmentations/big model/higher resolution) does not work?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/305801",
  "author_name": "",
  "post_date": "2022-02-07T01:03:05.871618900Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>When I run my inference on small resolution (x1.5) I got a higher LB score compared to a bigger one (x2.5) on the exact same weights, and this happened for most of my trained model.</p>\n<p>During dataset investigation, I see that there are some positive cases that are not annotated in the ground-truth, so the detection produced by the model is considered a FP case. </p>\n<p>Bigger models and resolutions tend to have lower precision in favor of improving recall, but if annotations are missing TP cases, those models have no improvement in recall while having lower precision. Maybe that's the cause of inconsistent results between CV/LB and the unexpected lower scores of bigger/more complex models? </p>\n<p>Here is a sample:<br>\nPrediction: <a href=\"https://drive.google.com/file/d/1wzN4OEx5Sn6ul-fIn5emrtApHFH6jKA_/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1wzN4OEx5Sn6ul-fIn5emrtApHFH6jKA_/view?usp=sharing</a><br>\nAnnotation: <a href=\"https://drive.google.com/file/d/1k3u5qIrITBvQgywDRNZwNdZrQmVfvbK5/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1k3u5qIrITBvQgywDRNZwNdZrQmVfvbK5/view?usp=sharing</a></p>",
  "messages": [
    {
      "id": "1679086",
      "postDate": "02/07/2022 01:03:05",
      "content": "<p>When I run my inference on small resolution (x1.5) I got a higher LB score compared to a bigger one (x2.5) on the exact same weights, and this happened for most of my trained model.</p>\n<p>During dataset investigation, I see that there are some positive cases that are not annotated in the ground-truth, so the detection produced by the model is considered a FP case. </p>\n<p>Bigger models and resolutions tend to have lower precision in favor of improving recall, but if annotations are missing TP cases, those models have no improvement in recall while having lower precision. Maybe that's the cause of inconsistent results between CV/LB and the unexpected lower scores of bigger/more complex models? </p>\n<p>Here is a sample:<br>\nPrediction: <a href=\"https://drive.google.com/file/d/1wzN4OEx5Sn6ul-fIn5emrtApHFH6jKA_/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1wzN4OEx5Sn6ul-fIn5emrtApHFH6jKA_/view?usp=sharing</a><br>\nAnnotation: <a href=\"https://drive.google.com/file/d/1k3u5qIrITBvQgywDRNZwNdZrQmVfvbK5/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1k3u5qIrITBvQgywDRNZwNdZrQmVfvbK5/view?usp=sharing</a></p>",
      "rawMarkdown": "When I run my inference on small resolution (x1.5) I got a higher LB score compared to a bigger one (x2.5) on the exact same weights, and this happened for most of my trained model.\n\nDuring dataset investigation, I see that there are some positive cases that are not annotated in the ground-truth, so the detection produced by the model is considered a FP case. \n\nBigger models and resolutions tend to have lower precision in favor of improving recall, but if annotations are missing TP cases, those models have no improvement in recall while having lower precision. Maybe that's the cause of inconsistent results between CV/LB and the unexpected lower scores of bigger/more complex models? \n\nHere is a sample:\nPrediction: https://drive.google.com/file/d/1wzN4OEx5Sn6ul-fIn5emrtApHFH6jKA_/view?usp=sharing\nAnnotation: https://drive.google.com/file/d/1k3u5qIrITBvQgywDRNZwNdZrQmVfvbK5/view?usp=sharing",
      "votes": null
    },
    {
      "id": "1679225",
      "postDate": "02/07/2022 04:53:49",
      "content": "<p>Welcome in competition…</p>\n<p>This topic was touched many times here …</p>",
      "rawMarkdown": "Welcome in competition…\n\nThis topic was touched many times here …",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1679225,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/07/2022 04:53:49",
      "content": "<p>Welcome in competition…</p>\n<p>This topic was touched many times here …</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1679086": "When I run my inference on small resolution (x1.5) I got a higher LB score compared to a bigger one (x2.5) on the exact same weights, and this happened for most of my trained model.\n\nDuring dataset investigation, I see that there are some positive cases that are not annotated in the ground-truth, so the detection produced by the model is considered a FP case. \n\nBigger models and resolutions tend to have lower precision in favor of improving recall, but if annotations are missing TP cases, those models have no improvement in recall while having lower precision. Maybe that's the cause of inconsistent results between CV/LB and the unexpected lower scores of bigger/more complex models? \n\nHere is a sample:\nPrediction: https://drive.google.com/file/d/1wzN4OEx5Sn6ul-fIn5emrtApHFH6jKA_/view?usp=sharing\nAnnotation: https://drive.google.com/file/d/1k3u5qIrITBvQgywDRNZwNdZrQmVfvbK5/view?usp=sharing",
    "1679225": "Welcome in competition…\n\nThis topic was touched many times here …"
  },
  "source": "meta"
}