{
  "id": 307173,
  "title": "Visualized models with low scores revealed some labeling errors in the official test set",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307173",
  "author_name": "",
  "post_date": "2022-02-13T03:55:20.841191900Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>Body</h1>\n<p>In the latest training, we adjusted the parameters to get what we thought was the our best model for visualization:</p>\n<p>However, our score was not better than the original one. After visualization, we found that it could actually correct some errors in the Label of the original data set. Some individuals that were obviously starfish but not marked could also be found, that is to say, He was able to find some starfish that hadn't been marked before.</p>\n<p>So I feel that maybe sometimes the performance of the model is higher than the score is because there are a few errors in the official data set.</p>\n<p>No other meaning and thoughts, just a small discovery of training</p>\n<h1>indicators</h1>\n<p><a href=\"https://postimg.cc/0zm3ch71\" target=\"_blank\"><img src=\"https://i.postimg.cc/7ZQkYk6h/8a099a96b37c2df50f0e1b8a8f02fd6.png\" alt=\"8a099a96b37c2df50f0e1b8a8f02fd6.png\"></a></p>\n<h1>visualization:</h1>\n<p><a href=\"https://postimg.cc/MnX8Wskn\" target=\"_blank\"><img src=\"https://i.postimg.cc/Gmx3xNqx/5578878036c2da1c9a6a44634605968.png\" alt=\"5578878036c2da1c9a6a44634605968.png\"></a><br>\n<a href=\"https://postimg.cc/WD68xv9g\" target=\"_blank\"><img src=\"https://i.postimg.cc/fb1PtWm2/8f16e72d460070251281d7622ab6574.png\" alt=\"8f16e72d460070251281d7622ab6574.png\"></a></p>\n<h2><a href=\"https://postimg.cc/HVc2d5Zp\" target=\"_blank\"><img src=\"https://i.postimg.cc/KvQs7Dtg/65448dcf88cd42095c1eadcd79aecc2.png\" alt=\"65448dcf88cd42095c1eadcd79aecc2.png\"></a></h2>\n<p><a href=\"https://postimg.cc/NKRkrtvS\" target=\"_blank\"><img src=\"https://i.postimg.cc/DwMjHfjv/593a7088276fb51bce51fbf1e14ccd7.png\" alt=\"593a7088276fb51bce51fbf1e14ccd7.png\"></a><br>\n<a href=\"https://postimg.cc/rzzGNxLS\" target=\"_blank\"><img src=\"https://i.postimg.cc/mZydRyX6/a40dd00c72dc9592f3adb3bbdfc4ef2.png\" alt=\"a40dd00c72dc9592f3adb3bbdfc4ef2.png\"></a></p>\n<p>During this time, we have encountered some problems<br>\nAt first we tried YoloX, but it didn't work very well;<br>\nThen we tried V5, and the effect of V5 made us reach 0.665;<br>\nWe used Ensemble[V5 &amp; X] and found that it was better without YoloX;<br>\nWe also tried many, for example SAHI | COPY &amp; Paste | YoloV5 &amp; BigSize<br>\nYoloV5 &amp; BigSize;</p>",
  "messages": [
    {
      "id": "1687677",
      "postDate": "02/13/2022 03:55:20",
      "content": "<h1>Body</h1>\n<p>In the latest training, we adjusted the parameters to get what we thought was the our best model for visualization:</p>\n<p>However, our score was not better than the original one. After visualization, we found that it could actually correct some errors in the Label of the original data set. Some individuals that were obviously starfish but not marked could also be found, that is to say, He was able to find some starfish that hadn't been marked before.</p>\n<p>So I feel that maybe sometimes the performance of the model is higher than the score is because there are a few errors in the official data set.</p>\n<p>No other meaning and thoughts, just a small discovery of training</p>\n<h1>indicators</h1>\n<p><a href=\"https://postimg.cc/0zm3ch71\" target=\"_blank\"><img src=\"https://i.postimg.cc/7ZQkYk6h/8a099a96b37c2df50f0e1b8a8f02fd6.png\" alt=\"8a099a96b37c2df50f0e1b8a8f02fd6.png\"></a></p>\n<h1>visualization:</h1>\n<p><a href=\"https://postimg.cc/MnX8Wskn\" target=\"_blank\"><img src=\"https://i.postimg.cc/Gmx3xNqx/5578878036c2da1c9a6a44634605968.png\" alt=\"5578878036c2da1c9a6a44634605968.png\"></a><br>\n<a href=\"https://postimg.cc/WD68xv9g\" target=\"_blank\"><img src=\"https://i.postimg.cc/fb1PtWm2/8f16e72d460070251281d7622ab6574.png\" alt=\"8f16e72d460070251281d7622ab6574.png\"></a></p>\n<h2><a href=\"https://postimg.cc/HVc2d5Zp\" target=\"_blank\"><img src=\"https://i.postimg.cc/KvQs7Dtg/65448dcf88cd42095c1eadcd79aecc2.png\" alt=\"65448dcf88cd42095c1eadcd79aecc2.png\"></a></h2>\n<p><a href=\"https://postimg.cc/NKRkrtvS\" target=\"_blank\"><img src=\"https://i.postimg.cc/DwMjHfjv/593a7088276fb51bce51fbf1e14ccd7.png\" alt=\"593a7088276fb51bce51fbf1e14ccd7.png\"></a><br>\n<a href=\"https://postimg.cc/rzzGNxLS\" target=\"_blank\"><img src=\"https://i.postimg.cc/mZydRyX6/a40dd00c72dc9592f3adb3bbdfc4ef2.png\" alt=\"a40dd00c72dc9592f3adb3bbdfc4ef2.png\"></a></p>\n<p>During this time, we have encountered some problems<br>\nAt first we tried YoloX, but it didn't work very well;<br>\nThen we tried V5, and the effect of V5 made us reach 0.665;<br>\nWe used Ensemble[V5 &amp; X] and found that it was better without YoloX;<br>\nWe also tried many, for example SAHI | COPY &amp; Paste | YoloV5 &amp; BigSize<br>\nYoloV5 &amp; BigSize;</p>",
      "rawMarkdown": "# Body\nIn the latest training, we adjusted the parameters to get what we thought was the our best model for visualization:\n\nHowever, our score was not better than the original one. After visualization, we found that it could actually correct some errors in the Label of the original data set. Some individuals that were obviously starfish but not marked could also be found, that is to say, He was able to find some starfish that hadn't been marked before.\n\nSo I feel that maybe sometimes the performance of the model is higher than the score is because there are a few errors in the official data set.\n\nNo other meaning and thoughts, just a small discovery of training\n\n# indicators\n[![8a099a96b37c2df50f0e1b8a8f02fd6.png](https://i.postimg.cc/7ZQkYk6h/8a099a96b37c2df50f0e1b8a8f02fd6.png)](https://postimg.cc/0zm3ch71)\n\n# visualization:\n[![5578878036c2da1c9a6a44634605968.png](https://i.postimg.cc/Gmx3xNqx/5578878036c2da1c9a6a44634605968.png)](https://postimg.cc/MnX8Wskn)\n[![8f16e72d460070251281d7622ab6574.png](https://i.postimg.cc/fb1PtWm2/8f16e72d460070251281d7622ab6574.png)](https://postimg.cc/WD68xv9g)\n[![65448dcf88cd42095c1eadcd79aecc2.png](https://i.postimg.cc/KvQs7Dtg/65448dcf88cd42095c1eadcd79aecc2.png)](https://postimg.cc/HVc2d5Zp)\n----------------------------------------------------------------------------\n[![593a7088276fb51bce51fbf1e14ccd7.png](https://i.postimg.cc/DwMjHfjv/593a7088276fb51bce51fbf1e14ccd7.png)](https://postimg.cc/NKRkrtvS)\n[![a40dd00c72dc9592f3adb3bbdfc4ef2.png](https://i.postimg.cc/mZydRyX6/a40dd00c72dc9592f3adb3bbdfc4ef2.png)](https://postimg.cc/rzzGNxLS)\n\n\n\n\n\nDuring this time, we have encountered some problems\nAt first we tried YoloX, but it didn't work very well;\nThen we tried V5, and the effect of V5 made us reach 0.665;\nWe used Ensemble[V5 & X] and found that it was better without YoloX;\nWe also tried many, for example SAHI | COPY & Paste | YoloV5 & BigSize\nYoloV5 & BigSize;",
      "votes": null
    },
    {
      "id": "1687884",
      "postDate": "02/13/2022 08:19:38",
      "content": "<p>Does copy paste affect the leaderboard and your CV?</p>",
      "rawMarkdown": "Does copy paste affect the leaderboard and your CV?",
      "votes": null
    },
    {
      "id": "1687885",
      "postDate": "02/13/2022 08:20:30",
      "content": "<p>copypaste在cv上有用吗？</p>",
      "rawMarkdown": "copypaste在cv上有用吗？",
      "votes": null
    },
    {
      "id": "1687904",
      "postDate": "02/13/2022 08:32:42",
      "content": "<p>We have tried to achieve similar data enhancement through Mixup method on YoloX, but the effect is not obvious. Maybe it is because the paste between small targets sometimes has occlusion problem, so it does not get good results.</p>",
      "rawMarkdown": "We have tried to achieve similar data enhancement through Mixup method on YoloX, but the effect is not obvious. Maybe it is because the paste between small targets sometimes has occlusion problem, so it does not get good results.",
      "votes": null
    },
    {
      "id": "1689126",
      "postDate": "02/14/2022 03:57:24",
      "content": "<p>We do not try it in a right way because of the limit of DDL. But experiments shows that open the parameter [--image weights] in yolov5 will make model training faster、 more stable and get higher LB score. So I guess that taking a special copy-paste methods by the error rate of the objects may make sense(only a guess).</p>",
      "rawMarkdown": "We do not try it in a right way because of the limit of DDL. But experiments shows that open the parameter [--image weights] in yolov5 will make model training faster、 more stable and get higher LB score. So I guess that taking a special copy-paste methods by the error rate of the objects may make sense(only a guess).",
      "votes": null
    },
    {
      "id": "1690294",
      "postDate": "02/14/2022 20:48:20",
      "content": "<p>I have tried it just now. The copy paste methods do not work on my model, I only try that once.</p>",
      "rawMarkdown": "I have tried it just now. The copy paste methods do not work on my model, I only try that once.",
      "votes": null
    },
    {
      "id": "1690367",
      "postDate": "02/14/2022 22:32:13",
      "content": "<p>kind of late , but I use your model(one posted by Good Moon) and look at all the frame with FN , then fix bad/missing annotations + remove any frames with single annotation that look like it just randomly place. Then retrain the model for 3 epoch after with only the frames that had false negative after fixing them and the score go up from .66 to .703</p>\n<p>There probably more bad frames in the dataset that also need to be fix, I guess that one of the secret to higher LB. </p>\n<p>Also this probably why cv and lb detract , if you fix the dataset it probably make more sense </p>",
      "rawMarkdown": "kind of late , but I use your model(one posted by Good Moon) and look at all the frame with FN , then fix bad/missing annotations + remove any frames with single annotation that look like it just randomly place. Then retrain the model for 3 epoch after with only the frames that had false negative after fixing them and the score go up from .66 to .703\n\nThere probably more bad frames in the dataset that also need to be fix, I guess that one of the secret to higher LB. \n\nAlso this probably why cv and lb detract , if you fix the dataset it probably make more sense",
      "votes": null
    },
    {
      "id": "1690411",
      "postDate": "02/15/2022 00:01:02",
      "content": "<p>Haha, never late, friend, any sharing is welcome.  This is my first complete competition , I share my single model and hope it can let people to do more nice work.</p>",
      "rawMarkdown": "Haha, never late, friend, any sharing is welcome.  This is my first complete competition , I share my single model and hope it can let people to do more nice work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1687884,
      "author_name": "klawensliu",
      "author_url": "",
      "post_date": "02/13/2022 08:19:38",
      "content": "<p>Does copy paste affect the leaderboard and your CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1687885,
          "author_name": "klawensliu",
          "author_url": "",
          "post_date": "02/13/2022 08:20:30",
          "content": "<p>copypaste在cv上有用吗？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689126,
          "author_name": "freshair1996",
          "author_url": "",
          "post_date": "02/14/2022 03:57:24",
          "content": "<p>We do not try it in a right way because of the limit of DDL. But experiments shows that open the parameter [--image weights] in yolov5 will make model training faster、 more stable and get higher LB score. So I guess that taking a special copy-paste methods by the error rate of the objects may make sense(only a guess).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690294,
          "author_name": "freshair1996",
          "author_url": "",
          "post_date": "02/14/2022 20:48:20",
          "content": "<p>I have tried it just now. The copy paste methods do not work on my model, I only try that once.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690367,
          "author_name": "truonghuymai",
          "author_url": "",
          "post_date": "02/14/2022 22:32:13",
          "content": "<p>kind of late , but I use your model(one posted by Good Moon) and look at all the frame with FN , then fix bad/missing annotations + remove any frames with single annotation that look like it just randomly place. Then retrain the model for 3 epoch after with only the frames that had false negative after fixing them and the score go up from .66 to .703</p>\n<p>There probably more bad frames in the dataset that also need to be fix, I guess that one of the secret to higher LB. </p>\n<p>Also this probably why cv and lb detract , if you fix the dataset it probably make more sense </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690411,
          "author_name": "freshair1996",
          "author_url": "",
          "post_date": "02/15/2022 00:01:02",
          "content": "<p>Haha, never late, friend, any sharing is welcome.  This is my first complete competition , I share my single model and hope it can let people to do more nice work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1687904,
      "author_name": "yizhenglinsz",
      "author_url": "",
      "post_date": "02/13/2022 08:32:42",
      "content": "<p>We have tried to achieve similar data enhancement through Mixup method on YoloX, but the effect is not obvious. Maybe it is because the paste between small targets sometimes has occlusion problem, so it does not get good results.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1687677": "# Body\nIn the latest training, we adjusted the parameters to get what we thought was the our best model for visualization:\n\nHowever, our score was not better than the original one. After visualization, we found that it could actually correct some errors in the Label of the original data set. Some individuals that were obviously starfish but not marked could also be found, that is to say, He was able to find some starfish that hadn't been marked before.\n\nSo I feel that maybe sometimes the performance of the model is higher than the score is because there are a few errors in the official data set.\n\nNo other meaning and thoughts, just a small discovery of training\n\n# indicators\n[![8a099a96b37c2df50f0e1b8a8f02fd6.png](https://i.postimg.cc/7ZQkYk6h/8a099a96b37c2df50f0e1b8a8f02fd6.png)](https://postimg.cc/0zm3ch71)\n\n# visualization:\n[![5578878036c2da1c9a6a44634605968.png](https://i.postimg.cc/Gmx3xNqx/5578878036c2da1c9a6a44634605968.png)](https://postimg.cc/MnX8Wskn)\n[![8f16e72d460070251281d7622ab6574.png](https://i.postimg.cc/fb1PtWm2/8f16e72d460070251281d7622ab6574.png)](https://postimg.cc/WD68xv9g)\n[![65448dcf88cd42095c1eadcd79aecc2.png](https://i.postimg.cc/KvQs7Dtg/65448dcf88cd42095c1eadcd79aecc2.png)](https://postimg.cc/HVc2d5Zp)\n----------------------------------------------------------------------------\n[![593a7088276fb51bce51fbf1e14ccd7.png](https://i.postimg.cc/DwMjHfjv/593a7088276fb51bce51fbf1e14ccd7.png)](https://postimg.cc/NKRkrtvS)\n[![a40dd00c72dc9592f3adb3bbdfc4ef2.png](https://i.postimg.cc/mZydRyX6/a40dd00c72dc9592f3adb3bbdfc4ef2.png)](https://postimg.cc/rzzGNxLS)\n\n\n\n\n\nDuring this time, we have encountered some problems\nAt first we tried YoloX, but it didn't work very well;\nThen we tried V5, and the effect of V5 made us reach 0.665;\nWe used Ensemble[V5 & X] and found that it was better without YoloX;\nWe also tried many, for example SAHI | COPY & Paste | YoloV5 & BigSize\nYoloV5 & BigSize;",
    "1687884": "Does copy paste affect the leaderboard and your CV?",
    "1687885": "copypaste在cv上有用吗？",
    "1687904": "We have tried to achieve similar data enhancement through Mixup method on YoloX, but the effect is not obvious. Maybe it is because the paste between small targets sometimes has occlusion problem, so it does not get good results.",
    "1689126": "We do not try it in a right way because of the limit of DDL. But experiments shows that open the parameter [--image weights] in yolov5 will make model training faster、 more stable and get higher LB score. So I guess that taking a special copy-paste methods by the error rate of the objects may make sense(only a guess).",
    "1690294": "I have tried it just now. The copy paste methods do not work on my model, I only try that once.",
    "1690367": "kind of late , but I use your model(one posted by Good Moon) and look at all the frame with FN , then fix bad/missing annotations + remove any frames with single annotation that look like it just randomly place. Then retrain the model for 3 epoch after with only the frames that had false negative after fixing them and the score go up from .66 to .703\n\nThere probably more bad frames in the dataset that also need to be fix, I guess that one of the secret to higher LB. \n\nAlso this probably why cv and lb detract , if you fix the dataset it probably make more sense",
    "1690411": "Haha, never late, friend, any sharing is welcome.  This is my first complete competition , I share my single model and hope it can let people to do more nice work."
  },
  "source": "meta"
}