{
  "id": 296170,
  "title": "How to handle unlabeled COTS data ? ",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/296170",
  "author_name": "",
  "post_date": "2021-12-20T07:46:55.460281500Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I'm trying to use unlabeled COTS data with <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">this split method</a> using YOLOX. To handle unlabeled data, if there is no bbox, assign [] and 0 in  bbox and area of image_annotations. </p>\n<p>But, when start to train, I'm facing below error.</p>\n<p>ERROR | yolox.core.launch:98 - An error has been caught in function 'launch', process 'MainProcess' (617), thread 'MainThread' (139651519448960) &amp; IndexError: list index out of range</p>\n<p>Has anyone tried to handle unlabeled COTS data ? </p>\n<p>If assign [0.0, 0.0, 0.0, 0.0] bbox instead [] bbox, is there any potential to hurt model performance or negative effect ? </p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1623750",
      "postDate": "12/20/2021 07:46:55",
      "content": "<p>Hi all,</p>\n<p>I'm trying to use unlabeled COTS data with <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">this split method</a> using YOLOX. To handle unlabeled data, if there is no bbox, assign [] and 0 in  bbox and area of image_annotations. </p>\n<p>But, when start to train, I'm facing below error.</p>\n<p>ERROR | yolox.core.launch:98 - An error has been caught in function 'launch', process 'MainProcess' (617), thread 'MainThread' (139651519448960) &amp; IndexError: list index out of range</p>\n<p>Has anyone tried to handle unlabeled COTS data ? </p>\n<p>If assign [0.0, 0.0, 0.0, 0.0] bbox instead [] bbox, is there any potential to hurt model performance or negative effect ? </p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi all,\n\nI'm trying to use unlabeled COTS data with [this split method](https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences) using YOLOX. To handle unlabeled data, if there is no bbox, assign [] and 0 in  bbox and area of image_annotations. \n\nBut, when start to train, I'm facing below error.\n\nERROR | yolox.core.launch:98 - An error has been caught in function 'launch', process 'MainProcess' (617), thread 'MainThread' (139651519448960) & IndexError: list index out of range\n\nHas anyone tried to handle unlabeled COTS data ? \n\nIf assign [0.0, 0.0, 0.0, 0.0] bbox instead [] bbox, is there any potential to hurt model performance or negative effect ? \n\nThanks.",
      "votes": null
    },
    {
      "id": "1623751",
      "postDate": "12/20/2021 07:54:19",
      "content": "<p>You have to take only data with bboxes ….</p>\n<pre><code>import ast\ndef num_boxes(annotations):\n    annotations = ast.literal_eval(annotations)\n    return len(annotations)\ndf['num_bbox'] = df['annotations'].apply(lambda x: num_boxes(x))\ndf = df.query(\"num_bbox&gt;0\")\n</code></pre>",
      "rawMarkdown": "You have to take only data with bboxes ....\n\n```\nimport ast\ndef num_boxes(annotations):\n    annotations = ast.literal_eval(annotations)\n    return len(annotations)\ndf['num_bbox'] = df['annotations'].apply(lambda x: num_boxes(x))\ndf = df.query(\"num_bbox>0\")\n```",
      "votes": null
    },
    {
      "id": "1623757",
      "postDate": "12/20/2021 08:03:11",
      "content": "<p>Is there not any potential to improve model performance by using unlabeled data ? </p>",
      "rawMarkdown": "Is there not any potential to improve model performance by using unlabeled data ?",
      "votes": null
    },
    {
      "id": "1623765",
      "postDate": "12/20/2021 08:12:30",
      "content": "<p>I was looking for starfish in unlabeled data … there is no (probably) object instances. </p>",
      "rawMarkdown": "I was looking for starfish in unlabeled data ... there is no (probably) object instances.",
      "votes": null
    },
    {
      "id": "1623910",
      "postDate": "12/20/2021 11:33:56",
      "content": "<p>Actually, there are object instances unlabeled. You can find them at the moment labels suddenly disappear/appear in the same sequence. </p>",
      "rawMarkdown": "Actually, there are object instances unlabeled. You can find them at the moment labels suddenly disappear/appear in the same sequence.",
      "votes": null
    },
    {
      "id": "1623930",
      "postDate": "12/20/2021 11:45:29",
      "content": "<p><a href=\"https://www.kaggle.com/joon9502\" target=\"_blank\">@joon9502</a>  but in labeled data … I agree </p>",
      "rawMarkdown": "joon9502  but in labeled data ... I agree",
      "votes": null
    },
    {
      "id": "1623931",
      "postDate": "12/20/2021 11:46:25",
      "content": "<p><a href=\"https://www.kaggle.com/seongwook93\" target=\"_blank\">@seongwook93</a> I have not followed you firstly.<br>\nI agree … you can improve model using unlabeled images - backgroud images can reduce False Positives (FP)</p>",
      "rawMarkdown": "seongwook93 I have not followed you firstly.\nI agree ... you can improve model using unlabeled images - backgroud images can reduce False Positives (FP)",
      "votes": null
    },
    {
      "id": "1631771",
      "postDate": "12/28/2021 19:00:06",
      "content": "<ol>\n<li>train with +ve images only</li>\n<li>test with the -ve images (i.e. those without labels). if there is no FP, you can ignore them. if there is FP, just label the FP with new \"others class\" and add the image to training set</li>\n</ol>",
      "rawMarkdown": "1. train with +ve images only\n2. test with the -ve images (i.e. those without labels). if there is no FP, you can ignore them. if there is FP, just label the FP with new \"others class\" and add the image to training set",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1623751,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "12/20/2021 07:54:19",
      "content": "<p>You have to take only data with bboxes ….</p>\n<pre><code>import ast\ndef num_boxes(annotations):\n    annotations = ast.literal_eval(annotations)\n    return len(annotations)\ndf['num_bbox'] = df['annotations'].apply(lambda x: num_boxes(x))\ndf = df.query(\"num_bbox&gt;0\")\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1623757,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "12/20/2021 08:03:11",
          "content": "<p>Is there not any potential to improve model performance by using unlabeled data ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623765,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/20/2021 08:12:30",
          "content": "<p>I was looking for starfish in unlabeled data … there is no (probably) object instances. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623910,
          "author_name": "joon9502",
          "author_url": "",
          "post_date": "12/20/2021 11:33:56",
          "content": "<p>Actually, there are object instances unlabeled. You can find them at the moment labels suddenly disappear/appear in the same sequence. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623930,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/20/2021 11:45:29",
          "content": "<p><a href=\"https://www.kaggle.com/joon9502\" target=\"_blank\">@joon9502</a>  but in labeled data … I agree </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1623931,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/20/2021 11:46:25",
          "content": "<p><a href=\"https://www.kaggle.com/seongwook93\" target=\"_blank\">@seongwook93</a> I have not followed you firstly.<br>\nI agree … you can improve model using unlabeled images - backgroud images can reduce False Positives (FP)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1631771,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/28/2021 19:00:06",
      "content": "<ol>\n<li>train with +ve images only</li>\n<li>test with the -ve images (i.e. those without labels). if there is no FP, you can ignore them. if there is FP, just label the FP with new \"others class\" and add the image to training set</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1623750": "Hi all,\n\nI'm trying to use unlabeled COTS data with [this split method](https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences) using YOLOX. To handle unlabeled data, if there is no bbox, assign [] and 0 in  bbox and area of image_annotations. \n\nBut, when start to train, I'm facing below error.\n\nERROR | yolox.core.launch:98 - An error has been caught in function 'launch', process 'MainProcess' (617), thread 'MainThread' (139651519448960) & IndexError: list index out of range\n\nHas anyone tried to handle unlabeled COTS data ? \n\nIf assign [0.0, 0.0, 0.0, 0.0] bbox instead [] bbox, is there any potential to hurt model performance or negative effect ? \n\nThanks.",
    "1623751": "You have to take only data with bboxes ....\n\n```\nimport ast\ndef num_boxes(annotations):\n    annotations = ast.literal_eval(annotations)\n    return len(annotations)\ndf['num_bbox'] = df['annotations'].apply(lambda x: num_boxes(x))\ndf = df.query(\"num_bbox>0\")\n```",
    "1623757": "Is there not any potential to improve model performance by using unlabeled data ?",
    "1623765": "I was looking for starfish in unlabeled data ... there is no (probably) object instances.",
    "1623910": "Actually, there are object instances unlabeled. You can find them at the moment labels suddenly disappear/appear in the same sequence.",
    "1623930": "joon9502  but in labeled data ... I agree",
    "1623931": "seongwook93 I have not followed you firstly.\nI agree ... you can improve model using unlabeled images - backgroud images can reduce False Positives (FP)",
    "1631771": "1. train with +ve images only\n2. test with the -ve images (i.e. those without labels). if there is no FP, you can ignore them. if there is FP, just label the FP with new \"others class\" and add the image to training set"
  },
  "source": "meta"
}