{
  "id": 290446,
  "title": "Error while reading annotation file livecell_coco_train.json using pycocotools",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/290446",
  "author_name": "",
  "post_date": "2021-11-24T17:47:13.908708700Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello </p>\n<p>While trying to include SH-SHY5Y cell line data for training, I tried to read the annotation file for the train set in the LIVECell dataset. However, an error is raised  while reading the annotation file. I already sent a GitHub issue (<a href=\"https://github.com/sartorius-research/LIVECell/issues/9\" target=\"_blank\">#9</a>) in the LIVECell repository, but it seems a previous issue (<a href=\"https://github.com/sartorius-research/LIVECell/issues/4\" target=\"_blank\">#4</a>) reported the same error, they solved the issue by downloading and reading the file a second time using the same method described here. From this solution it seems that the original annotation file was corrupted, but I am not sure if this is the case. Does anyone has been able to read this file without errors? or do you know another way to read the annotation file? </p>\n<pre><code>from pycocotools.coco import COCO\n\nannFile = '../input/sartorius-cell-instance-segmentation/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_train.json'\n\ncoco = COCO(annFile)\n</code></pre>\n<p>The following error is raised: </p>\n<pre><code>loading annotations into memory...\nDone (t=28.31s)\ncreating index...\n\n---------------------------------------------------------------------------\nTypeError                                 Traceback (most recent call last)\n/tmp/ipykernel_57/539935055.py in &lt;module&gt;\n     12     json.dump(dataset, f)\n     13 \n---&gt; 14 check_ANNT = COCO(ANNT_TRAIN_COPY)\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in __init__(self, annotation_file)\n     87             print('Done (t={:0.2f}s)'.format(time.time()- tic))\n     88             self.dataset = dataset\n---&gt; 89             self.createIndex()\n     90 \n     91     def createIndex(self):\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in createIndex(self)\n     96         if 'annotations' in self.dataset:\n     97             for ann in self.dataset['annotations']:\n---&gt; 98                 imgToAnns[ann['image_id']].append(ann)\n     99                 anns[ann['id']] = ann\n    100 \n\nTypeError: string indices must be integers\n</code></pre>",
  "messages": [
    {
      "id": "1594320",
      "postDate": "11/24/2021 17:47:13",
      "content": "<p>Hello </p>\n<p>While trying to include SH-SHY5Y cell line data for training, I tried to read the annotation file for the train set in the LIVECell dataset. However, an error is raised  while reading the annotation file. I already sent a GitHub issue (<a href=\"https://github.com/sartorius-research/LIVECell/issues/9\" target=\"_blank\">#9</a>) in the LIVECell repository, but it seems a previous issue (<a href=\"https://github.com/sartorius-research/LIVECell/issues/4\" target=\"_blank\">#4</a>) reported the same error, they solved the issue by downloading and reading the file a second time using the same method described here. From this solution it seems that the original annotation file was corrupted, but I am not sure if this is the case. Does anyone has been able to read this file without errors? or do you know another way to read the annotation file? </p>\n<pre><code>from pycocotools.coco import COCO\n\nannFile = '../input/sartorius-cell-instance-segmentation/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_train.json'\n\ncoco = COCO(annFile)\n</code></pre>\n<p>The following error is raised: </p>\n<pre><code>loading annotations into memory...\nDone (t=28.31s)\ncreating index...\n\n---------------------------------------------------------------------------\nTypeError                                 Traceback (most recent call last)\n/tmp/ipykernel_57/539935055.py in &lt;module&gt;\n     12     json.dump(dataset, f)\n     13 \n---&gt; 14 check_ANNT = COCO(ANNT_TRAIN_COPY)\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in __init__(self, annotation_file)\n     87             print('Done (t={:0.2f}s)'.format(time.time()- tic))\n     88             self.dataset = dataset\n---&gt; 89             self.createIndex()\n     90 \n     91     def createIndex(self):\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in createIndex(self)\n     96         if 'annotations' in self.dataset:\n     97             for ann in self.dataset['annotations']:\n---&gt; 98                 imgToAnns[ann['image_id']].append(ann)\n     99                 anns[ann['id']] = ann\n    100 \n\nTypeError: string indices must be integers\n</code></pre>",
      "rawMarkdown": "Hello \n\nWhile trying to include SH-SHY5Y cell line data for training, I tried to read the annotation file for the train set in the LIVECell dataset. However, an error is raised  while reading the annotation file. I already sent a GitHub issue ([#9](https://github.com/sartorius-research/LIVECell/issues/9)) in the LIVECell repository, but it seems a previous issue ([#4](https://github.com/sartorius-research/LIVECell/issues/4)) reported the same error, they solved the issue by downloading and reading the file a second time using the same method described here. From this solution it seems that the original annotation file was corrupted, but I am not sure if this is the case. Does anyone has been able to read this file without errors? or do you know another way to read the annotation file? \n\n```\nfrom pycocotools.coco import COCO\n\nannFile = '../input/sartorius-cell-instance-segmentation/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_train.json'\n\ncoco = COCO(annFile)\n```\n\n\nThe following error is raised: \n\n```\nloading annotations into memory...\nDone (t=28.31s)\ncreating index...\n\n---------------------------------------------------------------------------\nTypeError                                 Traceback (most recent call last)\n/tmp/ipykernel_57/539935055.py in <module>\n     12     json.dump(dataset, f)\n     13 \n---> 14 check_ANNT = COCO(ANNT_TRAIN_COPY)\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in __init__(self, annotation_file)\n     87             print('Done (t={:0.2f}s)'.format(time.time()- tic))\n     88             self.dataset = dataset\n---> 89             self.createIndex()\n     90 \n     91     def createIndex(self):\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in createIndex(self)\n     96         if 'annotations' in self.dataset:\n     97             for ann in self.dataset['annotations']:\n---> 98                 imgToAnns[ann['image_id']].append(ann)\n     99                 anns[ann['id']] = ann\n    100 \n\nTypeError: string indices must be integers\n```",
      "votes": null
    },
    {
      "id": "1594344",
      "postDate": "11/24/2021 18:05:06",
      "content": "<p>I read these annotations just as json and then converted and saved in appropriate format manually.</p>",
      "rawMarkdown": "I read these annotations just as json and then converted and saved in appropriate format manually.",
      "votes": null
    },
    {
      "id": "1594622",
      "postDate": "11/25/2021 03:06:14",
      "content": "<p>That problem is discussed <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "That problem is discussed [here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694)",
      "votes": null
    },
    {
      "id": "1595373",
      "postDate": "11/25/2021 16:27:28",
      "content": "<p>Thanks. I will try that too.</p>",
      "rawMarkdown": "Thanks. I will try that too.",
      "votes": null
    },
    {
      "id": "1595378",
      "postDate": "11/25/2021 16:31:45",
      "content": "<p>Thanks, I confirm that is exactly the same issue and will try the solution.</p>",
      "rawMarkdown": "Thanks, I confirm that is exactly the same issue and will try the solution.",
      "votes": null
    },
    {
      "id": "1603520",
      "postDate": "12/02/2021 14:53:50",
      "content": "<p>Do you manually generate LiveCell data after observing the data and merge it with the original data?</p>",
      "rawMarkdown": "Do you manually generate LiveCell data after observing the data and merge it with the original data?",
      "votes": null
    },
    {
      "id": "1612490",
      "postDate": "12/09/2021 02:47:03",
      "content": "<p>There is a solution to read the LIVECell annotation files in the discussion: <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694\" target=\"_blank\">LIVECELL_dataset_2021 annotationfiles are corrupt</a>. Using those files you could create the LIVECell train and validation sets and then merge them with the other sets. The only thing I found  out is that there is a <a href=\"https://www.tensorflow.org/versions/r2.0/api_docs/python/tf/data/Dataset#concatenate\" target=\"_blank\">Tensorflow method to concatenate two datasets</a>. </p>",
      "rawMarkdown": "There is a solution to read the LIVECell annotation files in the discussion: [LIVECELL_dataset_2021 annotationfiles are corrupt](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694). Using those files you could create the LIVECell train and validation sets and then merge them with the other sets. The only thing I found  out is that there is a [Tensorflow method to concatenate two datasets](https://www.tensorflow.org/versions/r2.0/api_docs/python/tf/data/Dataset#concatenate).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1594344,
      "author_name": "rednikotin",
      "author_url": "",
      "post_date": "11/24/2021 18:05:06",
      "content": "<p>I read these annotations just as json and then converted and saved in appropriate format manually.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1595373,
          "author_name": "aramos",
          "author_url": "",
          "post_date": "11/25/2021 16:27:28",
          "content": "<p>Thanks. I will try that too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1603520,
          "author_name": "zaopolearning",
          "author_url": "",
          "post_date": "12/02/2021 14:53:50",
          "content": "<p>Do you manually generate LiveCell data after observing the data and merge it with the original data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1612490,
          "author_name": "aramos",
          "author_url": "",
          "post_date": "12/09/2021 02:47:03",
          "content": "<p>There is a solution to read the LIVECell annotation files in the discussion: <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694\" target=\"_blank\">LIVECELL_dataset_2021 annotationfiles are corrupt</a>. Using those files you could create the LIVECell train and validation sets and then merge them with the other sets. The only thing I found  out is that there is a <a href=\"https://www.tensorflow.org/versions/r2.0/api_docs/python/tf/data/Dataset#concatenate\" target=\"_blank\">Tensorflow method to concatenate two datasets</a>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1594622,
      "author_name": "tyaiga",
      "author_url": "",
      "post_date": "11/25/2021 03:06:14",
      "content": "<p>That problem is discussed <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1595378,
          "author_name": "aramos",
          "author_url": "",
          "post_date": "11/25/2021 16:31:45",
          "content": "<p>Thanks, I confirm that is exactly the same issue and will try the solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1594320": "Hello \n\nWhile trying to include SH-SHY5Y cell line data for training, I tried to read the annotation file for the train set in the LIVECell dataset. However, an error is raised  while reading the annotation file. I already sent a GitHub issue ([#9](https://github.com/sartorius-research/LIVECell/issues/9)) in the LIVECell repository, but it seems a previous issue ([#4](https://github.com/sartorius-research/LIVECell/issues/4)) reported the same error, they solved the issue by downloading and reading the file a second time using the same method described here. From this solution it seems that the original annotation file was corrupted, but I am not sure if this is the case. Does anyone has been able to read this file without errors? or do you know another way to read the annotation file? \n\n```\nfrom pycocotools.coco import COCO\n\nannFile = '../input/sartorius-cell-instance-segmentation/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_train.json'\n\ncoco = COCO(annFile)\n```\n\n\nThe following error is raised: \n\n```\nloading annotations into memory...\nDone (t=28.31s)\ncreating index...\n\n---------------------------------------------------------------------------\nTypeError                                 Traceback (most recent call last)\n/tmp/ipykernel_57/539935055.py in <module>\n     12     json.dump(dataset, f)\n     13 \n---> 14 check_ANNT = COCO(ANNT_TRAIN_COPY)\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in __init__(self, annotation_file)\n     87             print('Done (t={:0.2f}s)'.format(time.time()- tic))\n     88             self.dataset = dataset\n---> 89             self.createIndex()\n     90 \n     91     def createIndex(self):\n\n/opt/conda/lib/python3.7/site-packages/pycocotools/coco.py in createIndex(self)\n     96         if 'annotations' in self.dataset:\n     97             for ann in self.dataset['annotations']:\n---> 98                 imgToAnns[ann['image_id']].append(ann)\n     99                 anns[ann['id']] = ann\n    100 \n\nTypeError: string indices must be integers\n```",
    "1594344": "I read these annotations just as json and then converted and saved in appropriate format manually.",
    "1594622": "That problem is discussed [here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694)",
    "1595373": "Thanks. I will try that too.",
    "1595378": "Thanks, I confirm that is exactly the same issue and will try the solution.",
    "1603520": "Do you manually generate LiveCell data after observing the data and merge it with the original data?",
    "1612490": "There is a solution to read the LIVECell annotation files in the discussion: [LIVECELL_dataset_2021 annotationfiles are corrupt](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/284694). Using those files you could create the LIVECell train and validation sets and then merge them with the other sets. The only thing I found  out is that there is a [Tensorflow method to concatenate two datasets](https://www.tensorflow.org/versions/r2.0/api_docs/python/tf/data/Dataset#concatenate)."
  },
  "source": "meta"
}