{
  "id": 284694,
  "title": "LIVECELL_dataset_2021 annotationfiles are corrupt",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/284694",
  "author_name": "Yamame🐟",
  "post_date": "2021-11-02T02:04:58.259000",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi everyone!<br>\nI tried to train my model using the LIVECELL_dataset_2021 provided by this comp.<br>\nHowever, when I read the .json files for any of the annotations using pycocotools, I get an error message like the below link and cannot parse .json files:</p>\n<p><a href=\"https://github.com/sartorius-research/LIVECell/issues/4\" target=\"_blank\">Unable to parse annotations using pycocotools package #4</a></p>\n<p>I followed the issue and re-downloaded the file from <a href=\"https://github.com/sartorius-research/LIVECell\" target=\"_blank\">repo</a>, then the error was resolved.</p>\n<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> check this, please.</p>",
  "messages": [
    {
      "id": 1567553,
      "postDate": "2021-11-02T02:04:58.260Z",
      "content": "<p>Hi everyone!<br>\nI tried to train my model using the LIVECELL_dataset_2021 provided by this comp.<br>\nHowever, when I read the .json files for any of the annotations using pycocotools, I get an error message like the below link and cannot parse .json files:</p>\n<p><a href=\"https://github.com/sartorius-research/LIVECell/issues/4\" target=\"_blank\">Unable to parse annotations using pycocotools package #4</a></p>\n<p>I followed the issue and re-downloaded the file from <a href=\"https://github.com/sartorius-research/LIVECell\" target=\"_blank\">repo</a>, then the error was resolved.</p>\n<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> check this, please.</p>",
      "rawMarkdown": "Hi everyone!\nI tried to train my model using the LIVECELL_dataset_2021 provided by this comp.\nHowever, when I read the .json files for any of the annotations using pycocotools, I get an error message like the below link and cannot parse .json files:\n\n[Unable to parse annotations using pycocotools package #4](https://github.com/sartorius-research/LIVECell/issues/4)\n\nI followed the issue and re-downloaded the file from [repo](https://github.com/sartorius-research/LIVECell), then the error was resolved.\n\n@addisonhoward check this, please.",
      "votes": 14
    },
    {
      "id": 1569736,
      "postDate": "2021-11-03T17:21:57.013Z",
      "content": "<p>With the provided \"LIVECell_dataset_2021\" data and with</p>\n<pre><code>from pycocotools.coco import COCO\nimport json\n\npath = \"whatever/path_to/file.json\"\n</code></pre>\n<p>Instead of a</p>\n<pre><code>coco = COCO(path)\n</code></pre>\n<p>use the following workaround:</p>\n<pre><code>s = json.load(open(path, 'r'))\ns[\"annotations\"] = list(s[\"annotations\"].values())\ncoco = COCO()\ncoco.dataset = s\ncoco.createIndex()\n</code></pre>\n<p>it is likely to work fine at least for the <code>pycocotools==2.0.2</code></p>",
      "rawMarkdown": "With the provided \"LIVECell_dataset_2021\" data and with\n```Python\nfrom pycocotools.coco import COCO\nimport json\n\npath = \"whatever/path_to/file.json\"\n```\nInstead of a\n```Python\ncoco = COCO(path)\n```\nuse the following workaround:\n```Python\ns = json.load(open(path, 'r'))\ns[\"annotations\"] = list(s[\"annotations\"].values())\ncoco = COCO()\ncoco.dataset = s\ncoco.createIndex()\n```\nit is likely to work fine at least for the `pycocotools==2.0.2`",
      "votes": 4
    },
    {
      "id": 1571061,
      "postDate": "2021-11-04T15:09:36.293Z",
      "content": "<p>Hi, thanks for pointing this out Yamame, I will see if we can update the annotations to the same version as is currently available at our LIVECell repo. Please use the workarounds provided or our official dataset resource in the meantime.</p>\n<p><a href=\"https://sartorius-research.github.io/LIVECell/\" target=\"_blank\">https://sartorius-research.github.io/LIVECell/</a></p>",
      "rawMarkdown": "Hi, thanks for pointing this out Yamame, I will see if we can update the annotations to the same version as is currently available at our LIVECell repo. Please use the workarounds provided or our official dataset resource in the meantime.\n\nhttps://sartorius-research.github.io/LIVECell/",
      "votes": 1
    },
    {
      "id": 1568924,
      "postDate": "2021-11-03T06:26:59.217Z",
      "content": "<p>So originally according to the COCO format the annotations are present in the form of 'list of dictionaries' where each element of the list is a dictionary(a single annotation). But in LiveCell2021 we have a dictionary of dictionaries whose keys are strings (annotation_id starting from '2', '3' and so on). Just a slight pre-processing after loading the json file should fix the issue.</p>",
      "rawMarkdown": "So originally according to the COCO format the annotations are present in the form of 'list of dictionaries' where each element of the list is a dictionary(a single annotation). But in LiveCell2021 we have a dictionary of dictionaries whose keys are strings (annotation_id starting from '2', '3' and so on). Just a slight pre-processing after loading the json file should fix the issue."
    },
    {
      "id": 1568015,
      "postDate": "2021-11-02T12:39:11.353Z",
      "content": "<p>I confirm seeing the same issue. </p>\n<p>As a workaround this is the command I used to get a working file:<br>\n<code>wget https://livecell-dataset.s3.eu-central-1.amazonaws.com/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_test.json</code></p>",
      "rawMarkdown": "I confirm seeing the same issue. \n\nAs a workaround this is the command I used to get a working file:\n`wget https://livecell-dataset.s3.eu-central-1.amazonaws.com/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_test.json`",
      "replies": [
        {
          "id": 1568218,
          "postDate": "2021-11-02T15:38:07.433Z",
          "content": "<p>And you can use this dataset w/o any problem ?</p>",
          "rawMarkdown": "And you can use this dataset w/o any problem ?"
        },
        {
          "id": 1568237,
          "postDate": "2021-11-02T15:58:00.657Z",
          "content": "<p><a href=\"https://www.kaggle.com/kfk42kfk\" target=\"_blank\">@kfk42kfk</a> It loaded into a pycocotools dataset correctly.</p>",
          "rawMarkdown": "@kfk42kfk It loaded into a pycocotools dataset correctly.",
          "votes": 1
        },
        {
          "id": 1568681,
          "postDate": "2021-11-03T01:13:12.663Z",
          "content": "<p>I can train my model using mmdetection</p>",
          "rawMarkdown": "I can train my model using mmdetection",
          "votes": 1
        },
        {
          "id": 1603497,
          "postDate": "2021-12-02T14:45:39.187Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1603512,
          "postDate": "2021-12-02T14:50:19.113Z",
          "content": "<p>So the loaded LiveCell can be combined with the raw data for training?</p>",
          "rawMarkdown": "So the loaded LiveCell can be combined with the raw data for training?"
        },
        {
          "id": 1628104,
          "postDate": "2021-12-24T14:14:14.630Z",
          "content": "<p>If I use these files uploaded to my private dataset the full path to tiff files seems wrong:</p>\n<p>random_instance = np.random.choice(len(train_ds))<br>\nprint('random instance:',random_instance)<br>\nd = train_ds[random_instance]<br>\nprint(d['file_name'],d['image_id'])<br>\nimg = tifffile.imread(d['file_name'])</p>\n<p>output:<br>\nrandom instance: 2553<br>\n./SKOV3_Phase_H4_1_01d08h00m_1.tif 1328760<br>\nFileNotFoundError: [Errno 2] No such file or directory: '/kaggle/working/SKOV3_Phase_H4_1_01d08h00m_1.tif'</p>",
          "rawMarkdown": "If I use these files uploaded to my private dataset the full path to tiff files seems wrong:\n\nrandom_instance = np.random.choice(len(train_ds))\nprint('random instance:',random_instance)\nd = train_ds[random_instance]\nprint(d['file_name'],d['image_id'])\nimg = tifffile.imread(d['file_name'])\n\n\noutput:\nrandom instance: 2553\n./SKOV3_Phase_H4_1_01d08h00m_1.tif 1328760\nFileNotFoundError: [Errno 2] No such file or directory: '/kaggle/working/SKOV3_Phase_H4_1_01d08h00m_1.tif'"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1569736,
      "author_name": "Sentinel-1",
      "author_url": "",
      "post_date": "2021-11-03T17:21:57.013000",
      "content": "<p>With the provided \"LIVECell_dataset_2021\" data and with</p>\n<pre><code>from pycocotools.coco import COCO\nimport json\n\npath = \"whatever/path_to/file.json\"\n</code></pre>\n<p>Instead of a</p>\n<pre><code>coco = COCO(path)\n</code></pre>\n<p>use the following workaround:</p>\n<pre><code>s = json.load(open(path, 'r'))\ns[\"annotations\"] = list(s[\"annotations\"].values())\ncoco = COCO()\ncoco.dataset = s\ncoco.createIndex()\n</code></pre>\n<p>it is likely to work fine at least for the <code>pycocotools==2.0.2</code></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1571061,
      "author_name": "CorporateResearchSartorius",
      "author_url": "",
      "post_date": "2021-11-04T15:09:36.293000",
      "content": "<p>Hi, thanks for pointing this out Yamame, I will see if we can update the annotations to the same version as is currently available at our LIVECell repo. Please use the workarounds provided or our official dataset resource in the meantime.</p>\n<p><a href=\"https://sartorius-research.github.io/LIVECell/\" target=\"_blank\">https://sartorius-research.github.io/LIVECell/</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1568924,
      "author_name": "Inumellonium",
      "author_url": "",
      "post_date": "2021-11-03T06:26:59.217000",
      "content": "<p>So originally according to the COCO format the annotations are present in the form of 'list of dictionaries' where each element of the list is a dictionary(a single annotation). But in LiveCell2021 we have a dictionary of dictionaries whose keys are strings (annotation_id starting from '2', '3' and so on). Just a slight pre-processing after loading the json file should fix the issue.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1568015,
      "author_name": "Slawek Biel",
      "author_url": "",
      "post_date": "2021-11-02T12:39:11.353000",
      "content": "<p>I confirm seeing the same issue. </p>\n<p>As a workaround this is the command I used to get a working file:<br>\n<code>wget https://livecell-dataset.s3.eu-central-1.amazonaws.com/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_test.json</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1568218,
          "author_name": "Furkan K",
          "author_url": "",
          "post_date": "2021-11-02T15:38:07.433000",
          "content": "<p>And you can use this dataset w/o any problem ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1568237,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-11-02T15:58:00.657000",
          "content": "<p><a href=\"https://www.kaggle.com/kfk42kfk\" target=\"_blank\">@kfk42kfk</a> It loaded into a pycocotools dataset correctly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1568681,
          "author_name": "Yamame🐟",
          "author_url": "",
          "post_date": "2021-11-03T01:13:12.663000",
          "content": "<p>I can train my model using mmdetection</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1603497,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-12-02T14:45:39.187000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1603512,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2021-12-02T14:50:19.113000",
          "content": "<p>So the loaded LiveCell can be combined with the raw data for training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1628104,
          "author_name": "Boris Polishchuk",
          "author_url": "",
          "post_date": "2021-12-24T14:14:14.630000",
          "content": "<p>If I use these files uploaded to my private dataset the full path to tiff files seems wrong:</p>\n<p>random_instance = np.random.choice(len(train_ds))<br>\nprint('random instance:',random_instance)<br>\nd = train_ds[random_instance]<br>\nprint(d['file_name'],d['image_id'])<br>\nimg = tifffile.imread(d['file_name'])</p>\n<p>output:<br>\nrandom instance: 2553<br>\n./SKOV3_Phase_H4_1_01d08h00m_1.tif 1328760<br>\nFileNotFoundError: [Errno 2] No such file or directory: '/kaggle/working/SKOV3_Phase_H4_1_01d08h00m_1.tif'</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1567553": "Hi everyone!\nI tried to train my model using the LIVECELL_dataset_2021 provided by this comp.\nHowever, when I read the .json files for any of the annotations using pycocotools, I get an error message like the below link and cannot parse .json files:\n\n[Unable to parse annotations using pycocotools package #4](https://github.com/sartorius-research/LIVECell/issues/4)\n\nI followed the issue and re-downloaded the file from [repo](https://github.com/sartorius-research/LIVECell), then the error was resolved.\n\n@addisonhoward check this, please.",
    "1569736": "With the provided \"LIVECell_dataset_2021\" data and with\n```Python\nfrom pycocotools.coco import COCO\nimport json\n\npath = \"whatever/path_to/file.json\"\n```\nInstead of a\n```Python\ncoco = COCO(path)\n```\nuse the following workaround:\n```Python\ns = json.load(open(path, 'r'))\ns[\"annotations\"] = list(s[\"annotations\"].values())\ncoco = COCO()\ncoco.dataset = s\ncoco.createIndex()\n```\nit is likely to work fine at least for the `pycocotools==2.0.2`",
    "1571061": "Hi, thanks for pointing this out Yamame, I will see if we can update the annotations to the same version as is currently available at our LIVECell repo. Please use the workarounds provided or our official dataset resource in the meantime.\n\nhttps://sartorius-research.github.io/LIVECell/",
    "1568924": "So originally according to the COCO format the annotations are present in the form of 'list of dictionaries' where each element of the list is a dictionary(a single annotation). But in LiveCell2021 we have a dictionary of dictionaries whose keys are strings (annotation_id starting from '2', '3' and so on). Just a slight pre-processing after loading the json file should fix the issue.",
    "1568015": "I confirm seeing the same issue. \n\nAs a workaround this is the command I used to get a working file:\n`wget https://livecell-dataset.s3.eu-central-1.amazonaws.com/LIVECell_dataset_2021/annotations/LIVECell/livecell_coco_test.json`"
  }
}