{
  "id": 228604,
  "title": "Help with submission",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/228604",
  "author_name": "Matthias",
  "post_date": "2021-03-25T13:39:57.472000",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I am trying to submit a test prediction, in which I predict class 0 with prop 1.0 for all cells. Somehow I get a submission error.<br>\nHow exactly do I generate the prediction string for each mask? At the moment generate the cell mask with the HPA segmentation that gives an array containing all masks for a given image. From the array I create the binary masks and prediction string for one image in the following way:</p>\n<pre><code>    predstring=\"\"\n    for cellid in range(1,ncells+1):\n        cell_mask_binary=np.copy(cell_mask)\n        cell_mask_binary[cell_mask_binary&lt;cellid]=0\n        cell_mask_binary[cell_mask_binary&gt;cellid]=0\n        cell_mask_binary=cell_mask_binary.astype(bool)\n        encoded_mask=encode_binary_mask(cell_mask_binary)\n        predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n</code></pre>\n<p>Is the mistake propably in the prediction string?</p>",
  "messages": [
    {
      "id": 1252199,
      "postDate": "2021-03-25T13:47:59.610Z",
      "content": "<p>Hi,</p>\n<p>can you share a prediction string for a image? </p>\n<p>If you do like this<br>\npredstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"</p>\n<p>it could be, that your last character of an image is a \" \". What would result in a Submission Scoring Error</p>",
      "rawMarkdown": "Hi,\n\ncan you share a prediction string for a image? \n\nIf you do like this\npredstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n\nit could be, that your last character of an image is a \" \". What would result in a Submission Scoring Error",
      "votes": 1,
      "replies": [
        {
          "id": 1252214,
          "postDate": "2021-03-25T13:58:28.970Z",
          "content": "<p>Hi Luca,</p>\n<p>you are right, my prediction strings all end with \" \". Thanks for your suggestion, I will try to get rid of that \" \" and try to submit again! </p>",
          "rawMarkdown": "Hi Luca,\n\nyou are right, my prediction strings all end with \" \". Thanks for your suggestion, I will try to get rid of that \" \" and try to submit again! "
        },
        {
          "id": 1252294,
          "postDate": "2021-03-25T15:07:09.783Z",
          "content": "<p>Unfortunately I still get an error. My full function looks like this</p>\n<pre><code>def create_submission(ids):\n    submission=pd.DataFrame()\n    submissioncolumns=pd.read_csv(\"../input/hpa-single-cell-image-classification/sample_submission.csv\").columns\n    count=0\n    print(\"processing \"+str(len(ids))+\" images\")\n    for image_id in ids:\n        count = count+1\n        print(\"do image nr \"+str(count)+\"/\"+str(len(ids)))\n        mt, er, nu, images = build_image_names(image_id)\n        # For nuclei\n        nuc_segmentations = segmentator.pred_nuclei(images[2])\n        # For full cells\n        cell_segmentations = segmentator.pred_cells(images)\n        nuclei_mask, cell_mask = label_cell(nuc_segmentations[0], cell_segmentations[0])\n        ncells=np.amax(cell_mask) \n        predstring=\"\"\n        im = Image.open('../input/hpa-single-cell-image-classification/test/'+image_id+\"_green.png\")\n        width, height = im.size\n        for cellid in range(1,ncells+1):\n            cell_mask_binary=np.copy(cell_mask)\n            cell_mask_binary[cell_mask_binary&lt;cellid]=0\n            cell_mask_binary[cell_mask_binary&gt;cellid]=0\n            cell_mask_binary=cell_mask_binary.astype(bool)\n            encoded_mask=encode_binary_mask(cell_mask_binary)\n            predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n        predstring = predstring[:-1]\n        imageprediction=pd.DataFrame({\"ID\":[image_id],\n                        \"ImageWidth\":[int(width)],\n                        \"ImageHeight\":[int(height)],\n                        \"PredictionString\":[predstring]})\n        submission=submission.append(imageprediction, ignore_index = True)\n\n    return submission\n</code></pre>\n<p>Do you see any other problems there? Thanks a lot in advance!</p>",
          "rawMarkdown": "Unfortunately I still get an error. My full function looks like this\n\n```\ndef create_submission(ids):\n    submission=pd.DataFrame()\n    submissioncolumns=pd.read_csv(\"../input/hpa-single-cell-image-classification/sample_submission.csv\").columns\n    count=0\n    print(\"processing \"+str(len(ids))+\" images\")\n    for image_id in ids:\n        count = count+1\n        print(\"do image nr \"+str(count)+\"/\"+str(len(ids)))\n        mt, er, nu, images = build_image_names(image_id)\n        # For nuclei\n        nuc_segmentations = segmentator.pred_nuclei(images[2])\n        # For full cells\n        cell_segmentations = segmentator.pred_cells(images)\n        nuclei_mask, cell_mask = label_cell(nuc_segmentations[0], cell_segmentations[0])\n        ncells=np.amax(cell_mask) \n        predstring=\"\"\n        im = Image.open('../input/hpa-single-cell-image-classification/test/'+image_id+\"_green.png\")\n        width, height = im.size\n        for cellid in range(1,ncells+1):\n            cell_mask_binary=np.copy(cell_mask)\n            cell_mask_binary[cell_mask_binary<cellid]=0\n            cell_mask_binary[cell_mask_binary>cellid]=0\n            cell_mask_binary=cell_mask_binary.astype(bool)\n            encoded_mask=encode_binary_mask(cell_mask_binary)\n            predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n        predstring = predstring[:-1]\n        imageprediction=pd.DataFrame({\"ID\":[image_id],\n                        \"ImageWidth\":[int(width)],\n                        \"ImageHeight\":[int(height)],\n                        \"PredictionString\":[predstring]})\n        submission=submission.append(imageprediction, ignore_index = True)\n            \n    return submission\n```\n\nDo you see any other problems there? Thanks a lot in advance!"
        },
        {
          "id": 1252381,
          "postDate": "2021-03-25T16:30:50.160Z",
          "content": "<p>the code looks ok. Where do you get the list ids?</p>",
          "rawMarkdown": "the code looks ok. Where do you get the list ids?"
        },
        {
          "id": 1252528,
          "postDate": "2021-03-25T18:47:50.370Z",
          "content": "<p>I uploaded all ids for the public train set here:<br>\n<a href=\"https://www.kaggle.com/maheller/ids-train-public\" target=\"_blank\">https://www.kaggle.com/maheller/ids-train-public</a><br>\nand loop over them (here over the first 5)</p>\n<pre><code>ids=pd.read_csv('../input/ids-train-public/ids_train.csv')['ID'][0:5]\noutput = create_submission(ids)\noutput.to_csv('submission.csv', index=False)\n</code></pre>",
          "rawMarkdown": "I uploaded all ids for the public train set here:\nhttps://www.kaggle.com/maheller/ids-train-public\nand loop over them (here over the first 5)\n\n```\nids=pd.read_csv('../input/ids-train-public/ids_train.csv')['ID'][0:5]\noutput = create_submission(ids)\noutput.to_csv('submission.csv', index=False)\n```"
        },
        {
          "id": 1253019,
          "postDate": "2021-03-26T09:46:30.710Z",
          "content": "<p>In order to get a valid submission, you have to make a prediction for the public and private test set. </p>\n<p>when you press the submit button. there are going to be more data in sample_submission.csv. So its bedder to take the IDs from there.</p>\n<p>Did you do that?</p>",
          "rawMarkdown": "In order to get a valid submission, you have to make a prediction for the public and private test set. \n\nwhen you press the submit button. there are going to be more data in sample_submission.csv. So its bedder to take the IDs from there.\n\nDid you do that?"
        },
        {
          "id": 1253025,
          "postDate": "2021-03-26T09:54:32.047Z",
          "content": "<p>No I did not, because I thought if I skip some ids, i.e. I don't put them on the csv file, they are just not counted. But that iss not true? If I don't make a prediction for ALL labels the submission is not valid?</p>",
          "rawMarkdown": "No I did not, because I thought if I skip some ids, i.e. I don't put them on the csv file, they are just not counted. But that iss not true? If I don't make a prediction for ALL labels the submission is not valid?"
        },
        {
          "id": 1253059,
          "postDate": "2021-03-26T10:34:33.973Z",
          "content": "<p>you have to put all IDs in sample_submission.csv in you prediction.</p>\n<p>look at <a href=\"https://www.kaggle.com/gody7334/local-submission\" target=\"_blank\">https://www.kaggle.com/gody7334/local-submission</a></p>",
          "rawMarkdown": "you have to put all IDs in sample_submission.csv in you prediction.\n\nlook at https://www.kaggle.com/gody7334/local-submission"
        },
        {
          "id": 1254099,
          "postDate": "2021-03-27T10:51:50.650Z",
          "content": "<p>It finally worked, thanks a lot!</p>",
          "rawMarkdown": "It finally worked, thanks a lot!"
        }
      ]
    },
    {
      "id": 1252187,
      "postDate": "2021-03-25T13:39:57.473Z",
      "content": "<p>Hi all,</p>\n<p>I am trying to submit a test prediction, in which I predict class 0 with prop 1.0 for all cells. Somehow I get a submission error.<br>\nHow exactly do I generate the prediction string for each mask? At the moment generate the cell mask with the HPA segmentation that gives an array containing all masks for a given image. From the array I create the binary masks and prediction string for one image in the following way:</p>\n<pre><code>    predstring=\"\"\n    for cellid in range(1,ncells+1):\n        cell_mask_binary=np.copy(cell_mask)\n        cell_mask_binary[cell_mask_binary&lt;cellid]=0\n        cell_mask_binary[cell_mask_binary&gt;cellid]=0\n        cell_mask_binary=cell_mask_binary.astype(bool)\n        encoded_mask=encode_binary_mask(cell_mask_binary)\n        predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n</code></pre>\n<p>Is the mistake propably in the prediction string?</p>",
      "rawMarkdown": "Hi all,\n\nI am trying to submit a test prediction, in which I predict class 0 with prop 1.0 for all cells. Somehow I get a submission error.\nHow exactly do I generate the prediction string for each mask? At the moment generate the cell mask with the HPA segmentation that gives an array containing all masks for a given image. From the array I create the binary masks and prediction string for one image in the following way:\n\n        predstring=\"\"\n        for cellid in range(1,ncells+1):\n            cell_mask_binary=np.copy(cell_mask)\n            cell_mask_binary[cell_mask_binary<cellid]=0\n            cell_mask_binary[cell_mask_binary>cellid]=0\n            cell_mask_binary=cell_mask_binary.astype(bool)\n            encoded_mask=encode_binary_mask(cell_mask_binary)\n            predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n\nIs the mistake propably in the prediction string?\n\n",
      "votes": 1
    },
    {
      "id": 1252211,
      "postDate": "2021-03-25T13:57:47.950Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1252199,
      "author_name": "LucaMTB",
      "author_url": "",
      "post_date": "2021-03-25T13:47:59.610000",
      "content": "<p>Hi,</p>\n<p>can you share a prediction string for a image? </p>\n<p>If you do like this<br>\npredstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"</p>\n<p>it could be, that your last character of an image is a \" \". What would result in a Submission Scoring Error</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1252214,
          "author_name": "Matthias",
          "author_url": "",
          "post_date": "2021-03-25T13:58:28.970000",
          "content": "<p>Hi Luca,</p>\n<p>you are right, my prediction strings all end with \" \". Thanks for your suggestion, I will try to get rid of that \" \" and try to submit again! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1252294,
          "author_name": "Matthias",
          "author_url": "",
          "post_date": "2021-03-25T15:07:09.783000",
          "content": "<p>Unfortunately I still get an error. My full function looks like this</p>\n<pre><code>def create_submission(ids):\n    submission=pd.DataFrame()\n    submissioncolumns=pd.read_csv(\"../input/hpa-single-cell-image-classification/sample_submission.csv\").columns\n    count=0\n    print(\"processing \"+str(len(ids))+\" images\")\n    for image_id in ids:\n        count = count+1\n        print(\"do image nr \"+str(count)+\"/\"+str(len(ids)))\n        mt, er, nu, images = build_image_names(image_id)\n        # For nuclei\n        nuc_segmentations = segmentator.pred_nuclei(images[2])\n        # For full cells\n        cell_segmentations = segmentator.pred_cells(images)\n        nuclei_mask, cell_mask = label_cell(nuc_segmentations[0], cell_segmentations[0])\n        ncells=np.amax(cell_mask) \n        predstring=\"\"\n        im = Image.open('../input/hpa-single-cell-image-classification/test/'+image_id+\"_green.png\")\n        width, height = im.size\n        for cellid in range(1,ncells+1):\n            cell_mask_binary=np.copy(cell_mask)\n            cell_mask_binary[cell_mask_binary&lt;cellid]=0\n            cell_mask_binary[cell_mask_binary&gt;cellid]=0\n            cell_mask_binary=cell_mask_binary.astype(bool)\n            encoded_mask=encode_binary_mask(cell_mask_binary)\n            predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n        predstring = predstring[:-1]\n        imageprediction=pd.DataFrame({\"ID\":[image_id],\n                        \"ImageWidth\":[int(width)],\n                        \"ImageHeight\":[int(height)],\n                        \"PredictionString\":[predstring]})\n        submission=submission.append(imageprediction, ignore_index = True)\n\n    return submission\n</code></pre>\n<p>Do you see any other problems there? Thanks a lot in advance!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1252381,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2021-03-25T16:30:50.160000",
          "content": "<p>the code looks ok. Where do you get the list ids?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1252528,
          "author_name": "Matthias",
          "author_url": "",
          "post_date": "2021-03-25T18:47:50.370000",
          "content": "<p>I uploaded all ids for the public train set here:<br>\n<a href=\"https://www.kaggle.com/maheller/ids-train-public\" target=\"_blank\">https://www.kaggle.com/maheller/ids-train-public</a><br>\nand loop over them (here over the first 5)</p>\n<pre><code>ids=pd.read_csv('../input/ids-train-public/ids_train.csv')['ID'][0:5]\noutput = create_submission(ids)\noutput.to_csv('submission.csv', index=False)\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253019,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2021-03-26T09:46:30.710000",
          "content": "<p>In order to get a valid submission, you have to make a prediction for the public and private test set. </p>\n<p>when you press the submit button. there are going to be more data in sample_submission.csv. So its bedder to take the IDs from there.</p>\n<p>Did you do that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253025,
          "author_name": "Matthias",
          "author_url": "",
          "post_date": "2021-03-26T09:54:32.047000",
          "content": "<p>No I did not, because I thought if I skip some ids, i.e. I don't put them on the csv file, they are just not counted. But that iss not true? If I don't make a prediction for ALL labels the submission is not valid?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253059,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2021-03-26T10:34:33.973000",
          "content": "<p>you have to put all IDs in sample_submission.csv in you prediction.</p>\n<p>look at <a href=\"https://www.kaggle.com/gody7334/local-submission\" target=\"_blank\">https://www.kaggle.com/gody7334/local-submission</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254099,
          "author_name": "Matthias",
          "author_url": "",
          "post_date": "2021-03-27T10:51:50.650000",
          "content": "<p>It finally worked, thanks a lot!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1252211,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-25T13:57:47.950000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1252199": "Hi,\n\ncan you share a prediction string for a image? \n\nIf you do like this\npredstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n\nit could be, that your last character of an image is a \" \". What would result in a Submission Scoring Error",
    "1252187": "Hi all,\n\nI am trying to submit a test prediction, in which I predict class 0 with prop 1.0 for all cells. Somehow I get a submission error.\nHow exactly do I generate the prediction string for each mask? At the moment generate the cell mask with the HPA segmentation that gives an array containing all masks for a given image. From the array I create the binary masks and prediction string for one image in the following way:\n\n        predstring=\"\"\n        for cellid in range(1,ncells+1):\n            cell_mask_binary=np.copy(cell_mask)\n            cell_mask_binary[cell_mask_binary<cellid]=0\n            cell_mask_binary[cell_mask_binary>cellid]=0\n            cell_mask_binary=cell_mask_binary.astype(bool)\n            encoded_mask=encode_binary_mask(cell_mask_binary)\n            predstring+=\"0 1.0 \"+str(encoded_mask, \"utf-8\")+\" \"\n\nIs the mistake propably in the prediction string?\n\n",
    "1252211": ""
  }
}