{
  "id": 225473,
  "title": "Build \"Submission file\"",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/225473",
  "author_name": "GitMach",
  "post_date": "2021-03-12T12:10:14.871000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi, <br>\nI'm struggling with the submission file, specially for creating the \"prediction string'<br>\nBelow, few lines from my dataframe \"prepare_to_submission\" </p>\n<pre><code>id    width   weight  label   confidence\n0    9b1d4b27-6946-4b86-a818-8e91029c3dfa    1728    1728    0.0 0.3484996557235718\n1    5af9bb81-719a-4374-9a12-5e5665329df6    1728    1728    0.0 0.9666309952735901\n2    8adb17b7-ce2c-4721-bf34-c1806c72b4d4    2048    2048    4.0 0.19799888134002686\n3    c6379d83-6b05-4a29-8f56-05a381d94d2a    3072    3072    0.0 0.2289401888847351\n4    277b3f6d-099b-4b6d-8592-a06ad6f52beb    2048    2048    0.0 0.0970616340637207\n</code></pre>\n<p>So, I've the ID, width, weight, label, and confidence</p>\n<p>When I try to encode I get this string :</p>\n<pre><code>b'OQAAAGIAAAAxAAAAZAAAADQAAABiAAAAMgAAADcAAAAtAAAANgAAADkAAA......'\n</code></pre>\n<p>which doesn't look pretty much with the one in the sample_submission</p>\n<p>What I've tried :</p>\n<pre><code>for c in np_sub[:,[0,1,2,4,5]][:,[0,3,4]]: ## Where np_sub is == prepare_to_submission.values\n    print(np.ascontiguousarray(c))\n    base64_str = base64.b64encode(np.ascontiguousarray(c))\n    print(base64_str)\n</code></pre>\n<p>My question is simple, can I get the prediction_string from this dataframe or do I need to use the \"pycocotools\"  and  mask as mutils?<br>\nIf you have any ideas, you're welcome, any help is truly appreciated</p>",
  "messages": [
    {
      "id": 1235802,
      "postDate": "2021-03-12T13:30:37.403Z",
      "content": "<p>Hello,<br>\nThe prediction string for each test image is expected to be in the format:</p>\n<blockquote>\n  <p>LabelA1 ConfidenceA1 EncodedMaskA1 LabelA2 ConfidenceA2 EncodedMaskA2</p>\n</blockquote>\n<p>where an EncodedMask is the encoded binary mask of one cell. You can extract locations of single cells in an image using the <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">cell segmentation tool</a>. The <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\">evaluation page</a> has more details about how the binary mask of a cell is expected to be encoded and also some code that may help.<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/notebook\" target=\"_blank\"> Darek's notebook</a> where he extracts the encoded masks for each cell in the test images may also be useful. You could also just use this dataset, but this would score 0 in the private test set.</p>",
      "rawMarkdown": "Hello,\nThe prediction string for each test image is expected to be in the format:\n> LabelA1 ConfidenceA1 EncodedMaskA1 LabelA2 ConfidenceA2 EncodedMaskA2\n\nwhere an EncodedMask is the encoded binary mask of one cell. You can extract locations of single cells in an image using the [cell segmentation tool](https://github.com/CellProfiling/HPA-Cell-Segmentation). The [evaluation page](https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation) has more details about how the binary mask of a cell is expected to be encoded and also some code that may help.[ Darek's notebook](https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/notebook) where he extracts the encoded masks for each cell in the test images may also be useful. You could also just use this dataset, but this would score 0 in the private test set.",
      "votes": 1
    },
    {
      "id": 1235721,
      "postDate": "2021-03-12T12:10:14.873Z",
      "content": "<p>Hi, <br>\nI'm struggling with the submission file, specially for creating the \"prediction string'<br>\nBelow, few lines from my dataframe \"prepare_to_submission\" </p>\n<pre><code>id    width   weight  label   confidence\n0    9b1d4b27-6946-4b86-a818-8e91029c3dfa    1728    1728    0.0 0.3484996557235718\n1    5af9bb81-719a-4374-9a12-5e5665329df6    1728    1728    0.0 0.9666309952735901\n2    8adb17b7-ce2c-4721-bf34-c1806c72b4d4    2048    2048    4.0 0.19799888134002686\n3    c6379d83-6b05-4a29-8f56-05a381d94d2a    3072    3072    0.0 0.2289401888847351\n4    277b3f6d-099b-4b6d-8592-a06ad6f52beb    2048    2048    0.0 0.0970616340637207\n</code></pre>\n<p>So, I've the ID, width, weight, label, and confidence</p>\n<p>When I try to encode I get this string :</p>\n<pre><code>b'OQAAAGIAAAAxAAAAZAAAADQAAABiAAAAMgAAADcAAAAtAAAANgAAADkAAA......'\n</code></pre>\n<p>which doesn't look pretty much with the one in the sample_submission</p>\n<p>What I've tried :</p>\n<pre><code>for c in np_sub[:,[0,1,2,4,5]][:,[0,3,4]]: ## Where np_sub is == prepare_to_submission.values\n    print(np.ascontiguousarray(c))\n    base64_str = base64.b64encode(np.ascontiguousarray(c))\n    print(base64_str)\n</code></pre>\n<p>My question is simple, can I get the prediction_string from this dataframe or do I need to use the \"pycocotools\"  and  mask as mutils?<br>\nIf you have any ideas, you're welcome, any help is truly appreciated</p>",
      "rawMarkdown": "Hi, \nI'm struggling with the submission file, specially for creating the \"prediction string'\nBelow, few lines from my dataframe \"prepare_to_submission\" \n```\nid\twidth\tweight\tlabel\tconfidence\n0\t9b1d4b27-6946-4b86-a818-8e91029c3dfa\t1728\t1728\t0.0\t0.3484996557235718\n1\t5af9bb81-719a-4374-9a12-5e5665329df6\t1728\t1728\t0.0\t0.9666309952735901\n2\t8adb17b7-ce2c-4721-bf34-c1806c72b4d4\t2048\t2048\t4.0\t0.19799888134002686\n3\tc6379d83-6b05-4a29-8f56-05a381d94d2a\t3072\t3072\t0.0\t0.2289401888847351\n4\t277b3f6d-099b-4b6d-8592-a06ad6f52beb\t2048\t2048\t0.0\t0.0970616340637207\n```\nSo, I've the ID, width, weight, label, and confidence\n\nWhen I try to encode I get this string :\n\n```\nb'OQAAAGIAAAAxAAAAZAAAADQAAABiAAAAMgAAADcAAAAtAAAANgAAADkAAA......'\n\n```\nwhich doesn't look pretty much with the one in the sample_submission\n\nWhat I've tried :\n\n```\nfor c in np_sub[:,[0,1,2,4,5]][:,[0,3,4]]: ## Where np_sub is == prepare_to_submission.values\n    print(np.ascontiguousarray(c))\n    base64_str = base64.b64encode(np.ascontiguousarray(c))\n    print(base64_str)\n```\n\nMy question is simple, can I get the prediction_string from this dataframe or do I need to use the \"pycocotools\"  and  mask as mutils?\nIf you have any ideas, you're welcome, any help is truly appreciated\n\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1235802,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-03-12T13:30:37.403000",
      "content": "<p>Hello,<br>\nThe prediction string for each test image is expected to be in the format:</p>\n<blockquote>\n  <p>LabelA1 ConfidenceA1 EncodedMaskA1 LabelA2 ConfidenceA2 EncodedMaskA2</p>\n</blockquote>\n<p>where an EncodedMask is the encoded binary mask of one cell. You can extract locations of single cells in an image using the <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">cell segmentation tool</a>. The <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\">evaluation page</a> has more details about how the binary mask of a cell is expected to be encoded and also some code that may help.<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/notebook\" target=\"_blank\"> Darek's notebook</a> where he extracts the encoded masks for each cell in the test images may also be useful. You could also just use this dataset, but this would score 0 in the private test set.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1235802": "Hello,\nThe prediction string for each test image is expected to be in the format:\n> LabelA1 ConfidenceA1 EncodedMaskA1 LabelA2 ConfidenceA2 EncodedMaskA2\n\nwhere an EncodedMask is the encoded binary mask of one cell. You can extract locations of single cells in an image using the [cell segmentation tool](https://github.com/CellProfiling/HPA-Cell-Segmentation). The [evaluation page](https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation) has more details about how the binary mask of a cell is expected to be encoded and also some code that may help.[ Darek's notebook](https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/notebook) where he extracts the encoded masks for each cell in the test images may also be useful. You could also just use this dataset, but this would score 0 in the private test set.",
    "1235721": "Hi, \nI'm struggling with the submission file, specially for creating the \"prediction string'\nBelow, few lines from my dataframe \"prepare_to_submission\" \n```\nid\twidth\tweight\tlabel\tconfidence\n0\t9b1d4b27-6946-4b86-a818-8e91029c3dfa\t1728\t1728\t0.0\t0.3484996557235718\n1\t5af9bb81-719a-4374-9a12-5e5665329df6\t1728\t1728\t0.0\t0.9666309952735901\n2\t8adb17b7-ce2c-4721-bf34-c1806c72b4d4\t2048\t2048\t4.0\t0.19799888134002686\n3\tc6379d83-6b05-4a29-8f56-05a381d94d2a\t3072\t3072\t0.0\t0.2289401888847351\n4\t277b3f6d-099b-4b6d-8592-a06ad6f52beb\t2048\t2048\t0.0\t0.0970616340637207\n```\nSo, I've the ID, width, weight, label, and confidence\n\nWhen I try to encode I get this string :\n\n```\nb'OQAAAGIAAAAxAAAAZAAAADQAAABiAAAAMgAAADcAAAAtAAAANgAAADkAAA......'\n\n```\nwhich doesn't look pretty much with the one in the sample_submission\n\nWhat I've tried :\n\n```\nfor c in np_sub[:,[0,1,2,4,5]][:,[0,3,4]]: ## Where np_sub is == prepare_to_submission.values\n    print(np.ascontiguousarray(c))\n    base64_str = base64.b64encode(np.ascontiguousarray(c))\n    print(base64_str)\n```\n\nMy question is simple, can I get the prediction_string from this dataframe or do I need to use the \"pycocotools\"  and  mask as mutils?\nIf you have any ideas, you're welcome, any help is truly appreciated\n\n"
  }
}