{
  "id": 67604,
  "title": "Welcome!",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/67604",
  "author_name": "Emma Lundberg",
  "post_date": "2018-10-03T21:37:35.502000",
  "votes": 22,
  "comment_count": 86,
  "views": 0,
  "content": "<p>Welcome to the Human Protein Atlas Image Classification Challenge. Help us classify subcellular protein patterns and win great prizes, provided by Leica Microsystems and Nvidia.</p>\n\n<p>We wish you all the best and look forward to see how the Kaggle community will tackle this problem and what new models will be developed. We will do our best to answer any questions the community might have.</p>\n\n<p>A note about the data:\nThere are two versions of the data: PNGs and TIFFs. The PNG sets are much smaller than the TIFF sets. The PNG sets have been downscaled to 512x512 pixels and converted to 8 bit, while the TIFF sets keep the full 2048x2048 or 3072x3072 pixels but have been converted to 8 bit, from the 16 bit of the raw data.</p>\n\n<p>When running the evaluation of the special prize, we will use the TIFF test set. For the main competition the evaluation procedure doesn't differ between PNG and TIFF sets.</p>\n\n<p>More information can be found on the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/data\">Data page</a> and on the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions\">Special Prize Instructions page</a>.</p>\n\n<p>Have fun!</p>\n\n<p>Update Dec 6: There are 15 samples that have a different pixel size (4096, 4096), see the attached files for a list of these.</p>",
  "messages": [
    {
      "id": 398306,
      "postDate": "2018-10-03T21:37:35.503Z",
      "content": "<p>Welcome to the Human Protein Atlas Image Classification Challenge. Help us classify subcellular protein patterns and win great prizes, provided by Leica Microsystems and Nvidia.</p>\n\n<p>We wish you all the best and look forward to see how the Kaggle community will tackle this problem and what new models will be developed. We will do our best to answer any questions the community might have.</p>\n\n<p>A note about the data:\nThere are two versions of the data: PNGs and TIFFs. The PNG sets are much smaller than the TIFF sets. The PNG sets have been downscaled to 512x512 pixels and converted to 8 bit, while the TIFF sets keep the full 2048x2048 or 3072x3072 pixels but have been converted to 8 bit, from the 16 bit of the raw data.</p>\n\n<p>When running the evaluation of the special prize, we will use the TIFF test set. For the main competition the evaluation procedure doesn't differ between PNG and TIFF sets.</p>\n\n<p>More information can be found on the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/data\">Data page</a> and on the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions\">Special Prize Instructions page</a>.</p>\n\n<p>Have fun!</p>\n\n<p>Update Dec 6: There are 15 samples that have a different pixel size (4096, 4096), see the attached files for a list of these.</p>",
      "rawMarkdown": "Welcome to the Human Protein Atlas Image Classification Challenge. Help us classify subcellular protein patterns and win great prizes, provided by Leica Microsystems and Nvidia.\n\nWe wish you all the best and look forward to see how the Kaggle community will tackle this problem and what new models will be developed. We will do our best to answer any questions the community might have.\n\nA note about the data:\nThere are two versions of the data: PNGs and TIFFs. The PNG sets are much smaller than the TIFF sets. The PNG sets have been downscaled to 512x512 pixels and converted to 8 bit, while the TIFF sets keep the full 2048x2048 or 3072x3072 pixels but have been converted to 8 bit, from the 16 bit of the raw data.\n\nWhen running the evaluation of the special prize, we will use the TIFF test set. For the main competition the evaluation procedure doesn't differ between PNG and TIFF sets.\n\nMore information can be found on the [Data page](https://www.kaggle.com/c/human-protein-atlas-image-classification/data) and on the [Special Prize Instructions page](https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions).\n\nHave fun!\n\nUpdate Dec 6: There are 15 samples that have a different pixel size (4096, 4096), see the attached files for a list of these.",
      "votes": 22
    },
    {
      "id": 432897,
      "postDate": "2018-12-04T12:48:55Z",
      "content": "<p>Are kaggle or the organisers going to comment on the data leak?</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395</a></p>",
      "rawMarkdown": "Are kaggle or the organisers going to comment on the data leak?\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395",
      "votes": 1,
      "replies": [
        {
          "id": 434648,
          "postDate": "2018-12-06T18:17:24.997Z",
          "content": "<p>Yes, <a href=\"/philculliton\">@philculliton</a> did answer in the thread.</p>",
          "rawMarkdown": "Yes, @philculliton did answer in the thread."
        }
      ]
    },
    {
      "id": 423743,
      "postDate": "2018-11-19T01:11:38.620Z",
      "content": "<p>In the data set there're many classes has co-occurence with the other classes. For example, One image is labeled as  both Nucleoplasm and Cytosol. So, how to understand by looking at images that why it is labeled to the both classes? If I see an image only labeled with the Nucleoplasm or Cytosol by combining all 4 channels, I can't find any visual similarities. If I see an image classified with the both classes then what should I look into so it can explain that the image is belong to both classes. </p>",
      "rawMarkdown": "In the data set there're many classes has co-occurence with the other classes. For example, One image is labeled as  both Nucleoplasm and Cytosol. So, how to understand by looking at images that why it is labeled to the both classes? If I see an image only labeled with the Nucleoplasm or Cytosol by combining all 4 channels, I can't find any visual similarities. If I see an image classified with the both classes then what should I look into so it can explain that the image is belong to both classes. ",
      "votes": 1,
      "replies": [
        {
          "id": 423803,
          "postDate": "2018-11-19T04:51:02.790Z",
          "content": "<p>I literary tried that yesterday. From my understanding, Cytosol is the same as Cytoplasm, but not the same as Nucleoplasm. Unfortunately, my current model can easily surpass my poor classification. I cannot imagine how EVE players can classify images so accurately with such noisy data. Some green labels can hardly be seen by human eyes, some (0a022f5e-bbca-11e8-b2bc-ac1f6b6435d0) are corrupted, and some do not even have a cell-like structure (0c1d7284-bbc1-11e8-b2bb-ac1f6b6435d0).</p>",
          "rawMarkdown": "I literary tried that yesterday. From my understanding, Cytosol is the same as Cytoplasm, but not the same as Nucleoplasm. Unfortunately, my current model can easily surpass my poor classification. I cannot imagine how EVE players can classify images so accurately with such noisy data. Some green labels can hardly be seen by human eyes, some (0a022f5e-bbca-11e8-b2bc-ac1f6b6435d0) are corrupted, and some do not even have a cell-like structure (0c1d7284-bbc1-11e8-b2bb-ac1f6b6435d0).",
          "votes": 1
        },
        {
          "id": 433422,
          "postDate": "2018-12-05T04:02:04.937Z",
          "content": "<p>It is often easier to find the visual simalarities/differences when not combining all four channels, in particular the yellow is hard to distinguish from the green. The dictionary in the HPA can help you understand the different patterns <a href=\"https://www.proteinatlas.org/learn/dictionary/cell\">https://www.proteinatlas.org/learn/dictionary/cell</a>.</p>\n\n<p>For the example of cytosol, nucleoplasm vs cytosol+nucleoplasm you can see the differences by looking at the overlap with the 'red' and 'blue' markers respectively. The blue is a marker for the nucleus, so if there is green staining within the blue area, it should have the label ‘nucleoplasm’. The red is microtubules that branch out through the cytosol, so if there is green staining within the red  area, it should have the label ‘cytosol’. If there is green staining in both the red and green area, it should have both labels.</p>",
          "rawMarkdown": "It is often easier to find the visual simalarities/differences when not combining all four channels, in particular the yellow is hard to distinguish from the green. The dictionary in the HPA can help you understand the different patterns https://www.proteinatlas.org/learn/dictionary/cell.\n\nFor the example of cytosol, nucleoplasm vs cytosol+nucleoplasm you can see the differences by looking at the overlap with the 'red' and 'blue' markers respectively. The blue is a marker for the nucleus, so if there is green staining within the blue area, it should have the label ‘nucleoplasm’. The red is microtubules that branch out through the cytosol, so if there is green staining within the red  area, it should have the label ‘cytosol’. If there is green staining in both the red and green area, it should have both labels.",
          "votes": 1
        }
      ]
    },
    {
      "id": 419779,
      "postDate": "2018-11-12T14:51:33.037Z",
      "content": "<p>I found that there are also 4096x4096 px images in the training dataset. 7 to be exact.\nId's:\n0858d008-bb9e-11e8-b2b9-ac1f6b6435d0\n16196e56-bba2-11e8-b2b9-ac1f6b6435d0\n1625b570-bbc6-11e8-b2bc-ac1f6b6435d0\n28be7d5e-bbbf-11e8-b2ba-ac1f6b6435d0\n558dfed4-bba3-11e8-b2b9-ac1f6b6435d0\nc3322376-bbc0-11e8-b2bb-ac1f6b6435d0\nf2998bbe-bbab-11e8-b2ba-ac1f6b6435d0</p>\n\n<p>This lead to a bug that cost me some calculation and debugging time. Consider adding this information to your Post above and the competition description where you list the image sizes.</p>\n\n<p>Cheers,\nBasti</p>",
      "rawMarkdown": "I found that there are also 4096x4096 px images in the training dataset. 7 to be exact.\nId's:\n0858d008-bb9e-11e8-b2b9-ac1f6b6435d0\n16196e56-bba2-11e8-b2b9-ac1f6b6435d0\n1625b570-bbc6-11e8-b2bc-ac1f6b6435d0\n28be7d5e-bbbf-11e8-b2ba-ac1f6b6435d0\n558dfed4-bba3-11e8-b2b9-ac1f6b6435d0\nc3322376-bbc0-11e8-b2bb-ac1f6b6435d0\nf2998bbe-bbab-11e8-b2ba-ac1f6b6435d0\n\nThis lead to a bug that cost me some calculation and debugging time. Consider adding this information to your Post above and the competition description where you list the image sizes.\n\nCheers,\nBasti",
      "votes": 1,
      "replies": [
        {
          "id": 435285,
          "postDate": "2018-12-07T19:50:58.940Z",
          "content": "<p>Hi Basti,</p>\n\n<p>Thanks for bringing this to our attention. We've analysed the shape of the complete data set now and updated the original post above with csv files of all tiff files with 4096x4096 px.</p>",
          "rawMarkdown": "Hi Basti,\n\nThanks for bringing this to our attention. We've analysed the shape of the complete data set now and updated the original post above with csv files of all tiff files with 4096x4096 px.",
          "votes": 2
        }
      ]
    },
    {
      "id": 414929,
      "postDate": "2018-11-03T23:10:13.027Z",
      "content": "<p>I can't upload and submit data in these two days. I tried almost all methods and couldn't solve them. What should I do?</p>",
      "rawMarkdown": "I can't upload and submit data in these two days. I tried almost all methods and couldn't solve them. What should I do?",
      "votes": 1
    },
    {
      "id": 411984,
      "postDate": "2018-10-29T10:51:45.017Z",
      "content": "<p>Hi Emma,\nWhile there exist most of the works done for protein subcellular classification using textual data, why we have to do same with images. How cost effective is images compare to text data prediction. What is the intention behind giving these image dataset for the same classification which is done using textual data already? As it will be questionable by the reviewer when a paper is submitted for publication. Any response would be appreciable :) Thanks</p>",
      "rawMarkdown": "Hi Emma,\nWhile there exist most of the works done for protein subcellular classification using textual data, why we have to do same with images. How cost effective is images compare to text data prediction. What is the intention behind giving these image dataset for the same classification which is done using textual data already? As it will be questionable by the reviewer when a paper is submitted for publication. Any response would be appreciable :) Thanks",
      "votes": 1,
      "replies": [
        {
          "id": 422960,
          "postDate": "2018-11-17T06:21:24.163Z",
          "content": "<p>Hi Tulasi,</p>\n\n<p>Predicting protein subcellular localization from textual data (here I am assuming that you mean the amino acid sequence of the protein, or nucleic acid sequence of the gene encoding the protein) is nowhere near a solved problem, unfortunately. It is possible to reliably predict localization with signal peptides directing them to certain places (mitochondria, nuclear localization signals etc) and transmembrane regions. But there are no methods (to my knowledge) that can predict anywhere near the number of locations and substructures included in this challenge (for example distinguish between if a protein is in the nucleoli or in the nucleoli fibrillar center). Prediction of multiple locations from sequences is also a very hard task, most likely because this cellular phenomena is influenced by post-translational regulation and interactions.</p>\n\n<p>Best,\nEmma</p>",
          "rawMarkdown": "Hi Tulasi,\n\nPredicting protein subcellular localization from textual data (here I am assuming that you mean the amino acid sequence of the protein, or nucleic acid sequence of the gene encoding the protein) is nowhere near a solved problem, unfortunately. It is possible to reliably predict localization with signal peptides directing them to certain places (mitochondria, nuclear localization signals etc) and transmembrane regions. But there are no methods (to my knowledge) that can predict anywhere near the number of locations and substructures included in this challenge (for example distinguish between if a protein is in the nucleoli or in the nucleoli fibrillar center). Prediction of multiple locations from sequences is also a very hard task, most likely because this cellular phenomena is influenced by post-translational regulation and interactions.\n\nBest,\nEmma"
        }
      ]
    },
    {
      "id": 399804,
      "postDate": "2018-10-06T20:07:54.690Z",
      "content": "<p>Looking forward to participating! Do you think it is possible to get top performing results using the PNG set or would this only be possible by training on the TIFF set?</p>",
      "rawMarkdown": "Looking forward to participating! Do you think it is possible to get top performing results using the PNG set or would this only be possible by training on the TIFF set?",
      "votes": 1,
      "replies": [
        {
          "id": 400450,
          "postDate": "2018-10-08T10:47:59.320Z",
          "content": "<p>Hello Carlo!</p>\n\n<p>The TIFFs should in theory contain more information than the PNGs which are downscaled to 512x512px from the same TIFFs. We would expect that the PNGs are of good enough quality to get a top performing result but we do not want to promise anything. </p>\n\n<p>Looking at the PNGs they are at least good enough for human annotation.</p>",
          "rawMarkdown": "Hello Carlo!\n\nThe TIFFs should in theory contain more information than the PNGs which are downscaled to 512x512px from the same TIFFs. We would expect that the PNGs are of good enough quality to get a top performing result but we do not want to promise anything. \n\nLooking at the PNGs they are at least good enough for human annotation.",
          "votes": 1
        },
        {
          "id": 406966,
          "postDate": "2018-10-20T05:12:55.103Z",
          "content": "<p>The special prize is interesting, would love that GPU.  But I see no chance that I will try for it.\n1.  The level of detail and work to get the special prize submission is way too much for a hobby data scientist.\n2.  The special prize requires the prediction to use the TIFF files - OK that makes some sense, but the TIFF are dumbed down, so why bother.  In addition, I would need to use the TIFF to make sure things work - that would mean for me that I should go with TIFF now, cause stuff happens.\n3.  Without any way to benchmark speed, I am not sure I want to put in all that work and know have some idea of my speed compared to others.  Any potential way there can be a speed leaderboard?</p>\n\n<p>Do like the regular challenge and look forward to learning all about proteins. </p>",
          "rawMarkdown": "The special prize is interesting, would love that GPU.  But I see no chance that I will try for it.\n1.  The level of detail and work to get the special prize submission is way too much for a hobby data scientist.\n2.  The special prize requires the prediction to use the TIFF files - OK that makes some sense, but the TIFF are dumbed down, so why bother.  In addition, I would need to use the TIFF to make sure things work - that would mean for me that I should go with TIFF now, cause stuff happens.\n3.  Without any way to benchmark speed, I am not sure I want to put in all that work and know have some idea of my speed compared to others.  Any potential way there can be a speed leaderboard?\n\nDo like the regular challenge and look forward to learning all about proteins. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 398935,
      "postDate": "2018-10-04T20:29:25.887Z",
      "content": "<p>Thanks for your reply, Martin.\nIt is very clear now.</p>",
      "rawMarkdown": "Thanks for your reply, Martin.\nIt is very clear now.",
      "votes": 1
    },
    {
      "id": 401008,
      "postDate": "2018-10-09T09:11:02.100Z",
      "content": "<p>Following up on the questions about the hardware limits of the special prize evaluation, we want to clarify that we only limit the hardware of the prediction process that we will perform when evaluating submissions. Training can still be done without hardware restrictions, since this is done before submission.</p>\n\n<p>We understand that the hardware limits can still impact top performance, but we are interested to find models that balance performance, speed and necessary hardware resources. The reason why we have put the limits on the hardware is to match our smart-microscopy system workflow, as mentioned above.</p>\n\n<p>Thanks for the feedback and please continue the discussion on this topic. We want to hear what you think.</p>",
      "rawMarkdown": "Following up on the questions about the hardware limits of the special prize evaluation, we want to clarify that we only limit the hardware of the prediction process that we will perform when evaluating submissions. Training can still be done without hardware restrictions, since this is done before submission.\n\nWe understand that the hardware limits can still impact top performance, but we are interested to find models that balance performance, speed and necessary hardware resources. The reason why we have put the limits on the hardware is to match our smart-microscopy system workflow, as mentioned above.\n\nThanks for the feedback and please continue the discussion on this topic. We want to hear what you think.",
      "votes": 2,
      "replies": [
        {
          "id": 411758,
          "postDate": "2018-10-29T00:31:04.493Z",
          "content": "<p>Martin,</p>\n\n<p>On the requirements for the special prize it mentions that an entrypoint is needed for training in the docker image. Does this mean that you want included a way to train completely from scratch with default initialization? Or should the training session load pretrained weights? Are there any size limits to the zip? My saved model and weights is less than 10MB, but the 512x512 input I am using is over 75GB.</p>\n\n<p>The model I'm working on takes a long time to train, but predicts quickly. Even with a GPU it currently takes days of training. In development I'm working using Jupyter fine tuning the model and training sessions, trying to take the best notes I can but it may be difficult to reproduce exactly.</p>\n\n<p>For the prediction performance, it is pretty straightforward. I am using Keras so it would pretty much look like:</p>\n\n<ul>\n<li>define model functions, loading functions for the input path</li>\n<li>load hdf5 model with saved weights from the model path</li>\n<li>predict and save submission to output path</li>\n</ul>\n\n<p>Thanks,</p>\n\n<p>Brian</p>",
          "rawMarkdown": "Martin,\n\nOn the requirements for the special prize it mentions that an entrypoint is needed for training in the docker image. Does this mean that you want included a way to train completely from scratch with default initialization? Or should the training session load pretrained weights? Are there any size limits to the zip? My saved model and weights is less than 10MB, but the 512x512 input I am using is over 75GB.\n\nThe model I'm working on takes a long time to train, but predicts quickly. Even with a GPU it currently takes days of training. In development I'm working using Jupyter fine tuning the model and training sessions, trying to take the best notes I can but it may be difficult to reproduce exactly.\n\nFor the prediction performance, it is pretty straightforward. I am using Keras so it would pretty much look like:\n\n\n- define model functions, loading functions for the input path\n- load hdf5 model with saved weights from the model path\n- predict and save submission to output path\n\nThanks,\n\nBrian",
          "votes": 1
        },
        {
          "id": 449329,
          "postDate": "2019-01-03T01:22:29.663Z",
          "content": "<p>Hi Brian,</p>\n\n<p>does my answer here answer your questions?\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67604#449328\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67604#449328</a></p>",
          "rawMarkdown": "Hi Brian,\n\ndoes my answer here answer your questions?\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67604#449328"
        },
        {
          "id": 450462,
          "postDate": "2019-01-05T02:18:38.030Z",
          "content": "<p>It is still a little opaque:</p>\n\n<pre><code>- the last link points to this very discussion. Are you referring to the message above the link?\n- \"...the readme file in the submission...\" Where are the requirements on this file described? I don't see it anywhere. Is it in the \"submission instructions\". Can you give a link to these?\n- I am not 100% sure but I doubt that anyone can reproduce their exact submitted predictions with a new run of the code. There are random elements in all packages that do not seem reproducible. I have tried in fastai and was not able to reproduce things exactly. This has been discussed in the threads.\n</code></pre>",
          "rawMarkdown": "It is still a little opaque:\n\n    - the last link points to this very discussion. Are you referring to the message above the link?\n    - \"...the readme file in the submission...\" Where are the requirements on this file described? I don't see it anywhere. Is it in the \"submission instructions\". Can you give a link to these?\n    - I am not 100% sure but I doubt that anyone can reproduce their exact submitted predictions with a new run of the code. There are random elements in all packages that do not seem reproducible. I have tried in fastai and was not able to reproduce things exactly. This has been discussed in the threads."
        },
        {
          "id": 452309,
          "postDate": "2019-01-08T14:51:32.213Z",
          "content": "<p>Hi Pete,</p>\n\n<ul>\n<li>The last link should be a permalink to my answer to Maksim further down in this thread.</li>\n<li>Here is a link to the special prize submission intructions:\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions\">https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions</a></li>\n<li>In the instructions we write that the submitted model should maintain its score +/- 1 percentage point. Are those limits within reason for a reproduced model too?</li>\n</ul>",
          "rawMarkdown": "Hi Pete,\n\n- The last link should be a permalink to my answer to Maksim further down in this thread.\n- Here is a link to the special prize submission intructions:\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions\n- In the instructions we write that the submitted model should maintain its score +/- 1 percentage point. Are those limits within reason for a reproduced model too?",
          "votes": 1
        },
        {
          "id": 452363,
          "postDate": "2019-01-08T16:14:15.420Z",
          "content": "<p>Hi Martin,</p>\n\n<p>Where is the <code>submission_sample.zip</code> file?</p>",
          "rawMarkdown": "Hi Martin,\n\nWhere is the `submission_sample.zip` file?"
        },
        {
          "id": 452459,
          "postDate": "2019-01-08T18:52:18.313Z",
          "content": "<p>For me, the chances are remote (but not yet strictly zero) of being in the money. But I would say that 1% is maybe a bit small - it is 0.6 vs 0.606. With variations in the randoms, which appear unavoidable, one can get this type of variation. I guess it would be OK if you allow 2-3 tries to get this number.</p>\n\n<p>Regarding the instructions : this seems to apply only to the \"special prize\". Do they in fact apply to the main contest submission as well, which I assume is independent of the special prize?</p>",
          "rawMarkdown": "For me, the chances are remote (but not yet strictly zero) of being in the money. But I would say that 1% is maybe a bit small - it is 0.6 vs 0.606. With variations in the randoms, which appear unavoidable, one can get this type of variation. I guess it would be OK if you allow 2-3 tries to get this number.\n\nRegarding the instructions : this seems to apply only to the \"special prize\". Do they in fact apply to the main contest submission as well, which I assume is independent of the special prize?"
        },
        {
          "id": 463904,
          "postDate": "2019-01-30T21:27:26.703Z",
          "content": "<p>Hi Martin,</p>\n\n<p>Will late submissions be enabled for this competition? Wishing to review and improve my solution.</p>\n\n<p>Even when submitting an older submission, I get:\n\"Your submission was unsuccessful. Please try again, and if you continue to have issues, please let us know the details of your submission and competition on the Product Feedback forum.\"</p>\n\n<p>Thanks</p>",
          "rawMarkdown": "Hi Martin,\n\nWill late submissions be enabled for this competition? Wishing to review and improve my solution.\n\nEven when submitting an older submission, I get:\n\"Your submission was unsuccessful. Please try again, and if you continue to have issues, please let us know the details of your submission and competition on the Product Feedback forum.\"\n\nThanks"
        },
        {
          "id": 466257,
          "postDate": "2019-02-05T01:07:15.657Z",
          "content": "<p>Hello, </p>\n\n<p>Late Submissions should be enabled for this competition already. We've run a test and were able to successfully submit a sample CSV file to the Late Submissions page. Could you please try submitting again and let us know if the issue persists? Thank you!</p>",
          "rawMarkdown": "Hello, \n\nLate Submissions should be enabled for this competition already. We've run a test and were able to successfully submit a sample CSV file to the Late Submissions page. Could you please try submitting again and let us know if the issue persists? Thank you!"
        }
      ]
    },
    {
      "id": 398314,
      "postDate": "2018-10-03T22:34:02.707Z",
      "content": "<blockquote>\n  <p>Top performing teams are invited to make the model code and any external data used to generate their final submission available to ...</p>\n</blockquote>\n\n<p>_</p>\n\n<blockquote>\n  <p>The winner will be the fastest model that maintains its submitted F1 score in the main prize competition +/- 1 percentage point.</p>\n</blockquote>\n\n<p>That is a bit too optimistic, isn't it? Top submissions will need significantly more FLOPS than what 2 CPU cores can produce.</p>",
      "rawMarkdown": "&gt;Top performing teams are invited to make the model code and any external data used to generate their final submission available to ...\n\n_\n&gt; The winner will be the fastest model that maintains its submitted F1 score in the main prize competition +/- 1 percentage point.\n\nThat is a bit too optimistic, isn't it? Top submissions will need significantly more FLOPS than what 2 CPU cores can produce.",
      "votes": 2,
      "replies": [
        {
          "id": 398533,
          "postDate": "2018-10-04T08:38:06.097Z",
          "content": "<p>Seems this approach is also penalizing better performing models (as the performance needs to be carried over to the special prize category and it is likely that better performing models will require more compute).</p>",
          "rawMarkdown": "Seems this approach is also penalizing better performing models (as the performance needs to be carried over to the special prize category and it is likely that better performing models will require more compute).\n"
        },
        {
          "id": 413479,
          "postDate": "2018-11-01T02:48:27.903Z",
          "content": "<p>once you have top performance model, you can use neural architecture search or knowledge distillation method to find smaller model with similar accuracy. Good example is nasnet and mobile-nasnet and there are such tools in tensorflow.</p>",
          "rawMarkdown": "once you have top performance model, you can use neural architecture search or knowledge distillation method to find smaller model with similar accuracy. Good example is nasnet and mobile-nasnet and there are such tools in tensorflow.",
          "votes": 1
        },
        {
          "id": 413955,
          "postDate": "2018-11-01T21:26:54.480Z",
          "content": "<p>I focused from the beginning on a small model. So far using CPU only it can predict one 512x512 image in ~850ms. </p>",
          "rawMarkdown": "I focused from the beginning on a small model. So far using CPU only it can predict one 512x512 image in ~850ms. "
        }
      ]
    },
    {
      "id": 514526,
      "postDate": "2019-04-11T17:42:39.417Z",
      "content": "<p>Hi,\nCan you share how you converted 16-bit tiff files to the 8-bit images?</p>",
      "rawMarkdown": "Hi,\nCan you share how you converted 16-bit tiff files to the 8-bit images?"
    },
    {
      "id": 443203,
      "postDate": "2018-12-21T07:57:16.593Z",
      "content": "<p>Hello Emma,\nCould you please explain the rules concerning reproducibility of the final submission:\nShould the whole training process be completely reproducible or only the inference part?\nThank you</p>",
      "rawMarkdown": "Hello Emma,\nCould you please explain the rules concerning reproducibility of the final submission:\nShould the whole training process be completely reproducible or only the inference part?\nThank you",
      "replies": [
        {
          "id": 449328,
          "postDate": "2019-01-03T01:19:23.867Z",
          "content": "<p>Hi Maksim,</p>\n\n<p>The training process should be described in the readme file in the submission and there should be entry points that prepares the data, starts the training etc as described in the submission instructions. We should be able to train a model from the instructions and reproduce your prediction result on the test set.</p>\n\n<p>If external data or pre-trained models have been used, that should either be included in the submission, linked to, or made available to us upon request. You also need to make a post describing what external data or pre-trainined models have been used in the corresponding official discussion thread here in the forum.</p>",
          "rawMarkdown": "Hi Maksim,\n\nThe training process should be described in the readme file in the submission and there should be entry points that prepares the data, starts the training etc as described in the submission instructions. We should be able to train a model from the instructions and reproduce your prediction result on the test set.\n\nIf external data or pre-trained models have been used, that should either be included in the submission, linked to, or made available to us upon request. You also need to make a post describing what external data or pre-trainined models have been used in the corresponding official discussion thread here in the forum."
        }
      ]
    },
    {
      "id": 431055,
      "postDate": "2018-12-01T14:41:40.787Z",
      "content": "<p>Hi, Emma Lundberg.\nI have a question. If there have colorful sample image, no like these grey?</p>",
      "rawMarkdown": "Hi, Emma Lundberg.\nI have a question. If there have colorful sample image, no like these grey?",
      "replies": [
        {
          "id": 431370,
          "postDate": "2018-12-02T05:52:03.723Z",
          "content": "<p>I think we get to color between the lines :)</p>",
          "rawMarkdown": "I think we get to color between the lines :)"
        }
      ]
    },
    {
      "id": 427150,
      "postDate": "2018-11-24T18:46:58.733Z",
      "content": "<p>Hi Emma,</p>\n\n<p>I have a problem with downloading tiff images.\nWhen I download from google bucket via browser (tried Chrome and Firefox), anywhere in the middle there is some connection error (maybe short disconnect or smth like that), and downloading interrupts, and browsers are not clever enough to continue...</p>\n\n<p>I wonder if there is any other fool-proof possibility of downloading tiff images (e.g. torrent file)?</p>\n\n<p>Did anybody have a similar problem maybe?</p>\n\n<p>Regards,\nMaksim</p>",
      "rawMarkdown": "Hi Emma,\n\nI have a problem with downloading tiff images.\nWhen I download from google bucket via browser (tried Chrome and Firefox), anywhere in the middle there is some connection error (maybe short disconnect or smth like that), and downloading interrupts, and browsers are not clever enough to continue...\n\nI wonder if there is any other fool-proof possibility of downloading tiff images (e.g. torrent file)?\n\nDid anybody have a similar problem maybe?\n\nRegards,\nMaksim\n",
      "replies": [
        {
          "id": 427229,
          "postDate": "2018-11-24T22:24:19.783Z",
          "content": "<p>Perhaps the organizers could split these very large files into smaller pieces, maybe about 10 GB each, to help alleviate the problem of interrupted downloads.</p>",
          "rawMarkdown": "Perhaps the organizers could split these very large files into smaller pieces, maybe about 10 GB each, to help alleviate the problem of interrupted downloads."
        }
      ]
    },
    {
      "id": 425719,
      "postDate": "2018-11-22T02:29:33.053Z",
      "content": "<p>The note on data says \" The green filter should hence be used to predict the label, and the other filters are used as references\" If that is the case then in the test/predict samples should not be divided into different colors.  Can someone explain I am confused as the green one carries the information on the location of the protein.</p>",
      "rawMarkdown": "The note on data says \" The green filter should hence be used to predict the label, and the other filters are used as references\" If that is the case then in the test/predict samples should not be divided into different colors.  Can someone explain I am confused as the green one carries the information on the location of the protein.",
      "replies": [
        {
          "id": 426776,
          "postDate": "2018-11-23T20:20:18.600Z",
          "content": "<p>Hi Randhawp,\nFor every sample there are four images, also sometimes referred to as channels. The challenge is to predict the label of the protein image channel (i.e. green). For every sample there three other image channels show reference markers that outline certain structures of the cell (red - microtubules, blue - nucleus, yellow - endoplasmic reticulum). You can use these channels to get context about the cell for improved classification of the green label.\nBest,\nEmma</p>",
          "rawMarkdown": "Hi Randhawp,\nFor every sample there are four images, also sometimes referred to as channels. The challenge is to predict the label of the protein image channel (i.e. green). For every sample there three other image channels show reference markers that outline certain structures of the cell (red - microtubules, blue - nucleus, yellow - endoplasmic reticulum). You can use these channels to get context about the cell for improved classification of the green label.\nBest,\nEmma"
        }
      ]
    },
    {
      "id": 425056,
      "postDate": "2018-11-21T04:15:48.500Z",
      "content": "<p>Hello, \nIf we are doing this competition primarily for fun/learning, will it still be possible to submit entries for testing/test set labels revealed after the challenge has ended, just that they wouldn't count towards the competition?</p>",
      "rawMarkdown": "Hello, \nIf we are doing this competition primarily for fun/learning, will it still be possible to submit entries for testing/test set labels revealed after the challenge has ended, just that they wouldn't count towards the competition?",
      "replies": [
        {
          "id": 425229,
          "postDate": "2018-11-21T10:02:44.630Z",
          "content": "<p>Hi Kevin,\nI think this topic explains how it works with late submissions:\n<a href=\"https://www.kaggle.com/general/39808\">https://www.kaggle.com/general/39808</a></p>",
          "rawMarkdown": "Hi Kevin,\nI think this topic explains how it works with late submissions:\nhttps://www.kaggle.com/general/39808"
        },
        {
          "id": 425441,
          "postDate": "2018-11-21T15:56:32.607Z",
          "content": "<p>Thank you very much!</p>",
          "rawMarkdown": "Thank you very much!"
        }
      ]
    },
    {
      "id": 422943,
      "postDate": "2018-11-17T05:42:31.563Z",
      "content": "<p>Hi, Emma,</p>\n\n<p>Emma, I'd like to read the article but it appears to have a paywall. Is there a way around this? I'll pay for it but I'd prefer if there was a free version or some equivalent free version.</p>\n\n<p>Also, 0.71 was achieved by experts on this data. But who provided the ground truth for these experts? Other experts?</p>\n\n<p>Regards,\nPete</p>",
      "rawMarkdown": "Hi, Emma,\n\nEmma, I'd like to read the article but it appears to have a paywall. Is there a way around this? I'll pay for it but I'd prefer if there was a free version or some equivalent free version.\n\nAlso, 0.71 was achieved by experts on this data. But who provided the ground truth for these experts? Other experts?\n\nRegards,\nPete",
      "replies": [
        {
          "id": 422955,
          "postDate": "2018-11-17T06:09:46.740Z",
          "content": "<p>Hi Pete,</p>\n\n<p>You can access the article freely through the link at the bottom of this page <a href=\"https://www.proteinatlas.org/news/2018-08-20/mapping-of-cells-and-proteins-improved-with-combination-of-multiplayer-crowdsourcing-and-ai\">https://www.proteinatlas.org/news/2018-08-20/mapping-of-cells-and-proteins-improved-with-combination-of-multiplayer-crowdsourcing-and-ai</a></p>\n\n<p>The ground truth in the article is labels provided by three independent experts and curated based on several images of the same sample plus replicate samples in different cell lines (v14 of the Cell Atlas). The work described in this article led to the classification of additional labels that subsequently have been assessed, by the same experts in the same manner, and integrated to a newer release of the Human Protein Atlas (v18 of the Cell Atlas). For this article we tested experts individually for their performance on single images, which resulted in the macro-F1 of 0.71.</p>\n\n<p>The ground truth in this challenge is derived from the labels in v18 of the cell atlas. Note that the challenge consists of non-public images.</p>\n\n<p>Best,\nEmma</p>",
          "rawMarkdown": "Hi Pete,\n\nYou can access the article freely through the link at the bottom of this page https://www.proteinatlas.org/news/2018-08-20/mapping-of-cells-and-proteins-improved-with-combination-of-multiplayer-crowdsourcing-and-ai\n\nThe ground truth in the article is labels provided by three independent experts and curated based on several images of the same sample plus replicate samples in different cell lines (v14 of the Cell Atlas). The work described in this article led to the classification of additional labels that subsequently have been assessed, by the same experts in the same manner, and integrated to a newer release of the Human Protein Atlas (v18 of the Cell Atlas). For this article we tested experts individually for their performance on single images, which resulted in the macro-F1 of 0.71.\n\nThe ground truth in this challenge is derived from the labels in v18 of the cell atlas. Note that the challenge consists of non-public images.\n\nBest,\nEmma",
          "votes": 2
        },
        {
          "id": 423259,
          "postDate": "2018-11-17T19:51:30.503Z",
          "content": "<p>Got it! Thanks.</p>",
          "rawMarkdown": "Got it! Thanks."
        }
      ]
    },
    {
      "id": 421507,
      "postDate": "2018-11-15T04:44:33.277Z",
      "content": "<p>I've used 256x256 color images on my hardware before.</p>\n\n<p>I'm wondering what hardware stack would you need to flow 3072x3072 images through even a moderate size CNN? I can't even imagine.</p>",
      "rawMarkdown": "I've used 256x256 color images on my hardware before.\n\nI'm wondering what hardware stack would you need to flow 3072x3072 images through even a moderate size CNN? I can't even imagine."
    },
    {
      "id": 408161,
      "postDate": "2018-10-22T11:45:55.143Z",
      "content": "<blockquote>\n  <p>You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.</p>\n</blockquote>\n\n<p>Hi, is there an official thread about pre-trained model and outside dataset?</p>",
      "rawMarkdown": "&gt; You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.\n\nHi, is there an official thread about pre-trained model and outside dataset?",
      "replies": [
        {
          "id": 412186,
          "postDate": "2018-10-29T18:37:47.657Z",
          "content": "<p>Hi Hanke, now there is such an official thread. \n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>",
          "rawMarkdown": "Hi Hanke, now there is such an official thread. \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984"
        },
        {
          "id": 412564,
          "postDate": "2018-10-30T11:56:33.077Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 406524,
      "postDate": "2018-10-19T11:53:18.283Z",
      "content": "<p>Hi, I still have some trouble with the interpretation of what is being said about the filters. Is it correct to assume that only green filter files should be used for training and classification? </p>",
      "rawMarkdown": "Hi, I still have some trouble with the interpretation of what is being said about the filters. Is it correct to assume that only green filter files should be used for training and classification? ",
      "replies": [
        {
          "id": 406592,
          "postDate": "2018-10-19T13:37:25.850Z",
          "content": "<p>The image labels refers to what is being seen in the green filter. The other filters can be used as structural reference markers and can be used as input to your model should you want to.</p>",
          "rawMarkdown": "The image labels refers to what is being seen in the green filter. The other filters can be used as structural reference markers and can be used as input to your model should you want to."
        },
        {
          "id": 409080,
          "postDate": "2018-10-23T19:58:47.933Z",
          "content": "<p>There is no \"should\". </p>\n\n<p>It's up to us to analyse and decide which data is relevant or not. <br>\nIf you find out that yellow makes your models better, go for yellow. </p>",
          "rawMarkdown": "There is no \"should\". \n\nIt's up to us to analyse and decide which data is relevant or not.    \nIf you find out that yellow makes your models better, go for yellow. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 404308,
      "postDate": "2018-10-15T14:52:10.963Z",
      "content": "<p>What macro F-score human experts achieve when labeling proteins?</p>",
      "rawMarkdown": "What macro F-score human experts achieve when labeling proteins?",
      "replies": [
        {
          "id": 404327,
          "postDate": "2018-10-15T15:11:39.393Z",
          "content": "<p>Hi! The experts we have tested have performed at an average of macro F1 0.71. This was presented as part of this study <a href=\"https://www.nature.com/articles/nbt.4225\">https://www.nature.com/articles/nbt.4225</a>.</p>",
          "rawMarkdown": "Hi! The experts we have tested have performed at an average of macro F1 0.71. This was presented as part of this study https://www.nature.com/articles/nbt.4225.",
          "votes": 4
        },
        {
          "id": 404678,
          "postDate": "2018-10-16T07:38:57.247Z",
          "content": "<p>It will be interesting to see if anyone can surpass this result</p>",
          "rawMarkdown": "It will be interesting to see if anyone can surpass this result",
          "votes": 1
        },
        {
          "id": 405032,
          "postDate": "2018-10-16T18:58:50.940Z",
          "content": "<p>still some work to go it seems :) ...</p>",
          "rawMarkdown": "still some work to go it seems :) ..."
        },
        {
          "id": 406969,
          "postDate": "2018-10-20T05:18:57.893Z",
          "content": "<p>Emma - the experts you tested - their results are the ground truth for the train and test set?  I only did a quick read of the article you linked, I seem to read that a deep learning process and gamers resulted in the 0.71.  Was that deep learning process the ground truth for the train and test set?</p>",
          "rawMarkdown": "Emma - the experts you tested - their results are the ground truth for the train and test set?  I only did a quick read of the article you linked, I seem to read that a deep learning process and gamers resulted in the 0.71.  Was that deep learning process the ground truth for the train and test set?",
          "votes": 1
        },
        {
          "id": 410259,
          "postDate": "2018-10-25T17:34:04.903Z",
          "content": "<p>Hi PC Jimmmy!\nThe ground truth in the article is labels provided by three independent experts and curated based on several images of the same sample plus replicate samples in different cell lines (v14 of the Cell Atlas). The work described in this article led to the classification of additional labels that subsequently have been assessed, by the same experts in the same manner, and integrated to a newer release of the Human Protein Atlas (v18 of the Cell Atlas). For this article we tested experts individually for their performance on single images, which resulted in the macro-F1 of 0.71.</p>\n\n<p>The ground truth in this challenge is derived from the labels in v18 of the cell atlas. Note that the challenge consists of non-public images.</p>",
          "rawMarkdown": "Hi PC Jimmmy!\nThe ground truth in the article is labels provided by three independent experts and curated based on several images of the same sample plus replicate samples in different cell lines (v14 of the Cell Atlas). The work described in this article led to the classification of additional labels that subsequently have been assessed, by the same experts in the same manner, and integrated to a newer release of the Human Protein Atlas (v18 of the Cell Atlas). For this article we tested experts individually for their performance on single images, which resulted in the macro-F1 of 0.71.\n\nThe ground truth in this challenge is derived from the labels in v18 of the cell atlas. Note that the challenge consists of non-public images.",
          "votes": 3
        },
        {
          "id": 415796,
          "postDate": "2018-11-05T17:56:27.703Z",
          "content": "<p>It's not very clear whether this competition's dataset is \"perfect and some experts reach 0.71 of it\" or if this dataset is \"0.71\" F1 itself, meaning some samples are not quite right.</p>",
          "rawMarkdown": "It's not very clear whether this competition's dataset is \"perfect and some experts reach 0.71 of it\" or if this dataset is \"0.71\" F1 itself, meaning some samples are not quite right."
        },
        {
          "id": 420776,
          "postDate": "2018-11-14T05:18:08.120Z",
          "content": "<p>Emma, I'd like to read the article but it appears to have a paywall. Is there a way around this?</p>",
          "rawMarkdown": "Emma, I'd like to read the article but it appears to have a paywall. Is there a way around this?"
        },
        {
          "id": 452522,
          "postDate": "2019-01-08T21:08:32.677Z",
          "content": "<p>Yes there is! You can access the article freely through this link <a href=\"https://rdcu.be/4ReL\">Free full text</a></p>",
          "rawMarkdown": "Yes there is! You can access the article freely through this link [Free full text][1]\n\n  [1]: https://rdcu.be/4ReL \"link\""
        }
      ]
    },
    {
      "id": 401163,
      "postDate": "2018-10-09T14:35:53.973Z",
      "content": "<p>I'm sorry if someone already asked, are there any meanings in the order of Target?</p>",
      "rawMarkdown": "I'm sorry if someone already asked, are there any meanings in the order of Target?",
      "replies": [
        {
          "id": 401551,
          "postDate": "2018-10-10T09:01:35.807Z",
          "content": "<p>Hi! No, the order of labels in the Target column of the submission doesn't have any meaning.</p>",
          "rawMarkdown": "Hi! No, the order of labels in the Target column of the submission doesn't have any meaning.",
          "votes": 3
        },
        {
          "id": 401661,
          "postDate": "2018-10-10T13:00:56.360Z",
          "content": "<p>Thanks for your reply, Martin.</p>",
          "rawMarkdown": "Thanks for your reply, Martin."
        },
        {
          "id": 409073,
          "postDate": "2018-10-23T19:50:51.923Z",
          "content": "<p>Martin, looks like the order of the target ids matters in the submission. I confirmed it yesterday and others have too. Ref to the discussion here.  <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366</a></p>",
          "rawMarkdown": "Martin, looks like the order of the target ids matters in the submission. I confirmed it yesterday and others have too. Ref to the discussion here.  https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366",
          "votes": 4
        },
        {
          "id": 409075,
          "postDate": "2018-10-23T19:56:24.303Z",
          "content": "<p>Do not mistake \"order of the targets\" for \"order of the ids\". </p>",
          "rawMarkdown": "Do not mistake \"order of the targets\" for \"order of the ids\". \n\n"
        },
        {
          "id": 409147,
          "postDate": "2018-10-23T20:47:17.910Z",
          "content": "<p>You are right Daniel. Order of the Id's NOT the target predictions. </p>",
          "rawMarkdown": "You are right Daniel. Order of the Id's NOT the target predictions. "
        }
      ]
    },
    {
      "id": 400485,
      "postDate": "2018-10-08T11:54:41.253Z",
      "content": "<p>Hi, Emma, thank you for this competition!\nI do not understand what the meaning of green channel is? Is it mask made by people or another variant of cells photo with colorant? How to interpretate \"the protein of interest\"?</p>",
      "rawMarkdown": "Hi, Emma, thank you for this competition!\nI do not understand what the meaning of green channel is? Is it mask made by people or another variant of cells photo with colorant? How to interpretate \"the protein of interest\"?",
      "replies": [
        {
          "id": 400511,
          "postDate": "2018-10-08T12:44:40.503Z",
          "content": "<p>Hello!</p>\n\n<p>Protein of interest in this context refers to the protein that we are investigating in that experiment. For our experiments the protein of interest is the protein that we are trying to localize. The protein is stained using targeted antibodies and imaged using immunoflourescence microscopy. The resulting staining is  the green channel of each image.\nThe labels that are associated with each image refers to which pattern can be seen within the green channel.</p>\n\n<p>The other three channels are targeting specific cellular structures and can be used for reference when annotating to make it easier to identify the spatial orientation of the cells.</p>",
          "rawMarkdown": "Hello!\n\nProtein of interest in this context refers to the protein that we are investigating in that experiment. For our experiments the protein of interest is the protein that we are trying to localize. The protein is stained using targeted antibodies and imaged using immunoflourescence microscopy. The resulting staining is  the green channel of each image.\nThe labels that are associated with each image refers to which pattern can be seen within the green channel.\n\nThe other three channels are targeting specific cellular structures and can be used for reference when annotating to make it easier to identify the spatial orientation of the cells.",
          "votes": 2
        },
        {
          "id": 400521,
          "postDate": "2018-10-08T13:08:00.703Z",
          "content": "<p>Thank you for your answer, so people set up only the labels?</p>",
          "rawMarkdown": "Thank you for your answer, so people set up only the labels?",
          "votes": 1
        },
        {
          "id": 400543,
          "postDate": "2018-10-08T14:12:51.723Z",
          "content": "<p>Exactly! The images are manually acquired and annotated but the patterns themselves are from the antibodies.</p>",
          "rawMarkdown": "Exactly! The images are manually acquired and annotated but the patterns themselves are from the antibodies.",
          "votes": 1
        }
      ]
    },
    {
      "id": 399993,
      "postDate": "2018-10-07T10:11:12.397Z",
      "content": "<p>I have some trouble understanding what we should do : how is it possible that \"microtubules\" is both a channel (so present in all images) and a label (present or not in images)... </p>",
      "rawMarkdown": "I have some trouble understanding what we should do : how is it possible that \"microtubules\" is both a channel (so present in all images) and a label (present or not in images)... ",
      "replies": [
        {
          "id": 400412,
          "postDate": "2018-10-08T09:17:42.300Z",
          "content": "<p>In our experiment protocol one of the reference markers we use stain the microtubules. So one of the filters of a sample (red) will always show the microtubules. The green filter, the protein of interest, varies between samples though. This will show different organelles in different samples. Sometimes the protein of interest might be located in the microtubules and the green filter will show this. So in those cases there is usually a perfect overlap between the green and the red filter patterns.</p>",
          "rawMarkdown": "In our experiment protocol one of the reference markers we use stain the microtubules. So one of the filters of a sample (red) will always show the microtubules. The green filter, the protein of interest, varies between samples though. This will show different organelles in different samples. Sometimes the protein of interest might be located in the microtubules and the green filter will show this. So in those cases there is usually a perfect overlap between the green and the red filter patterns."
        }
      ]
    },
    {
      "id": 399778,
      "postDate": "2018-10-06T18:35:01.080Z",
      "content": "<p>nyc</p>",
      "rawMarkdown": "nyc"
    },
    {
      "id": 399200,
      "postDate": "2018-10-05T11:57:13.450Z",
      "content": "<p>Hi Emma</p>\n\n<p>Can I just check I understand the problem correctly. </p>\n\n<p>We are predicting where, in a particular cell type,  a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right? Can we assume the proportions of proteins, cell types and labels in the datasets are typical?</p>\n\n<p>Perhaps the problem is obvious from the dataset but the zip file seems to be corrupted. I've tried downloading it twice.</p>",
      "rawMarkdown": "Hi Emma\n\nCan I just check I understand the problem correctly. \n\nWe are predicting where, in a particular cell type,  a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right? Can we assume the proportions of proteins, cell types and labels in the datasets are typical?\n\nPerhaps the problem is obvious from the dataset but the zip file seems to be corrupted. I've tried downloading it twice.",
      "replies": [
        {
          "id": 399247,
          "postDate": "2018-10-05T13:33:14.633Z",
          "content": "<p>Hi!</p>\n\n<blockquote>\n  <p>We are predicting where, in a particular cell type, a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right?</p>\n</blockquote>\n\n<p>Basically yes. The degree of variation in morphology between different cell types will vary though. Some cell types are more similar to each other than others.</p>\n\n<p>I won't say the proportions of the datasets. You can count the occurrence of the labels in the training set to get an idea about the label proportions.</p>\n\n<p>I'm sorry to hear about the trouble with the zip file. I tested downloading the test.zip archive earlier today and that worked. Is anyone else having issues with downloads?</p>",
          "rawMarkdown": "Hi!\n\n&gt; We are predicting where, in a particular cell type, a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right?\n\nBasically yes. The degree of variation in morphology between different cell types will vary though. Some cell types are more similar to each other than others.\n\nI won't say the proportions of the datasets. You can count the occurrence of the labels in the training set to get an idea about the label proportions.\n\nI'm sorry to hear about the trouble with the zip file. I tested downloading the test.zip archive earlier today and that worked. Is anyone else having issues with downloads?"
        },
        {
          "id": 399309,
          "postDate": "2018-10-05T15:19:21.953Z",
          "content": "<p>Thanks Martin. I was using the 'download all' link. I shall try the test.zip file instead.</p>",
          "rawMarkdown": "Thanks Martin. I was using the 'download all' link. I shall try the test.zip file instead."
        }
      ]
    },
    {
      "id": 399196,
      "postDate": "2018-10-05T11:50:18.977Z",
      "content": "<p>Now, it is very clear</p>",
      "rawMarkdown": "Now, it is very clear"
    },
    {
      "id": 398704,
      "postDate": "2018-10-04T13:23:38.437Z",
      "content": "<p>Hi Emma,\nThank you for hosting this competition. It is a very interesting domain to apply Machine Learning.</p>\n\n<p>I'm assuming the CPU + 4Gb RAM restriction has to do with integrating our models in a \"smart-microscopy system\" as said in the Overview page of this competition. Am I right?</p>\n\n<p>Since this isn't a Kaggle Kernel CPU only competition, I'd like to understand from you and from Kaggle staff if, for this competition, we can have two separate models:</p>\n\n<ul>\n<li>one for top accuracy, using GPU hardware which is standard for image problems. This one competes for the normal money prize</li>\n<li>another one for fast prediction in CPU but lower accuracy, competing for the special prize</li>\n</ul>\n\n<p>If that's not possible, we will forcibly have to choose what to model for: best accuracy/fastest CPU model - normal prize/special prize.</p>\n\n<p>Looking forward for news from you.\nThanks</p>",
      "rawMarkdown": "Hi Emma,\nThank you for hosting this competition. It is a very interesting domain to apply Machine Learning.\n\nI'm assuming the CPU + 4Gb RAM restriction has to do with integrating our models in a \"smart-microscopy system\" as said in the Overview page of this competition. Am I right?\n\nSince this isn't a Kaggle Kernel CPU only competition, I'd like to understand from you and from Kaggle staff if, for this competition, we can have two separate models:\n\n - one for top accuracy, using GPU hardware which is standard for image problems. This one competes for the normal money prize\n - another one for fast prediction in CPU but lower accuracy, competing for the special prize\n\nIf that's not possible, we will forcibly have to choose what to model for: best accuracy/fastest CPU model - normal prize/special prize.\n\nLooking forward for news from you.\nThanks",
      "replies": [
        {
          "id": 398796,
          "postDate": "2018-10-04T15:37:53.630Z",
          "content": "<p>Hi Renan, thanks for your feedback!</p>\n\n<p>Yes, we have put the limits on the hardware to match our smart-microscopy system workflow.</p>\n\n<p>In the rules it says:</p>\n\n<blockquote>\n  <p>You may select up to 2 final submissions for judging.</p>\n</blockquote>\n\n<p>This should allow you to compete with two separate models. We're excited to see what different approaches the community will take.</p>\n\n<p>Martin, part of the Human Protein Atlas team</p>",
          "rawMarkdown": "Hi Renan, thanks for your feedback!\n\nYes, we have put the limits on the hardware to match our smart-microscopy system workflow.\n\nIn the rules it says:\n&gt; You may select up to 2 final submissions for judging.\n\nThis should allow you to compete with two separate models. We're excited to see what different approaches the community will take.\n\nMartin, part of the Human Protein Atlas team",
          "votes": 1
        }
      ]
    },
    {
      "id": 424265,
      "postDate": "2018-11-19T20:15:29.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 415813,
      "postDate": "2018-11-05T18:34:55.757Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 399442,
      "postDate": "2018-10-05T19:35:06.563Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 399488,
          "postDate": "2018-10-05T22:03:06.763Z",
          "content": "<p>Hi Rian,</p>\n\n<p>Submissions will be evaluated based on their macro F1 score. This means that if a sample has multiple labels, all of them need to be predicted for maximum score.</p>",
          "rawMarkdown": "Hi Rian,\n\nSubmissions will be evaluated based on their macro F1 score. This means that if a sample has multiple labels, all of them need to be predicted for maximum score.\n\n"
        },
        {
          "id": 415023,
          "postDate": "2018-11-04T07:24:38.720Z",
          "content": "<p>So when predicting the test samples for submission there can be multiple labels for each sample?</p>",
          "rawMarkdown": "So when predicting the test samples for submission there can be multiple labels for each sample?"
        },
        {
          "id": 415288,
          "postDate": "2018-11-04T20:21:11.847Z",
          "content": "<p>Yes, there can be multiple labels for each sample.</p>",
          "rawMarkdown": "Yes, there can be multiple labels for each sample.",
          "votes": 1
        },
        {
          "id": 463084,
          "postDate": "2019-01-29T11:27:03.270Z",
          "content": "<p>yes why not emma</p>",
          "rawMarkdown": "yes why not emma",
          "votes": 1
        }
      ]
    },
    {
      "id": 399415,
      "postDate": "2018-10-05T18:30:41.077Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 401999,
          "postDate": "2018-10-11T02:51:25.897Z",
          "content": "<p>Thank you Simon! Great to hear that we have frequent Atlas users willing to challenge the computational interpretation of the images.</p>",
          "rawMarkdown": "Thank you Simon! Great to hear that we have frequent Atlas users willing to challenge the computational interpretation of the images."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 432897,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2018-12-04T12:48:55",
      "content": "<p>Are kaggle or the organisers going to comment on the data leak?</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 434648,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-12-06T18:17:24.997000",
          "content": "<p>Yes, <a href=\"/philculliton\">@philculliton</a> did answer in the thread.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423743,
      "author_name": "Shalin",
      "author_url": "",
      "post_date": "2018-11-19T01:11:38.620000",
      "content": "<p>In the data set there're many classes has co-occurence with the other classes. For example, One image is labeled as  both Nucleoplasm and Cytosol. So, how to understand by looking at images that why it is labeled to the both classes? If I see an image only labeled with the Nucleoplasm or Cytosol by combining all 4 channels, I can't find any visual similarities. If I see an image classified with the both classes then what should I look into so it can explain that the image is belong to both classes. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 423803,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2018-11-19T04:51:02.790000",
          "content": "<p>I literary tried that yesterday. From my understanding, Cytosol is the same as Cytoplasm, but not the same as Nucleoplasm. Unfortunately, my current model can easily surpass my poor classification. I cannot imagine how EVE players can classify images so accurately with such noisy data. Some green labels can hardly be seen by human eyes, some (0a022f5e-bbca-11e8-b2bc-ac1f6b6435d0) are corrupted, and some do not even have a cell-like structure (0c1d7284-bbc1-11e8-b2bb-ac1f6b6435d0).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433422,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-12-05T04:02:04.937000",
          "content": "<p>It is often easier to find the visual simalarities/differences when not combining all four channels, in particular the yellow is hard to distinguish from the green. The dictionary in the HPA can help you understand the different patterns <a href=\"https://www.proteinatlas.org/learn/dictionary/cell\">https://www.proteinatlas.org/learn/dictionary/cell</a>.</p>\n\n<p>For the example of cytosol, nucleoplasm vs cytosol+nucleoplasm you can see the differences by looking at the overlap with the 'red' and 'blue' markers respectively. The blue is a marker for the nucleus, so if there is green staining within the blue area, it should have the label ‘nucleoplasm’. The red is microtubules that branch out through the cytosol, so if there is green staining within the red  area, it should have the label ‘cytosol’. If there is green staining in both the red and green area, it should have both labels.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 419779,
      "author_name": "Basti",
      "author_url": "",
      "post_date": "2018-11-12T14:51:33.037000",
      "content": "<p>I found that there are also 4096x4096 px images in the training dataset. 7 to be exact.\nId's:\n0858d008-bb9e-11e8-b2b9-ac1f6b6435d0\n16196e56-bba2-11e8-b2b9-ac1f6b6435d0\n1625b570-bbc6-11e8-b2bc-ac1f6b6435d0\n28be7d5e-bbbf-11e8-b2ba-ac1f6b6435d0\n558dfed4-bba3-11e8-b2b9-ac1f6b6435d0\nc3322376-bbc0-11e8-b2bb-ac1f6b6435d0\nf2998bbe-bbab-11e8-b2ba-ac1f6b6435d0</p>\n\n<p>This lead to a bug that cost me some calculation and debugging time. Consider adding this information to your Post above and the competition description where you list the image sizes.</p>\n\n<p>Cheers,\nBasti</p>",
      "votes": 1,
      "replies": [
        {
          "id": 435285,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2018-12-07T19:50:58.940000",
          "content": "<p>Hi Basti,</p>\n\n<p>Thanks for bringing this to our attention. We've analysed the shape of the complete data set now and updated the original post above with csv files of all tiff files with 4096x4096 px.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 414929,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2018-11-03T23:10:13.027000",
      "content": "<p>I can't upload and submit data in these two days. I tried almost all methods and couldn't solve them. What should I do?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 411984,
      "author_name": "tulasi",
      "author_url": "",
      "post_date": "2018-10-29T10:51:45.017000",
      "content": "<p>Hi Emma,\nWhile there exist most of the works done for protein subcellular classification using textual data, why we have to do same with images. How cost effective is images compare to text data prediction. What is the intention behind giving these image dataset for the same classification which is done using textual data already? As it will be questionable by the reviewer when a paper is submitted for publication. Any response would be appreciable :) Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 422960,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-11-17T06:21:24.163000",
          "content": "<p>Hi Tulasi,</p>\n\n<p>Predicting protein subcellular localization from textual data (here I am assuming that you mean the amino acid sequence of the protein, or nucleic acid sequence of the gene encoding the protein) is nowhere near a solved problem, unfortunately. It is possible to reliably predict localization with signal peptides directing them to certain places (mitochondria, nuclear localization signals etc) and transmembrane regions. But there are no methods (to my knowledge) that can predict anywhere near the number of locations and substructures included in this challenge (for example distinguish between if a protein is in the nucleoli or in the nucleoli fibrillar center). Prediction of multiple locations from sequences is also a very hard task, most likely because this cellular phenomena is influenced by post-translational regulation and interactions.</p>\n\n<p>Best,\nEmma</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 399804,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2018-10-06T20:07:54.690000",
      "content": "<p>Looking forward to participating! Do you think it is possible to get top performing results using the PNG set or would this only be possible by training on the TIFF set?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 400450,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2018-10-08T10:47:59.320000",
          "content": "<p>Hello Carlo!</p>\n\n<p>The TIFFs should in theory contain more information than the PNGs which are downscaled to 512x512px from the same TIFFs. We would expect that the PNGs are of good enough quality to get a top performing result but we do not want to promise anything. </p>\n\n<p>Looking at the PNGs they are at least good enough for human annotation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 406966,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2018-10-20T05:12:55.103000",
          "content": "<p>The special prize is interesting, would love that GPU.  But I see no chance that I will try for it.\n1.  The level of detail and work to get the special prize submission is way too much for a hobby data scientist.\n2.  The special prize requires the prediction to use the TIFF files - OK that makes some sense, but the TIFF are dumbed down, so why bother.  In addition, I would need to use the TIFF to make sure things work - that would mean for me that I should go with TIFF now, cause stuff happens.\n3.  Without any way to benchmark speed, I am not sure I want to put in all that work and know have some idea of my speed compared to others.  Any potential way there can be a speed leaderboard?</p>\n\n<p>Do like the regular challenge and look forward to learning all about proteins. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 398935,
      "author_name": "Renan Rodrigues dos Santos",
      "author_url": "",
      "post_date": "2018-10-04T20:29:25.887000",
      "content": "<p>Thanks for your reply, Martin.\nIt is very clear now.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 401008,
      "author_name": "Martin Hjelmare",
      "author_url": "",
      "post_date": "2018-10-09T09:11:02.100000",
      "content": "<p>Following up on the questions about the hardware limits of the special prize evaluation, we want to clarify that we only limit the hardware of the prediction process that we will perform when evaluating submissions. Training can still be done without hardware restrictions, since this is done before submission.</p>\n\n<p>We understand that the hardware limits can still impact top performance, but we are interested to find models that balance performance, speed and necessary hardware resources. The reason why we have put the limits on the hardware is to match our smart-microscopy system workflow, as mentioned above.</p>\n\n<p>Thanks for the feedback and please continue the discussion on this topic. We want to hear what you think.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 411758,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-10-29T00:31:04.493000",
          "content": "<p>Martin,</p>\n\n<p>On the requirements for the special prize it mentions that an entrypoint is needed for training in the docker image. Does this mean that you want included a way to train completely from scratch with default initialization? Or should the training session load pretrained weights? Are there any size limits to the zip? My saved model and weights is less than 10MB, but the 512x512 input I am using is over 75GB.</p>\n\n<p>The model I'm working on takes a long time to train, but predicts quickly. Even with a GPU it currently takes days of training. In development I'm working using Jupyter fine tuning the model and training sessions, trying to take the best notes I can but it may be difficult to reproduce exactly.</p>\n\n<p>For the prediction performance, it is pretty straightforward. I am using Keras so it would pretty much look like:</p>\n\n<ul>\n<li>define model functions, loading functions for the input path</li>\n<li>load hdf5 model with saved weights from the model path</li>\n<li>predict and save submission to output path</li>\n</ul>\n\n<p>Thanks,</p>\n\n<p>Brian</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 449329,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2019-01-03T01:22:29.663000",
          "content": "<p>Hi Brian,</p>\n\n<p>does my answer here answer your questions?\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67604#449328\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67604#449328</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450462,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-01-05T02:18:38.030000",
          "content": "<p>It is still a little opaque:</p>\n\n<pre><code>- the last link points to this very discussion. Are you referring to the message above the link?\n- \"...the readme file in the submission...\" Where are the requirements on this file described? I don't see it anywhere. Is it in the \"submission instructions\". Can you give a link to these?\n- I am not 100% sure but I doubt that anyone can reproduce their exact submitted predictions with a new run of the code. There are random elements in all packages that do not seem reproducible. I have tried in fastai and was not able to reproduce things exactly. This has been discussed in the threads.\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452309,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2019-01-08T14:51:32.213000",
          "content": "<p>Hi Pete,</p>\n\n<ul>\n<li>The last link should be a permalink to my answer to Maksim further down in this thread.</li>\n<li>Here is a link to the special prize submission intructions:\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions\">https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions</a></li>\n<li>In the instructions we write that the submitted model should maintain its score +/- 1 percentage point. Are those limits within reason for a reproduced model too?</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 452363,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2019-01-08T16:14:15.420000",
          "content": "<p>Hi Martin,</p>\n\n<p>Where is the <code>submission_sample.zip</code> file?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452459,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-01-08T18:52:18.313000",
          "content": "<p>For me, the chances are remote (but not yet strictly zero) of being in the money. But I would say that 1% is maybe a bit small - it is 0.6 vs 0.606. With variations in the randoms, which appear unavoidable, one can get this type of variation. I guess it would be OK if you allow 2-3 tries to get this number.</p>\n\n<p>Regarding the instructions : this seems to apply only to the \"special prize\". Do they in fact apply to the main contest submission as well, which I assume is independent of the special prize?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 463904,
          "author_name": "genethes",
          "author_url": "",
          "post_date": "2019-01-30T21:27:26.703000",
          "content": "<p>Hi Martin,</p>\n\n<p>Will late submissions be enabled for this competition? Wishing to review and improve my solution.</p>\n\n<p>Even when submitting an older submission, I get:\n\"Your submission was unsuccessful. Please try again, and if you continue to have issues, please let us know the details of your submission and competition on the Product Feedback forum.\"</p>\n\n<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 466257,
          "author_name": "Elizabeth Park",
          "author_url": "",
          "post_date": "2019-02-05T01:07:15.657000",
          "content": "<p>Hello, </p>\n\n<p>Late Submissions should be enabled for this competition already. We've run a test and were able to successfully submit a sample CSV file to the Late Submissions page. Could you please try submitting again and let us know if the issue persists? Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 398314,
      "author_name": "Miha Skalic",
      "author_url": "",
      "post_date": "2018-10-03T22:34:02.707000",
      "content": "<blockquote>\n  <p>Top performing teams are invited to make the model code and any external data used to generate their final submission available to ...</p>\n</blockquote>\n\n<p>_</p>\n\n<blockquote>\n  <p>The winner will be the fastest model that maintains its submitted F1 score in the main prize competition +/- 1 percentage point.</p>\n</blockquote>\n\n<p>That is a bit too optimistic, isn't it? Top submissions will need significantly more FLOPS than what 2 CPU cores can produce.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 398533,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2018-10-04T08:38:06.097000",
          "content": "<p>Seems this approach is also penalizing better performing models (as the performance needs to be carried over to the special prize category and it is likely that better performing models will require more compute).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413479,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-01T02:48:27.903000",
          "content": "<p>once you have top performance model, you can use neural architecture search or knowledge distillation method to find smaller model with similar accuracy. Good example is nasnet and mobile-nasnet and there are such tools in tensorflow.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 413955,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-01T21:26:54.480000",
          "content": "<p>I focused from the beginning on a small model. So far using CPU only it can predict one 512x512 image in ~850ms. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 514526,
      "author_name": "Jeppe",
      "author_url": "",
      "post_date": "2019-04-11T17:42:39.417000",
      "content": "<p>Hi,\nCan you share how you converted 16-bit tiff files to the 8-bit images?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 443203,
      "author_name": "Maksim Rodin",
      "author_url": "",
      "post_date": "2018-12-21T07:57:16.593000",
      "content": "<p>Hello Emma,\nCould you please explain the rules concerning reproducibility of the final submission:\nShould the whole training process be completely reproducible or only the inference part?\nThank you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 449328,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2019-01-03T01:19:23.867000",
          "content": "<p>Hi Maksim,</p>\n\n<p>The training process should be described in the readme file in the submission and there should be entry points that prepares the data, starts the training etc as described in the submission instructions. We should be able to train a model from the instructions and reproduce your prediction result on the test set.</p>\n\n<p>If external data or pre-trained models have been used, that should either be included in the submission, linked to, or made available to us upon request. You also need to make a post describing what external data or pre-trainined models have been used in the corresponding official discussion thread here in the forum.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 431055,
      "author_name": "BenChur",
      "author_url": "",
      "post_date": "2018-12-01T14:41:40.787000",
      "content": "<p>Hi, Emma Lundberg.\nI have a question. If there have colorful sample image, no like these grey?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 431370,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2018-12-02T05:52:03.723000",
          "content": "<p>I think we get to color between the lines :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427150,
      "author_name": "Maksim Rodin",
      "author_url": "",
      "post_date": "2018-11-24T18:46:58.733000",
      "content": "<p>Hi Emma,</p>\n\n<p>I have a problem with downloading tiff images.\nWhen I download from google bucket via browser (tried Chrome and Firefox), anywhere in the middle there is some connection error (maybe short disconnect or smth like that), and downloading interrupts, and browsers are not clever enough to continue...</p>\n\n<p>I wonder if there is any other fool-proof possibility of downloading tiff images (e.g. torrent file)?</p>\n\n<p>Did anybody have a similar problem maybe?</p>\n\n<p>Regards,\nMaksim</p>",
      "votes": 0,
      "replies": [
        {
          "id": 427229,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2018-11-24T22:24:19.783000",
          "content": "<p>Perhaps the organizers could split these very large files into smaller pieces, maybe about 10 GB each, to help alleviate the problem of interrupted downloads.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 425719,
      "author_name": "randhawp",
      "author_url": "",
      "post_date": "2018-11-22T02:29:33.053000",
      "content": "<p>The note on data says \" The green filter should hence be used to predict the label, and the other filters are used as references\" If that is the case then in the test/predict samples should not be divided into different colors.  Can someone explain I am confused as the green one carries the information on the location of the protein.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 426776,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-11-23T20:20:18.600000",
          "content": "<p>Hi Randhawp,\nFor every sample there are four images, also sometimes referred to as channels. The challenge is to predict the label of the protein image channel (i.e. green). For every sample there three other image channels show reference markers that outline certain structures of the cell (red - microtubules, blue - nucleus, yellow - endoplasmic reticulum). You can use these channels to get context about the cell for improved classification of the green label.\nBest,\nEmma</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 425056,
      "author_name": "Kevin Hu",
      "author_url": "",
      "post_date": "2018-11-21T04:15:48.500000",
      "content": "<p>Hello, \nIf we are doing this competition primarily for fun/learning, will it still be possible to submit entries for testing/test set labels revealed after the challenge has ended, just that they wouldn't count towards the competition?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 425229,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2018-11-21T10:02:44.630000",
          "content": "<p>Hi Kevin,\nI think this topic explains how it works with late submissions:\n<a href=\"https://www.kaggle.com/general/39808\">https://www.kaggle.com/general/39808</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425441,
          "author_name": "Kevin Hu",
          "author_url": "",
          "post_date": "2018-11-21T15:56:32.607000",
          "content": "<p>Thank you very much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422943,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-11-17T05:42:31.563000",
      "content": "<p>Hi, Emma,</p>\n\n<p>Emma, I'd like to read the article but it appears to have a paywall. Is there a way around this? I'll pay for it but I'd prefer if there was a free version or some equivalent free version.</p>\n\n<p>Also, 0.71 was achieved by experts on this data. But who provided the ground truth for these experts? Other experts?</p>\n\n<p>Regards,\nPete</p>",
      "votes": 0,
      "replies": [
        {
          "id": 422955,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-11-17T06:09:46.740000",
          "content": "<p>Hi Pete,</p>\n\n<p>You can access the article freely through the link at the bottom of this page <a href=\"https://www.proteinatlas.org/news/2018-08-20/mapping-of-cells-and-proteins-improved-with-combination-of-multiplayer-crowdsourcing-and-ai\">https://www.proteinatlas.org/news/2018-08-20/mapping-of-cells-and-proteins-improved-with-combination-of-multiplayer-crowdsourcing-and-ai</a></p>\n\n<p>The ground truth in the article is labels provided by three independent experts and curated based on several images of the same sample plus replicate samples in different cell lines (v14 of the Cell Atlas). The work described in this article led to the classification of additional labels that subsequently have been assessed, by the same experts in the same manner, and integrated to a newer release of the Human Protein Atlas (v18 of the Cell Atlas). For this article we tested experts individually for their performance on single images, which resulted in the macro-F1 of 0.71.</p>\n\n<p>The ground truth in this challenge is derived from the labels in v18 of the cell atlas. Note that the challenge consists of non-public images.</p>\n\n<p>Best,\nEmma</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423259,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-11-17T19:51:30.503000",
          "content": "<p>Got it! Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421507,
      "author_name": "Anthony",
      "author_url": "",
      "post_date": "2018-11-15T04:44:33.277000",
      "content": "<p>I've used 256x256 color images on my hardware before.</p>\n\n<p>I'm wondering what hardware stack would you need to flow 3072x3072 images through even a moderate size CNN? I can't even imagine.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 408161,
      "author_name": "Hanke Chen",
      "author_url": "",
      "post_date": "2018-10-22T11:45:55.143000",
      "content": "<blockquote>\n  <p>You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.</p>\n</blockquote>\n\n<p>Hi, is there an official thread about pre-trained model and outside dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 412186,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-10-29T18:37:47.657000",
          "content": "<p>Hi Hanke, now there is such an official thread. \n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 412564,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2018-10-30T11:56:33.077000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 406524,
      "author_name": "Maciej Rybinski",
      "author_url": "",
      "post_date": "2018-10-19T11:53:18.283000",
      "content": "<p>Hi, I still have some trouble with the interpretation of what is being said about the filters. Is it correct to assume that only green filter files should be used for training and classification? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 406592,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2018-10-19T13:37:25.850000",
          "content": "<p>The image labels refers to what is being seen in the green filter. The other filters can be used as structural reference markers and can be used as input to your model should you want to.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409080,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-10-23T19:58:47.933000",
          "content": "<p>There is no \"should\". </p>\n\n<p>It's up to us to analyse and decide which data is relevant or not. <br>\nIf you find out that yellow makes your models better, go for yellow. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 404308,
      "author_name": "pi-null-mezon",
      "author_url": "",
      "post_date": "2018-10-15T14:52:10.963000",
      "content": "<p>What macro F-score human experts achieve when labeling proteins?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 404327,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-10-15T15:11:39.393000",
          "content": "<p>Hi! The experts we have tested have performed at an average of macro F1 0.71. This was presented as part of this study <a href=\"https://www.nature.com/articles/nbt.4225\">https://www.nature.com/articles/nbt.4225</a>.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 404678,
          "author_name": "pi-null-mezon",
          "author_url": "",
          "post_date": "2018-10-16T07:38:57.247000",
          "content": "<p>It will be interesting to see if anyone can surpass this result</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 405032,
          "author_name": "Thomas Bourgeois",
          "author_url": "",
          "post_date": "2018-10-16T18:58:50.940000",
          "content": "<p>still some work to go it seems :) ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 406969,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2018-10-20T05:18:57.893000",
          "content": "<p>Emma - the experts you tested - their results are the ground truth for the train and test set?  I only did a quick read of the article you linked, I seem to read that a deep learning process and gamers resulted in the 0.71.  Was that deep learning process the ground truth for the train and test set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 410259,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2018-10-25T17:34:04.903000",
          "content": "<p>Hi PC Jimmmy!\nThe ground truth in the article is labels provided by three independent experts and curated based on several images of the same sample plus replicate samples in different cell lines (v14 of the Cell Atlas). The work described in this article led to the classification of additional labels that subsequently have been assessed, by the same experts in the same manner, and integrated to a newer release of the Human Protein Atlas (v18 of the Cell Atlas). For this article we tested experts individually for their performance on single images, which resulted in the macro-F1 of 0.71.</p>\n\n<p>The ground truth in this challenge is derived from the labels in v18 of the cell atlas. Note that the challenge consists of non-public images.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 415796,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-05T17:56:27.703000",
          "content": "<p>It's not very clear whether this competition's dataset is \"perfect and some experts reach 0.71 of it\" or if this dataset is \"0.71\" F1 itself, meaning some samples are not quite right.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420776,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-11-14T05:18:08.120000",
          "content": "<p>Emma, I'd like to read the article but it appears to have a paywall. Is there a way around this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452522,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2019-01-08T21:08:32.677000",
          "content": "<p>Yes there is! You can access the article freely through this link <a href=\"https://rdcu.be/4ReL\">Free full text</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 401163,
      "author_name": "nor",
      "author_url": "",
      "post_date": "2018-10-09T14:35:53.973000",
      "content": "<p>I'm sorry if someone already asked, are there any meanings in the order of Target?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 401551,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2018-10-10T09:01:35.807000",
          "content": "<p>Hi! No, the order of labels in the Target column of the submission doesn't have any meaning.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 401661,
          "author_name": "nor",
          "author_url": "",
          "post_date": "2018-10-10T13:00:56.360000",
          "content": "<p>Thanks for your reply, Martin.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409073,
          "author_name": "Venky Krishnamani",
          "author_url": "",
          "post_date": "2018-10-23T19:50:51.923000",
          "content": "<p>Martin, looks like the order of the target ids matters in the submission. I confirmed it yesterday and others have too. Ref to the discussion here.  <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 409075,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-10-23T19:56:24.303000",
          "content": "<p>Do not mistake \"order of the targets\" for \"order of the ids\". </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409147,
          "author_name": "Venky Krishnamani",
          "author_url": "",
          "post_date": "2018-10-23T20:47:17.910000",
          "content": "<p>You are right Daniel. Order of the Id's NOT the target predictions. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 400485,
      "author_name": "Aleksey Alekseev",
      "author_url": "",
      "post_date": "2018-10-08T11:54:41.253000",
      "content": "<p>Hi, Emma, thank you for this competition!\nI do not understand what the meaning of green channel is? Is it mask made by people or another variant of cells photo with colorant? How to interpretate \"the protein of interest\"?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 400511,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2018-10-08T12:44:40.503000",
          "content": "<p>Hello!</p>\n\n<p>Protein of interest in this context refers to the protein that we are investigating in that experiment. For our experiments the protein of interest is the protein that we are trying to localize. The protein is stained using targeted antibodies and imaged using immunoflourescence microscopy. The resulting staining is  the green channel of each image.\nThe labels that are associated with each image refers to which pattern can be seen within the green channel.</p>\n\n<p>The other three channels are targeting specific cellular structures and can be used for reference when annotating to make it easier to identify the spatial orientation of the cells.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 400521,
          "author_name": "Aleksey Alekseev",
          "author_url": "",
          "post_date": "2018-10-08T13:08:00.703000",
          "content": "<p>Thank you for your answer, so people set up only the labels?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 400543,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2018-10-08T14:12:51.723000",
          "content": "<p>Exactly! The images are manually acquired and annotated but the patterns themselves are from the antibodies.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 399993,
      "author_name": "Thomas Bourgeois",
      "author_url": "",
      "post_date": "2018-10-07T10:11:12.397000",
      "content": "<p>I have some trouble understanding what we should do : how is it possible that \"microtubules\" is both a channel (so present in all images) and a label (present or not in images)... </p>",
      "votes": 0,
      "replies": [
        {
          "id": 400412,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2018-10-08T09:17:42.300000",
          "content": "<p>In our experiment protocol one of the reference markers we use stain the microtubules. So one of the filters of a sample (red) will always show the microtubules. The green filter, the protein of interest, varies between samples though. This will show different organelles in different samples. Sometimes the protein of interest might be located in the microtubules and the green filter will show this. So in those cases there is usually a perfect overlap between the green and the red filter patterns.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 399778,
      "author_name": "Digbijay Panda",
      "author_url": "",
      "post_date": "2018-10-06T18:35:01.080000",
      "content": "<p>nyc</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 399200,
      "author_name": "Cugel",
      "author_url": "",
      "post_date": "2018-10-05T11:57:13.450000",
      "content": "<p>Hi Emma</p>\n\n<p>Can I just check I understand the problem correctly. </p>\n\n<p>We are predicting where, in a particular cell type,  a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right? Can we assume the proportions of proteins, cell types and labels in the datasets are typical?</p>\n\n<p>Perhaps the problem is obvious from the dataset but the zip file seems to be corrupted. I've tried downloading it twice.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 399247,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2018-10-05T13:33:14.633000",
          "content": "<p>Hi!</p>\n\n<blockquote>\n  <p>We are predicting where, in a particular cell type, a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right?</p>\n</blockquote>\n\n<p>Basically yes. The degree of variation in morphology between different cell types will vary though. Some cell types are more similar to each other than others.</p>\n\n<p>I won't say the proportions of the datasets. You can count the occurrence of the labels in the training set to get an idea about the label proportions.</p>\n\n<p>I'm sorry to hear about the trouble with the zip file. I tested downloading the test.zip archive earlier today and that worked. Is anyone else having issues with downloads?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 399309,
          "author_name": "Cugel",
          "author_url": "",
          "post_date": "2018-10-05T15:19:21.953000",
          "content": "<p>Thanks Martin. I was using the 'download all' link. I shall try the test.zip file instead.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 399196,
      "author_name": "Saikiran N. Pasikanti",
      "author_url": "",
      "post_date": "2018-10-05T11:50:18.977000",
      "content": "<p>Now, it is very clear</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 398704,
      "author_name": "Renan Rodrigues dos Santos",
      "author_url": "",
      "post_date": "2018-10-04T13:23:38.437000",
      "content": "<p>Hi Emma,\nThank you for hosting this competition. It is a very interesting domain to apply Machine Learning.</p>\n\n<p>I'm assuming the CPU + 4Gb RAM restriction has to do with integrating our models in a \"smart-microscopy system\" as said in the Overview page of this competition. Am I right?</p>\n\n<p>Since this isn't a Kaggle Kernel CPU only competition, I'd like to understand from you and from Kaggle staff if, for this competition, we can have two separate models:</p>\n\n<ul>\n<li>one for top accuracy, using GPU hardware which is standard for image problems. This one competes for the normal money prize</li>\n<li>another one for fast prediction in CPU but lower accuracy, competing for the special prize</li>\n</ul>\n\n<p>If that's not possible, we will forcibly have to choose what to model for: best accuracy/fastest CPU model - normal prize/special prize.</p>\n\n<p>Looking forward for news from you.\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 398796,
          "author_name": "Martin Hjelmare",
          "author_url": "",
          "post_date": "2018-10-04T15:37:53.630000",
          "content": "<p>Hi Renan, thanks for your feedback!</p>\n\n<p>Yes, we have put the limits on the hardware to match our smart-microscopy system workflow.</p>\n\n<p>In the rules it says:</p>\n\n<blockquote>\n  <p>You may select up to 2 final submissions for judging.</p>\n</blockquote>\n\n<p>This should allow you to compete with two separate models. We're excited to see what different approaches the community will take.</p>\n\n<p>Martin, part of the Human Protein Atlas team</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 424265,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-19T20:15:29.650000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415813,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-05T18:34:55.757000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 399442,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-05T19:35:06.563000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 399488,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-05T22:03:06.763000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415023,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-04T07:24:38.720000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415288,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-04T20:21:11.847000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 463084,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-29T11:27:03.270000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 399415,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-05T18:30:41.077000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 401999,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-11T02:51:25.897000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "398306": "Welcome to the Human Protein Atlas Image Classification Challenge. Help us classify subcellular protein patterns and win great prizes, provided by Leica Microsystems and Nvidia.\n\nWe wish you all the best and look forward to see how the Kaggle community will tackle this problem and what new models will be developed. We will do our best to answer any questions the community might have.\n\nA note about the data:\nThere are two versions of the data: PNGs and TIFFs. The PNG sets are much smaller than the TIFF sets. The PNG sets have been downscaled to 512x512 pixels and converted to 8 bit, while the TIFF sets keep the full 2048x2048 or 3072x3072 pixels but have been converted to 8 bit, from the 16 bit of the raw data.\n\nWhen running the evaluation of the special prize, we will use the TIFF test set. For the main competition the evaluation procedure doesn't differ between PNG and TIFF sets.\n\nMore information can be found on the [Data page](https://www.kaggle.com/c/human-protein-atlas-image-classification/data) and on the [Special Prize Instructions page](https://www.kaggle.com/c/human-protein-atlas-image-classification#Special-Prize-Instructions).\n\nHave fun!\n\nUpdate Dec 6: There are 15 samples that have a different pixel size (4096, 4096), see the attached files for a list of these.",
    "432897": "Are kaggle or the organisers going to comment on the data leak?\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395",
    "423743": "In the data set there're many classes has co-occurence with the other classes. For example, One image is labeled as  both Nucleoplasm and Cytosol. So, how to understand by looking at images that why it is labeled to the both classes? If I see an image only labeled with the Nucleoplasm or Cytosol by combining all 4 channels, I can't find any visual similarities. If I see an image classified with the both classes then what should I look into so it can explain that the image is belong to both classes. ",
    "419779": "I found that there are also 4096x4096 px images in the training dataset. 7 to be exact.\nId's:\n0858d008-bb9e-11e8-b2b9-ac1f6b6435d0\n16196e56-bba2-11e8-b2b9-ac1f6b6435d0\n1625b570-bbc6-11e8-b2bc-ac1f6b6435d0\n28be7d5e-bbbf-11e8-b2ba-ac1f6b6435d0\n558dfed4-bba3-11e8-b2b9-ac1f6b6435d0\nc3322376-bbc0-11e8-b2bb-ac1f6b6435d0\nf2998bbe-bbab-11e8-b2ba-ac1f6b6435d0\n\nThis lead to a bug that cost me some calculation and debugging time. Consider adding this information to your Post above and the competition description where you list the image sizes.\n\nCheers,\nBasti",
    "414929": "I can't upload and submit data in these two days. I tried almost all methods and couldn't solve them. What should I do?",
    "411984": "Hi Emma,\nWhile there exist most of the works done for protein subcellular classification using textual data, why we have to do same with images. How cost effective is images compare to text data prediction. What is the intention behind giving these image dataset for the same classification which is done using textual data already? As it will be questionable by the reviewer when a paper is submitted for publication. Any response would be appreciable :) Thanks",
    "399804": "Looking forward to participating! Do you think it is possible to get top performing results using the PNG set or would this only be possible by training on the TIFF set?",
    "398935": "Thanks for your reply, Martin.\nIt is very clear now.",
    "401008": "Following up on the questions about the hardware limits of the special prize evaluation, we want to clarify that we only limit the hardware of the prediction process that we will perform when evaluating submissions. Training can still be done without hardware restrictions, since this is done before submission.\n\nWe understand that the hardware limits can still impact top performance, but we are interested to find models that balance performance, speed and necessary hardware resources. The reason why we have put the limits on the hardware is to match our smart-microscopy system workflow, as mentioned above.\n\nThanks for the feedback and please continue the discussion on this topic. We want to hear what you think.",
    "398314": "&gt;Top performing teams are invited to make the model code and any external data used to generate their final submission available to ...\n\n_\n&gt; The winner will be the fastest model that maintains its submitted F1 score in the main prize competition +/- 1 percentage point.\n\nThat is a bit too optimistic, isn't it? Top submissions will need significantly more FLOPS than what 2 CPU cores can produce.",
    "514526": "Hi,\nCan you share how you converted 16-bit tiff files to the 8-bit images?",
    "443203": "Hello Emma,\nCould you please explain the rules concerning reproducibility of the final submission:\nShould the whole training process be completely reproducible or only the inference part?\nThank you",
    "431055": "Hi, Emma Lundberg.\nI have a question. If there have colorful sample image, no like these grey?",
    "427150": "Hi Emma,\n\nI have a problem with downloading tiff images.\nWhen I download from google bucket via browser (tried Chrome and Firefox), anywhere in the middle there is some connection error (maybe short disconnect or smth like that), and downloading interrupts, and browsers are not clever enough to continue...\n\nI wonder if there is any other fool-proof possibility of downloading tiff images (e.g. torrent file)?\n\nDid anybody have a similar problem maybe?\n\nRegards,\nMaksim\n",
    "425719": "The note on data says \" The green filter should hence be used to predict the label, and the other filters are used as references\" If that is the case then in the test/predict samples should not be divided into different colors.  Can someone explain I am confused as the green one carries the information on the location of the protein.",
    "425056": "Hello, \nIf we are doing this competition primarily for fun/learning, will it still be possible to submit entries for testing/test set labels revealed after the challenge has ended, just that they wouldn't count towards the competition?",
    "422943": "Hi, Emma,\n\nEmma, I'd like to read the article but it appears to have a paywall. Is there a way around this? I'll pay for it but I'd prefer if there was a free version or some equivalent free version.\n\nAlso, 0.71 was achieved by experts on this data. But who provided the ground truth for these experts? Other experts?\n\nRegards,\nPete",
    "421507": "I've used 256x256 color images on my hardware before.\n\nI'm wondering what hardware stack would you need to flow 3072x3072 images through even a moderate size CNN? I can't even imagine.",
    "408161": "&gt; You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.\n\nHi, is there an official thread about pre-trained model and outside dataset?",
    "406524": "Hi, I still have some trouble with the interpretation of what is being said about the filters. Is it correct to assume that only green filter files should be used for training and classification? ",
    "404308": "What macro F-score human experts achieve when labeling proteins?",
    "401163": "I'm sorry if someone already asked, are there any meanings in the order of Target?",
    "400485": "Hi, Emma, thank you for this competition!\nI do not understand what the meaning of green channel is? Is it mask made by people or another variant of cells photo with colorant? How to interpretate \"the protein of interest\"?",
    "399993": "I have some trouble understanding what we should do : how is it possible that \"microtubules\" is both a channel (so present in all images) and a label (present or not in images)... ",
    "399778": "nyc",
    "399200": "Hi Emma\n\nCan I just check I understand the problem correctly. \n\nWe are predicting where, in a particular cell type,  a particular protein is located, denoted by labels referring to cell features. So a particular protein will be in the same locations in similar cells, but at different locations in different cell types. Is that right? Can we assume the proportions of proteins, cell types and labels in the datasets are typical?\n\nPerhaps the problem is obvious from the dataset but the zip file seems to be corrupted. I've tried downloading it twice.",
    "399196": "Now, it is very clear",
    "398704": "Hi Emma,\nThank you for hosting this competition. It is a very interesting domain to apply Machine Learning.\n\nI'm assuming the CPU + 4Gb RAM restriction has to do with integrating our models in a \"smart-microscopy system\" as said in the Overview page of this competition. Am I right?\n\nSince this isn't a Kaggle Kernel CPU only competition, I'd like to understand from you and from Kaggle staff if, for this competition, we can have two separate models:\n\n - one for top accuracy, using GPU hardware which is standard for image problems. This one competes for the normal money prize\n - another one for fast prediction in CPU but lower accuracy, competing for the special prize\n\nIf that's not possible, we will forcibly have to choose what to model for: best accuracy/fastest CPU model - normal prize/special prize.\n\nLooking forward for news from you.\nThanks",
    "424265": "",
    "415813": "",
    "399442": "",
    "399415": ""
  }
}