{
  "id": 69984,
  "title": "Official pre-trained models and external data thread",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/69984",
  "author_name": "Martin Hjelmare",
  "post_date": "2018-10-29T15:35:35.428000",
  "votes": 34,
  "comment_count": 243,
  "views": 0,
  "content": "<p>Please post which pre-trained model(s) you are using in this thread.</p>\n\n<blockquote>\n  <p>EXTERNAL DATA\n  You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.</p>\n</blockquote>",
  "messages": [
    {
      "id": 430860,
      "postDate": "2018-12-01T04:28:06.257Z",
      "content": "<p>Hi! @FabSchreiber </p>\n\n<p>Here's list</p>\n\n<p>!! still contains duplicated images\n　RGB images(HPAv18.csv) :  77,878 sample NG\n　RGB images 77,864 sample  ...deleted id duplication,still contains duplicated images\n　RGBY images  77,430 sample\n　RGBY withoutUncertain:73,881 sample</p>\n\n<p>old csv list contains\n GeneID_ Dir_ImageURL Target(28 class)</p>\n\n<p>\"<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)</p>\n\n<p><strong>Modified:</strong> ..without duplicated images (Gene information lost,labels merged)\n　RGB_wodpl 75,040 sample\n　RGBY_wodpl  74,606 sample\n　RGBY withoutUncertain_wodpl:71,437 sample</p>\n\n<p>new csv list contains\n　Dir_ImageURL Target(28 class)</p>\n\n<p><strong>AddFiles:</strong> @Chase_the_Trane adviced me to open how I make those csv files.\n　 1. Parce XML and Download HPAv18 Image.html\n--&gt;total downloaded image 60GB\n　  2.  NoYellow512.txt\n--&gt;resize img 512x512 png 70GB, and I realized some sample has no yellow filter.\n　  3. Make Metadata for HPAv18 Image.html\n--&gt;duplicate image exist\n　 4. Clean HPAv18 dataset.html...How I merged the labels.</p>\n\n<p><strong>AddFiles2:</strong> add CellLine information \n*_withCellLine.csv</p>",
      "rawMarkdown": "Hi! @FabSchreiber \n\nHere's list\n\n!! still contains duplicated images\n　RGB images(HPAv18.csv) :  77,878 sample NG\n　RGB images 77,864 sample  ...deleted id duplication,still contains duplicated images\n　RGBY images  77,430 sample\n　RGBY withoutUncertain:73,881 sample\n\nold csv list contains\n GeneID_ Dir_ImageURL Target(28 class)\n\n\"http://v18.proteinatlas.org/images/\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)\n\n**Modified:** ..without duplicated images (Gene information lost,labels merged)\n　RGB\\_wodpl 75,040 sample\n　RGBY\\_wodpl  74,606 sample\n　RGBY withoutUncertain_wodpl:71,437 sample\n\nnew csv list contains\n　Dir_ImageURL Target(28 class)\n\n**AddFiles:** @Chase\\_the\\_Trane adviced me to open how I make those csv files.\n　 1. Parce XML and Download HPAv18 Image.html\n--&gt;total downloaded image 60GB\n　  2.  NoYellow512.txt\n--&gt;resize img 512x512 png 70GB, and I realized some sample has no yellow filter.\n　  3. Make Metadata for HPAv18 Image.html\n--&gt;duplicate image exist\n　 4. Clean HPAv18 dataset.html...How I merged the labels.\n\n**AddFiles2:** add CellLine information \n*\\_withCellLine.csv",
      "votes": 39,
      "replies": [
        {
          "id": 431235,
          "postDate": "2018-12-01T22:53:04.120Z",
          "content": "<p>I noticed a few duplicates here. Different ENSG id, but same image paths.</p>",
          "rawMarkdown": "I noticed a few duplicates here. Different ENSG id, but same image paths.",
          "votes": 1
        },
        {
          "id": 431287,
          "postDate": "2018-12-02T01:50:10.870Z",
          "content": "<p>Yes, I'm seeing only 75,040 unique paths.  For one extreme example, it looks like path 10580_1610_C1_1 is repeated 22 times and has four different label assignments.   So we have dups + some label noise.   It would be interesting if an expert would examine the associated Ensembl gene sets to determine how much of a cause for concern this might be.</p>",
          "rawMarkdown": "Yes, I'm seeing only 75,040 unique paths.  For one extreme example, it looks like path 10580\\_1610\\_C1\\_1 is repeated 22 times and has four different label assignments.   So we have dups + some label noise.   It would be interesting if an expert would examine the associated Ensembl gene sets to determine how much of a cause for concern this might be.",
          "votes": 1
        },
        {
          "id": 431291,
          "postDate": "2018-12-02T01:58:11.767Z",
          "content": "<p>&gt;Different ENSG id, but same image paths.\nYou are right. There are 4552 such duplicate rows in one file.\nI should have checked before sharering the list, sorry.</p>\n\n<p>I 'm now trying to remove duplicates, \nand found same image but different labels, \nso not simply remove duplicate rows but merge labels.</p>\n\n<p>e.g.\nENSG00000081853\nCytosol (GO:0005829);Nucleoli (GO:0005730);Nucleus (GO:0005634);Plasma membrane (GO:0005886);Vesicles (GO:0043231)</p>\n\n<p>ENSG00000240764\nCytosol (GO:0005829);Nucleoplasm (GO:0005654);Nucleus (GO:0005634);Plasma membrane (GO:0005886);Vesicles (GO:0043231)</p>",
          "rawMarkdown": "&gt;Different ENSG id, but same image paths.\nYou are right. There are 4552 such duplicate rows in one file.\nI should have checked before sharering the list, sorry.\n\nI 'm now trying to remove duplicates, \nand found same image but different labels, \nso not simply remove duplicate rows but merge labels.\n\ne.g.\nENSG00000081853\nCytosol (GO:0005829);Nucleoli (GO:0005730);Nucleus (GO:0005634);Plasma membrane (GO:0005886);Vesicles (GO:0043231)\n\nENSG00000240764\nCytosol (GO:0005829);Nucleoplasm (GO:0005654);Nucleus (GO:0005634);Plasma membrane (GO:0005886);Vesicles (GO:0043231)"
        },
        {
          "id": 431367,
          "postDate": "2018-12-02T05:30:06.990Z",
          "content": "<p>I only noticed because my latest download script checks and skips any files that exist already. For now I am leaving the noise in, having multiple rows with the same image id but different labels. I think it might help with overfitting.</p>",
          "rawMarkdown": "I only noticed because my latest download script checks and skips any files that exist already. For now I am leaving the noise in, having multiple rows with the same image id but different labels. I think it might help with overfitting."
        },
        {
          "id": 431484,
          "postDate": "2018-12-02T11:22:13.980Z",
          "content": "<p>Thanks Tomomi and Brian.  BTW, when typing a word that contains more than one underscore, use a backslash in front of each so markup does not convert the intervening text to italics.  I think this trick may work for other special characters as well, e..g \\&lt;;  these can alternatively be specified with an ampersand.</p>",
          "rawMarkdown": "Thanks Tomomi and Brian.  BTW, when typing a word that contains more than one underscore, use a backslash in front of each so markup does not convert the intervening text to italics.  I think this trick may work for other special characters as well, e..g \\&lt;;  these can alternatively be specified with an ampersand.",
          "votes": 1
        },
        {
          "id": 431523,
          "postDate": "2018-12-02T12:35:07.550Z",
          "content": "<p>Thank you @Russ. \\trick worked!</p>",
          "rawMarkdown": "Thank you @Russ. \\trick worked!"
        },
        {
          "id": 432922,
          "postDate": "2018-12-04T13:26:21.433Z",
          "content": "<p>Thanks for your script.\nNow I am running it to download the HPAv18 data,</p>",
          "rawMarkdown": "Thanks for your script.\nNow I am running it to download the HPAv18 data,",
          "votes": 1
        }
      ]
    },
    {
      "id": 412120,
      "postDate": "2018-10-29T15:35:35.430Z",
      "content": "<p>Please post which pre-trained model(s) you are using in this thread.</p>\n\n<blockquote>\n  <p>EXTERNAL DATA\n  You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.</p>\n</blockquote>",
      "rawMarkdown": "Please post which pre-trained model(s) you are using in this thread.\n\n&gt; EXTERNAL DATA\nYou may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.",
      "votes": 34
    },
    {
      "id": 436319,
      "postDate": "2018-12-10T05:52:40.370Z",
      "content": "<p>Those who read this comment, please help me.\nI want to use some open source code and need some advice about the usage...</p>\n\n<p>One month a go, <a href=\"/hengck23\">@hengck23</a> introduced this paper</p>\n\n<p>Learning unsupervised feature representations for single cell microscopy images with paired cell inpainting\nAlex Lu, Oren Z Kraus, Sam Cooper, Alan M Moses</p>\n\n<p>on [ideas and discussion] thread\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69955\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69955</a></p>\n\n<p>I found it very attractive, but source code was not released at the time.</p>\n\n<p>One week ago, source code has finally released!\n<a href=\"https://github.com/alexxijielu/paired_cell_inpainting#human-model\">https://github.com/alexxijielu/paired_cell_inpainting#human-model</a></p>\n\n<p>It contains \n- Cleaner HPA data downloading scripts.\n- Extracting single cell.\n- pre-trained weights \n- and more</p>\n\n<p>Their task setting(unsupervised) is different from us(supervised classification),\nbut can be used with some extra code.</p>\n\n<p>They released the code under [GPL-2.0]</p>\n\n<p>My question is, \n\"Is it OK to use this code and pre-trained weight as a baseline model?\"\n if it's against the rule \"Can I re-implement from scratch and imitate the idea?\"\n if it's within the rule \"Should I ask the author of this code about the usage in Github/Issues section beforehand?\"\nOr they might also participated in this competition already!?</p>",
      "rawMarkdown": "Those who read this comment, please help me.\nI want to use some open source code and need some advice about the usage...\n\n\nOne month a go, @hengck23 introduced this paper\n\nLearning unsupervised feature representations for single cell microscopy images with paired cell inpainting\nAlex Lu, Oren Z Kraus, Sam Cooper, Alan M Moses\n\non [ideas and discussion] thread\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69955\n\nI found it very attractive, but source code was not released at the time.\n\nOne week ago, source code has finally released!\nhttps://github.com/alexxijielu/paired_cell_inpainting#human-model\n\nIt contains \n- Cleaner HPA data downloading scripts.\n- Extracting single cell.\n- pre-trained weights \n- and more\n\nTheir task setting(unsupervised) is different from us(supervised classification),\nbut can be used with some extra code.\n\nThey released the code under [GPL-2.0]\n\nMy question is, \n\"Is it OK to use this code and pre-trained weight as a baseline model?\"\n if it's against the rule \"Can I re-implement from scratch and imitate the idea?\"\n if it's within the rule \"Should I ask the author of this code about the usage in Github/Issues section beforehand?\"\nOr they might also participated in this competition already!?",
      "votes": 13,
      "replies": [
        {
          "id": 436726,
          "postDate": "2018-12-10T20:36:46.483Z",
          "content": "<p>Thank you for sharing. My understanding is as long as the algorithm described in the paper is not patented we are okay to use the implementation (software) to compete whatever the license given is. Let's wait for the organizers to rule this.</p>",
          "rawMarkdown": "Thank you for sharing. My understanding is as long as the algorithm described in the paper is not patented we are okay to use the implementation (software) to compete whatever the license given is. Let's wait for the organizers to rule this.",
          "votes": 1
        },
        {
          "id": 436740,
          "postDate": "2018-12-10T21:24:03.863Z",
          "content": "<p>Hi! Original author of this preprint and code here! I was wondering where the sudden traffic on my GitHub came from!</p>\n\n<p>I'd be delighted if anyone ended up using my work to pre-train a model for this competition, and you're free to use my code if the Kaggle organizers rule this to be okay. It would be a really good way to validate the utility of our method, so I totally encourage it. We were hoping to put together a team from my lab, but as the competition deadline inches closer and we're all busy with end-of-semester academic stuff, we aren't sure if we can put together a fully polished result. </p>\n\n<p>Off the top of my head, here's some of the problems and limitations you'd need to solve:\n - Our method is designed to extract single-cell features, and you'd need to find a way to aggregate these into a full-image prediction. You could just take an average of the features and train a classifier on that, but that seems a bit unrefined - I'd love to see what kind of an end-to-end approach might be possible for this (a RCNN maybe?)\n - One major limitation is with phenotypes that don't penetrate well (e.g. mitotic spindle, etc.) In my preliminary exploration of the feature space, I think the model ends up learning to ignore phenotypes where maybe only one out of fifty cells expresses the phenotype. It sounds like fine-tuning might solve this problem though.\n - Preprocessing: to input single-cell crops into this method, you need to, of course, get single cell crops. The current method we have on the GitHub is very ad-hoc - we just run the human nuclei channel through an Otsu filter. I'm sure a better segmentation method would be more robust to this.</p>\n\n<p>Let me know if you have any questions/comments and I'll do my best to address them! </p>",
          "rawMarkdown": "Hi! Original author of this preprint and code here! I was wondering where the sudden traffic on my GitHub came from!\n\nI'd be delighted if anyone ended up using my work to pre-train a model for this competition, and you're free to use my code if the Kaggle organizers rule this to be okay. It would be a really good way to validate the utility of our method, so I totally encourage it. We were hoping to put together a team from my lab, but as the competition deadline inches closer and we're all busy with end-of-semester academic stuff, we aren't sure if we can put together a fully polished result. \n\nOff the top of my head, here's some of the problems and limitations you'd need to solve:\n - Our method is designed to extract single-cell features, and you'd need to find a way to aggregate these into a full-image prediction. You could just take an average of the features and train a classifier on that, but that seems a bit unrefined - I'd love to see what kind of an end-to-end approach might be possible for this (a RCNN maybe?)\n - One major limitation is with phenotypes that don't penetrate well (e.g. mitotic spindle, etc.) In my preliminary exploration of the feature space, I think the model ends up learning to ignore phenotypes where maybe only one out of fifty cells expresses the phenotype. It sounds like fine-tuning might solve this problem though.\n - Preprocessing: to input single-cell crops into this method, you need to, of course, get single cell crops. The current method we have on the GitHub is very ad-hoc - we just run the human nuclei channel through an Otsu filter. I'm sure a better segmentation method would be more robust to this.\n\nLet me know if you have any questions/comments and I'll do my best to address them! ",
          "votes": 23
        },
        {
          "id": 436751,
          "postDate": "2018-12-10T21:38:42.827Z",
          "content": "<p>Alex, thanks. You should join and set a strong baseline for us</p>",
          "rawMarkdown": "Alex, thanks. You should join and set a strong baseline for us",
          "votes": 3
        },
        {
          "id": 436788,
          "postDate": "2018-12-11T00:06:04.130Z",
          "content": "<p>Thank you!!!!!! Alex.\nThank you for allowing everyone to use your great model and sharing many ideas.</p>",
          "rawMarkdown": "Thank you!!!!!! Alex.\nThank you for allowing everyone to use your great model and sharing many ideas.",
          "votes": 1
        },
        {
          "id": 436794,
          "postDate": "2018-12-11T00:34:13.843Z",
          "content": "<p>No problem. :) If you end up using my pre-trained weights, use the one named \"filtered_model_weights.h5\", those are way better than the default because we filtered out all variable proteins in training it instead of treating the problem as a fully unsupervised one. I wish we had weights for stronger models to offer, but AlexNet was all we experimented with for now. In any case, I think it should be better than ImageNet pretrained weights, which I think most people are using, just because the weights have learned HPA-specific channel correlations. </p>",
          "rawMarkdown": "No problem. :) If you end up using my pre-trained weights, use the one named \"filtered_model_weights.h5\", those are way better than the default because we filtered out all variable proteins in training it instead of treating the problem as a fully unsupervised one. I wish we had weights for stronger models to offer, but AlexNet was all we experimented with for now. In any case, I think it should be better than ImageNet pretrained weights, which I think most people are using, just because the weights have learned HPA-specific channel correlations. ",
          "votes": 2
        },
        {
          "id": 436837,
          "postDate": "2018-12-11T02:26:59.413Z",
          "content": "<p>Great code, thanks for sharing. I've integrated it into my processing and generated over 1M single cell images to feed into a network:</p>\n\n<p><img src=\"http://brians.network/images/hpa_cells.png\" alt=\"Image\"></p>",
          "rawMarkdown": "Great code, thanks for sharing. I've integrated it into my processing and generated over 1M single cell images to feed into a network:\n\n![Image](http://brians.network/images/hpa_cells.png)",
          "votes": 1
        },
        {
          "id": 437012,
          "postDate": "2018-12-11T08:49:02.370Z",
          "content": "<p>I have experimented with single cell images and was not overly successful. Maybe that was because I did not aggregate the scores properly for an entire image, but honestly I think that it was unsuccessful because we have only image-level annotations. I looked at some images manually and noticed (especially for the very rare classes) that the structure of interest (such as cell joints or some other small classes) must be present in only one cell within the entire image for the entire image to get this class. Given that there are always multiple cells in an image, sometimes even &gt;20, you would need to reformulate the whole  task as a <strong>multiple instance learning</strong> problem.\nAlso I would be interested in understanding how you (@Alex Lu) managed to separate cells. Cells are of very different size and often cells are touching with their borders so if you crop around the nucleus of a cell you can end up with multiple cells in your crop. (But honestly I think this is much less of a problem than the multiple instance learning thing).</p>",
          "rawMarkdown": "I have experimented with single cell images and was not overly successful. Maybe that was because I did not aggregate the scores properly for an entire image, but honestly I think that it was unsuccessful because we have only image-level annotations. I looked at some images manually and noticed (especially for the very rare classes) that the structure of interest (such as cell joints or some other small classes) must be present in only one cell within the entire image for the entire image to get this class. Given that there are always multiple cells in an image, sometimes even &gt;20, you would need to reformulate the whole  task as a **multiple instance learning** problem.\nAlso I would be interested in understanding how you (@Alex Lu) managed to separate cells. Cells are of very different size and often cells are touching with their borders so if you crop around the nucleus of a cell you can end up with multiple cells in your crop. (But honestly I think this is much less of a problem than the multiple instance learning thing)."
        },
        {
          "id": 437171,
          "postDate": "2018-12-11T13:36:16.013Z",
          "content": "<p>We simply segmented the nuclei channel and extracted a fixed crop around the cell. You're right that it does sometime cause multiple cells in the same crop, but that generally doesn't seem to pose an issue for learning the feature representation - you can check out in the paper how we used the model generatively, because there's an example of where we inpaint a protein from a crop with two cells onto a single cell, I think.</p>\n\n<p>In general, the method we use seems to learn a feature representation robust to single cell variability, including morphology, cell line, etc. I think this will be important to the competition - knowing the domain, the organizers may try to confound your models by giving them unseen cell lines in the test stage, because this would really emphasize generalizable models. Imagine going from bone cells to skin cells or brain cells - they look very different, and you want to make sure your model hasn't overfit to the morphologies in your training set.</p>\n\n<p>You're right that the challenge is going to be aggregating these crops, though! But I don't think the use of single cell crops is fundamentally a bad idea - if you think about it, it's basically a pseudo-attention mechanism, that directs your network to the relevant parts of the image. MIL sounds like a good idea for solving some of the issues with poorly penetrating classes. </p>",
          "rawMarkdown": "We simply segmented the nuclei channel and extracted a fixed crop around the cell. You're right that it does sometime cause multiple cells in the same crop, but that generally doesn't seem to pose an issue for learning the feature representation - you can check out in the paper how we used the model generatively, because there's an example of where we inpaint a protein from a crop with two cells onto a single cell, I think.\n\nIn general, the method we use seems to learn a feature representation robust to single cell variability, including morphology, cell line, etc. I think this will be important to the competition - knowing the domain, the organizers may try to confound your models by giving them unseen cell lines in the test stage, because this would really emphasize generalizable models. Imagine going from bone cells to skin cells or brain cells - they look very different, and you want to make sure your model hasn't overfit to the morphologies in your training set.\n\nYou're right that the challenge is going to be aggregating these crops, though! But I don't think the use of single cell crops is fundamentally a bad idea - if you think about it, it's basically a pseudo-attention mechanism, that directs your network to the relevant parts of the image. MIL sounds like a good idea for solving some of the issues with poorly penetrating classes. ",
          "votes": 5
        }
      ]
    },
    {
      "id": 432870,
      "postDate": "2018-12-04T12:13:17.520Z",
      "content": "<p>Hi, @pete\nXML parsing is time consuming.\nAlternatively, you can download from csv list.</p>\n\n<p>(How I processed the list was posted few comments ago.\nMy code is far from efficient and safe, there's no Error handrring, though )</p>\n\n<p><strong>Note 12/13:</strong>\n*<em>Please use <a href=\"/davidwagnerkc\">@davidwagnerkc</a> 's downloading script !</em>*\nNot gray-scaled image hurts me:11b_Segment_and_Crop.html\nToo many class 27?:Check metadata for HPAv18.html</p>",
      "rawMarkdown": "Hi, @pete\nXML parsing is time consuming.\nAlternatively, you can download from csv list.\n\n(How I processed the list was posted few comments ago.\nMy code is far from efficient and safe, there's no Error handrring, though )\n\n**Note 12/13:**\n**Please use @davidwagnerkc 's downloading script !**\nNot gray-scaled image hurts me:11b_Segment_and_Crop.html\nToo many class 27?:Check metadata for HPAv18.html",
      "votes": 14,
      "replies": [
        {
          "id": 436787,
          "postDate": "2018-12-11T00:04:50.230Z",
          "content": "<p>Hi <a href=\"/tomomimoriyama\">@tomomimoriyama</a>, thanks for sharing this great scripts. After getting HPAv18 datasets, I found that the downloaded images are all having 3 channels. For example, <code>img_HPAv18_some_name_blue.jpg</code> has 3 channels and all channels have some values. Would you tell how you convert these 3-channel \"RGB\" image to a grayscale one? <code>0.299 R + 0.587 G + 0.114 B</code> is what I am doing now but some images ended up with very low visibility, not sure if I am doing something wrong.</p>\n\n<p>The following are two images, the first one is <code>11_203_h3_1_yellow</code> to grayscale, the second one is <code>11_203_h3_1_blue</code> to grayscale. Before converting to grayscale the blue one looks good, but after the conversion, it becomes very dark.</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/14886380/49769533-c762b500-fc9d-11e8-9570-fbcfaf4ca370.png\" alt=\"yellow\">\n<img src=\"https://user-images.githubusercontent.com/14886380/49769530-c762b500-fc9d-11e8-8580-e881e1605458.png\" alt=\"blue\"></p>",
          "rawMarkdown": "Hi @tomomimoriyama, thanks for sharing this great scripts. After getting HPAv18 datasets, I found that the downloaded images are all having 3 channels. For example, `img_HPAv18_some_name_blue.jpg` has 3 channels and all channels have some values. Would you tell how you convert these 3-channel \"RGB\" image to a grayscale one? `0.299 R + 0.587 G + 0.114 B` is what I am doing now but some images ended up with very low visibility, not sure if I am doing something wrong.\n\nThe following are two images, the first one is `11_203_h3_1_yellow` to grayscale, the second one is `11_203_h3_1_blue` to grayscale. Before converting to grayscale the blue one looks good, but after the conversion, it becomes very dark.\n\n![yellow][1]\n![blue][2]\n\n\n  [1]: https://user-images.githubusercontent.com/14886380/49769533-c762b500-fc9d-11e8-9570-fbcfaf4ca370.png\n  [2]: https://user-images.githubusercontent.com/14886380/49769530-c762b500-fc9d-11e8-8580-e881e1605458.png",
          "votes": 9
        },
        {
          "id": 436934,
          "postDate": "2018-12-11T05:42:47.477Z",
          "content": "<p>Hi <a href=\"/tomomimoriyama\">@tomomimoriyama</a>, </p>\n\n<p>Your scripts are working fine for me but the download time is very long. I have downloaded about 18k of the 74k  images in about 11 hours, maybe there are many kagglers accessing the protein atlas.</p>\n\n<p>Thanks</p>\n\n<p>Kickback</p>",
          "rawMarkdown": "Hi @tomomimoriyama, \n\nYour scripts are working fine for me but the download time is very long. I have downloaded about 18k of the 74k  images in about 11 hours, maybe there are many kagglers accessing the protein atlas.\n\nThanks\n\nKickback"
        },
        {
          "id": 436948,
          "postDate": "2018-12-11T06:24:13.087Z",
          "content": "<p>I used <a href=\"/tomomimoriyama\">@tomomimoriyama</a>'s csv file to download all RGBY images and resize to (512, 512) in under 8 hours if I remember correctly. Here is the code from my notebook, I haven't tested it, but it worked before in notebook form. Maybe map only <code>df[:1000].Id</code> to make sure it works and get a time estimate. I ran this on a MBP with 4 cores. </p>\n\n<p>Change where csv is read from and the <code>hpa_dir</code> you want to write to. </p>\n\n<p>EDIT: Thanks to @Mark Worrall I am actually downloading RGBY images now, was missing B channel before!\n```\nimport pandas as pd\nimport requests\nfrom multiprocessing import Pool\nfrom pathlib import Path\nfrom PIL import Image\nfrom io import BytesIO</p>\n\n<p>df = pd.read_csv('../HPAv18RBGY_wodpl.csv')</p>\n\n<p>def download(id_):\n    try:\n        hpa_dir = Path('../HPAv18')\n        hpa_dir.mkdir(parents=True, exist_ok=True)\n        image_dir, _, image_id = id_.partition('_')\n        url = f'<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>{image_dir}/{image_id}_blue_red_green_yellow.jpg'\n        r = requests.get(url)\n        image = Image.open(BytesIO(r.content)).resize((512, 512), Image.LANCZOS)\n        image.save(hpa_dir / f'{id_}.png', format='png')\n    except:\n        print(f'{id_} broke...')</p>\n\n<p>p = Pool()\np.map(download, df.Id)\n```</p>",
          "rawMarkdown": "I used @tomomimoriyama's csv file to download all RGBY images and resize to (512, 512) in under 8 hours if I remember correctly. Here is the code from my notebook, I haven't tested it, but it worked before in notebook form. Maybe map only `df[:1000].Id` to make sure it works and get a time estimate. I ran this on a MBP with 4 cores. \n\nChange where csv is read from and the `hpa_dir` you want to write to. \n\nEDIT: Thanks to @Mark Worrall I am actually downloading RGBY images now, was missing B channel before!\n```\nimport pandas as pd\nimport requests\nfrom multiprocessing import Pool\nfrom pathlib import Path\nfrom PIL import Image\nfrom io import BytesIO\n\ndf = pd.read_csv('../HPAv18RBGY_wodpl.csv')\n\ndef download(id_):\n    try:\n        hpa_dir = Path('../HPAv18')\n        hpa_dir.mkdir(parents=True, exist_ok=True)\n        image_dir, _, image_id = id_.partition('_')\n        url = f'http://v18.proteinatlas.org/images/{image_dir}/{image_id}_blue_red_green_yellow.jpg'\n        r = requests.get(url)\n        image = Image.open(BytesIO(r.content)).resize((512, 512), Image.LANCZOS)\n        image.save(hpa_dir / f'{id_}.png', format='png')\n    except:\n        print(f'{id_} broke...')\n\np = Pool()\np.map(download, df.Id)\n```\n",
          "votes": 7
        },
        {
          "id": 437011,
          "postDate": "2018-12-11T08:48:55.647Z",
          "content": "<p>A multi-processing version may works better for chinese guys who has a gfw :)</p>\n\n<pre><code>import os\nfrom multiprocessing.pool import Pool\nfrom tqdm import tqdm\nimport requests\nimport pandas as pd\ndef download(pid, sp, ep):\n    colors = ['red', 'green', 'blue', 'yellow']\n    DIR = \"/home/czy/dataset/human_protain/moredata/\"\n    v18_url = 'http://v18.proteinatlas.org/images/'\n    imgList = pd.read_csv(\"/home/czy/dataset/human_protain/moredata/HPAv18RBGY_wodpl.csv\")\n    for i in tqdm(imgList['Id'][sp:ep], postfix=pid):  # [:5] means downloard only first 5 samples, if it works, please remove it\n        img = i.split('_')\n        for color in colors:\n            img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n            img_name = i + \"_\" + color + \".jpg\"\n            img_url = v18_url + img_path\n            r = requests.get(img_url, allow_redirects=True)\n            open(DIR + img_name, 'wb').write(r.content)\n\ndef run_proc(name, sp, ep):\n    print('Run child process %s (%s) sp:%d ep: %d' % (name, os.getpid(), sp, ep))\n    download(name, sp, ep)\n    print('Run child process %s done' % (name))\n\nif __name__ == \"__main__\":\n    print('Parent process %s.' % os.getpid())\n    img_list = pd.read_csv(\"/home/czy/dataset/human_protain/moredata/HPAv18RBGY_wodpl.csv\")['Id']\n    list_len = len(img_list)\n    process_num = 100\n    p = Pool(process_num)\n    for i in range(process_num):\n        p.apply_async(run_proc, args=(str(i), int(i * list_len / process_num), int((i + 1) * list_len / process_num)))\n    print('Waiting for all subprocesses done...')\n    p.close()\n    p.join()\n    print('All subprocesses done.')\n</code></pre>",
          "rawMarkdown": "A multi-processing version may works better for chinese guys who has a gfw :)\n\n   \n\n    import os\n    from multiprocessing.pool import Pool\n    from tqdm import tqdm\n    import requests\n    import pandas as pd\n    def download(pid, sp, ep):\n        colors = ['red', 'green', 'blue', 'yellow']\n        DIR = \"/home/czy/dataset/human_protain/moredata/\"\n        v18_url = 'http://v18.proteinatlas.org/images/'\n        imgList = pd.read_csv(\"/home/czy/dataset/human_protain/moredata/HPAv18RBGY_wodpl.csv\")\n        for i in tqdm(imgList['Id'][sp:ep], postfix=pid):  # [:5] means downloard only first 5 samples, if it works, please remove it\n            img = i.split('_')\n            for color in colors:\n                img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n                img_name = i + \"_\" + color + \".jpg\"\n                img_url = v18_url + img_path\n                r = requests.get(img_url, allow_redirects=True)\n                open(DIR + img_name, 'wb').write(r.content)\n    \n    def run_proc(name, sp, ep):\n        print('Run child process %s (%s) sp:%d ep: %d' % (name, os.getpid(), sp, ep))\n        download(name, sp, ep)\n        print('Run child process %s done' % (name))\n    \n    if __name__ == \"__main__\":\n        print('Parent process %s.' % os.getpid())\n        img_list = pd.read_csv(\"/home/czy/dataset/human_protain/moredata/HPAv18RBGY_wodpl.csv\")['Id']\n        list_len = len(img_list)\n        process_num = 100\n        p = Pool(process_num)\n        for i in range(process_num):\n            p.apply_async(run_proc, args=(str(i), int(i * list_len / process_num), int((i + 1) * list_len / process_num)))\n        print('Waiting for all subprocesses done...')\n        p.close()\n        p.join()\n        print('All subprocesses done.')",
          "votes": 13
        },
        {
          "id": 437054,
          "postDate": "2018-12-11T09:59:09.780Z",
          "content": "<p>Thank u. for sharing. \ntiny remind, it`s protein not protain.\nbtw, how long could u download all the image files using ur script ?  I am suffering for my slow speed, is it because the EDU NET?  </p>",
          "rawMarkdown": "Thank u. for sharing. \ntiny remind, it`s protein not protain.\nbtw, how long could u download all the image files using ur script ?  I am suffering for my slow speed, is it because the EDU NET?  "
        },
        {
          "id": 437119,
          "postDate": "2018-12-11T11:57:46.903Z",
          "content": "<p>thanks for your reminding. i'm also in the EDU NET, i use 100 threads which takes me about 4 hours to download all the 74k+ images.</p>",
          "rawMarkdown": "thanks for your reminding. i'm also in the EDU NET, i use 100 threads which takes me about 4 hours to download all the 74k+ images."
        },
        {
          "id": 437120,
          "postDate": "2018-12-11T11:58:23.203Z",
          "content": "<p>I also faced with the issue of 3 channels in .jpg vs 1 in .png files. Moreover, if I convert .jpg images to grayscale and then train my model with extended dataset - I got very low score. </p>",
          "rawMarkdown": "I also faced with the issue of 3 channels in .jpg vs 1 in .png files. Moreover, if I convert .jpg images to grayscale and then train my model with extended dataset - I got very low score. ",
          "votes": 2
        },
        {
          "id": 437122,
          "postDate": "2018-12-11T12:01:33.090Z",
          "content": "<p>and I'm in the IPv6 net</p>",
          "rawMarkdown": "and I'm in the IPv6 net"
        },
        {
          "id": 437167,
          "postDate": "2018-12-11T13:29:41.613Z",
          "content": "<p>hello，@Hogger, I run your code to download the HPA data，but it failed. So I change process_num from 100 to 5, but get following result:  could you help me? thank you.</p>\n\n<p>Parent process 6352.\nWaiting for all subprocesses done...\nRun child process 0 (6356) sp:0 ep: 14921\nRun child process 1 (6357) sp:14921 ep: 29842\nRun child process 2 (6358) sp:29842 ep: 44763\nRun child process 3 (6359) sp:44763 ep: 59684\nRun child process 4 (6360) sp:59684 ep: 74606\n  0%|                                                                                    | 0/14921 [00:00\n\n</p><p>All subprocesses done.</p>",
          "rawMarkdown": "hello，@Hogger, I run your code to download the HPA data，but it failed. So I change process_num from 100 to 5, but get following result:  could you help me? thank you.\n\nParent process 6352.\nWaiting for all subprocesses done...\nRun child process 0 (6356) sp:0 ep: 14921\nRun child process 1 (6357) sp:14921 ep: 29842\nRun child process 2 (6358) sp:29842 ep: 44763\nRun child process 3 (6359) sp:44763 ep: 59684\nRun child process 4 (6360) sp:59684 ep: 74606\n  0%|                                                                                    | 0/14921 [00:00"
        },
        {
          "id": 437203,
          "postDate": "2018-12-11T14:37:30.520Z",
          "content": "<p>@ManyFoldCV\n<code>\nfrom PIL import Image\nImage.open('img_HPAv18_some_name_blue.jpg').convert('L')\n</code>\nGive that a shot. The object returned has a <code>.save()</code> method and also can be converted to <code>ndarray</code> with <code>np.array(pil_image)</code>.</p>",
          "rawMarkdown": "@ManyFoldCV\n```\nfrom PIL import Image\nImage.open('img_HPAv18_some_name_blue.jpg').convert('L')\n```\nGive that a shot. The object returned has a `.save()` method and also can be converted to `ndarray` with `np.array(pil_image)`.",
          "votes": 2
        },
        {
          "id": 437244,
          "postDate": "2018-12-11T15:53:05.280Z",
          "content": "<p>these are the normal output. if it stuck at 0. Try to access a single image by your internet browser,  if you even can't access the image by your browser, maybe its your internet that makes u can't access it. </p>",
          "rawMarkdown": "these are the normal output. if it stuck at 0. Try to access a single image by your internet browser,  if you even can't access the image by your browser, maybe its your internet that makes u can't access it. "
        },
        {
          "id": 437318,
          "postDate": "2018-12-11T17:54:03.577Z",
          "content": "<p><a href=\"/davidwagnerkc\">@davidwagnerkc</a>\nMuch appreciated!</p>",
          "rawMarkdown": "@davidwagnerkc\nMuch appreciated!"
        },
        {
          "id": 437386,
          "postDate": "2018-12-11T20:16:31.047Z",
          "content": "<p>I have made a few small modifications to @Hogger code:\n- images are resized to a specified size: <code>image_size</code>\n- images are saved as grayscale\n- images are saved as PNG</p>\n\n<p>```\nimport os\nimport errno\nfrom multiprocessing.pool import Pool\nfrom tqdm import tqdm\nimport requests\nimport pandas as pd\nfrom PIL import Image</p>\n\n<p>def download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path</p>\n\n<pre><code>        # Get the raw response from the url\n        r = requests.get(img_url, allow_redirects=True, stream=True)\n        r.raw.decode_content = True\n\n        # Use PIL to resize the image and to convert it to L\n        # (8-bit pixels, black and white)\n        im = Image.open(r.raw)\n        im = im.resize(image_size, Image.LANCZOS).convert('L')\n        im.save(os.path.join(save_dir, img_name), 'PNG')\n</code></pre>\n\n<p>if <strong>name</strong> == '<strong>main</strong>':\n    # Parameters\n    process_num = 24\n    image_size = (512, 512)\n    url = '<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>'\n    csv_path =  \"path/to/csv/HPAv18RBGY_wodpl.csv\"\n    save_dir = \"where/the/images/are/saved\"</p>\n\n<pre><code># Create the directory to save the images in case it doesn't exist\ntry:\n    os.makedirs(save_dir)\nexcept OSError as exc:\n    if exc.errno != errno.EEXIST:\n        raise\n    pass\n\nprint('Parent process %s.' % os.getpid())\nimg_list = pd.read_csv(csv_path)['Id']\nlist_len = len(img_list)\np = Pool(process_num)\nfor i in range(process_num):\n    start = int(i * list_len / process_num)\n    end = int((i + 1) * list_len / process_num)\n    process_images = img_list[start:end]\n    p.apply_async(\n        download, args=(str(i), process_images, url, save_dir, image_size)\n    )\nprint('Waiting for all subprocesses done...')\np.close()\np.join()\nprint('All subprocesses done.')\n</code></pre>\n\n<p>```</p>\n\n<p>The script from @Hogger downloads 100 images in about 16 seconds; this script increases that time to 27 seconds, not that bad of a trade for some convenience.  On my computer tqdm says it should take around 4 hours total. I'm also CPU bottlenecked because I'm training at the same time, so a good internet connection and a free CPU can probably decrease this time significantly.</p>",
          "rawMarkdown": "I have made a few small modifications to @Hogger code:\n- images are resized to a specified size: ``image_size``\n- images are saved as grayscale\n- images are saved as PNG\n\n```\nimport os\nimport errno\nfrom multiprocessing.pool import Pool\nfrom tqdm import tqdm\nimport requests\nimport pandas as pd\nfrom PIL import Image\n\ndef download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path\n\n            # Get the raw response from the url\n            r = requests.get(img_url, allow_redirects=True, stream=True)\n            r.raw.decode_content = True\n\n            # Use PIL to resize the image and to convert it to L\n            # (8-bit pixels, black and white)\n            im = Image.open(r.raw)\n            im = im.resize(image_size, Image.LANCZOS).convert('L')\n            im.save(os.path.join(save_dir, img_name), 'PNG')\n\nif __name__ == '__main__':\n    # Parameters\n    process_num = 24\n    image_size = (512, 512)\n    url = 'http://v18.proteinatlas.org/images/'\n    csv_path =  \"path/to/csv/HPAv18RBGY_wodpl.csv\"\n    save_dir = \"where/the/images/are/saved\"\n\n    # Create the directory to save the images in case it doesn't exist\n    try:\n        os.makedirs(save_dir)\n    except OSError as exc:\n        if exc.errno != errno.EEXIST:\n            raise\n        pass\n\n    print('Parent process %s.' % os.getpid())\n    img_list = pd.read_csv(csv_path)['Id']\n    list_len = len(img_list)\n    p = Pool(process_num)\n    for i in range(process_num):\n        start = int(i * list_len / process_num)\n        end = int((i + 1) * list_len / process_num)\n        process_images = img_list[start:end]\n        p.apply_async(\n            download, args=(str(i), process_images, url, save_dir, image_size)\n        )\n    print('Waiting for all subprocesses done...')\n    p.close()\n    p.join()\n    print('All subprocesses done.')\n```\n\nThe script from @Hogger downloads 100 images in about 16 seconds; this script increases that time to 27 seconds, not that bad of a trade for some convenience.  On my computer tqdm says it should take around 4 hours total. I'm also CPU bottlenecked because I'm training at the same time, so a good internet connection and a free CPU can probably decrease this time significantly.",
          "votes": 10
        },
        {
          "id": 437516,
          "postDate": "2018-12-12T04:25:54.587Z",
          "content": "<p>great job!</p>",
          "rawMarkdown": "great job!"
        },
        {
          "id": 437563,
          "postDate": "2018-12-12T06:04:36.213Z",
          "content": "<p>I am now running Hogger's code with David's tweaks, download is much faster but file sizes are also much larger than with <a href=\"/tomomimoriyama\">@tomomimoriyama</a> code, look like total down load will be 250-260 GB. I commented out </p>\n\n<p>im = im.resize(image_size, Image.LANCZOS).convert('L')</p>\n\n<p>because I don't want gray scale images. I am curious as to why David Silva wants gray scale vs. color images.</p>\n\n<p>Kickback</p>",
          "rawMarkdown": "I am now running Hogger's code with David's tweaks, download is much faster but file sizes are also much larger than with @tomomimoriyama code, look like total down load will be 250-260 GB. I commented out \n\n im = im.resize(image_size, Image.LANCZOS).convert('L')\n\nbecause I don't want gray scale images. I am curious as to why David Silva wants gray scale vs. color images.\n\nKickback\n"
        },
        {
          "id": 437586,
          "postDate": "2018-12-12T06:44:53.357Z",
          "content": "<p>@kickback All 74k RGBY images takes 30GB disk space on my machine being resized to 512 x 512 RGB PNGs. Silva's code saves as grayscale because he is saving four separate images for each of the channel, in a similar manner to how the competition data has been provided.</p>",
          "rawMarkdown": "@kickback All 74k RGBY images takes 30GB disk space on my machine being resized to 512 x 512 RGB PNGs. Silva's code saves as grayscale because he is saving four separate images for each of the channel, in a similar manner to how the competition data has been provided.",
          "votes": 1
        },
        {
          "id": 437730,
          "postDate": "2018-12-12T11:43:04.167Z",
          "content": "<p>Hi <a href=\"/davidwagnerkc\">@davidwagnerkc</a>,</p>\n\n<p>Thanks for your script above though it's a little confusing following this thread. Does your script download RGB images as the url has <code>'red_green_yellow'</code> in the name?</p>\n\n<p>Thanks</p>",
          "rawMarkdown": "Hi @davidwagnerkc,\n\nThanks for your script above though it's a little confusing following this thread. Does your script download RGB images as the url has `'red_green_yellow'` in the name?\n\nThanks"
        },
        {
          "id": 437831,
          "postDate": "2018-12-12T15:30:12.987Z",
          "content": "<p>@Mark Worrall That might be a bug! Haha. I must have just forgot to add <em>blue</em> wherever that is supposed to go in the color order. I'll fix it and edit the post. Thanks!</p>\n\n<p>EDIT: I updated the code above if anybody is actually using it. </p>",
          "rawMarkdown": "@Mark Worrall That might be a bug! Haha. I must have just forgot to add _blue_ wherever that is supposed to go in the color order. I'll fix it and edit the post. Thanks!\n\nEDIT: I updated the code above if anybody is actually using it. "
        },
        {
          "id": 437848,
          "postDate": "2018-12-12T16:03:08.400Z",
          "content": "<p>No worries, I fixed it up to get each channel color individually anyway. Thank you for this!</p>",
          "rawMarkdown": "No worries, I fixed it up to get each channel color individually anyway. Thank you for this!"
        },
        {
          "id": 438010,
          "postDate": "2018-12-12T23:51:55.793Z",
          "content": "<p>I ran the code with the one line commented out and after 16 hours I have downloaded 292k items that consume 728 GB of space. I have no idea what I really have but some of the images look pretty interesting!</p>\n\n<p>Time for some resizing.......</p>",
          "rawMarkdown": "I ran the code with the one line commented out and after 16 hours I have downloaded 292k items that consume 728 GB of space. I have no idea what I really have but some of the images look pretty interesting!\n\nTime for some resizing......."
        },
        {
          "id": 438012,
          "postDate": "2018-12-12T23:56:58.443Z",
          "content": "<p>@ManyFoldCV\nThank you for your information that downloaded image has 3 channels.\nI apology for those who affected by my mistake.I'm very sorry.</p>\n\n<p><a href=\"/davidwagnerkc\">@davidwagnerkc</a>,@Hogger\nThank you for your great code!</p>\n\n<p>@Mark Worrall, @kickback\nto be Grayscale is very important!\nas <a href=\"/davidwagnerkc\">@davidwagnerkc</a> said [how the competition data has been provided]</p>\n\n<p>In my case, \nI've been using 3 channels image without noticing it,\nscinse I use Iafoss's kernel and it loads image like</p>\n\n<pre><code>def open_rgby(path,id): \n    colors = ['red','green','blue','yellow']\n    flags = cv2.IMREAD_GRAYSCALE\n    img = [cv2.imread(os.path.join(path, id+'_'+color+'.png'), flags).astype(np.float32)/255\n           for color in colors]\n    return np.stack(img, axis=-1)\n</code></pre>\n\n<p>Visualizing it make no difference from me..</p>\n\n<p>But, from yesterday,\nI'm trying @Alex 's SingleCell Extraction.\nI found rgb-png-converted one do harm but grayscale-converted is fine.\nI attached html file.\nOnly rgb-png-converted image, detected nuclei(right image) is not clear.\nResizing is not affected the result.</p>",
          "rawMarkdown": "@ManyFoldCV\nThank you for your information that downloaded image has 3 channels.\nI apology for those who affected by my mistake.I'm very sorry.\n\n@davidwagnerkc,@Hogger\nThank you for your great code!\n\n@Mark Worrall, @kickback\nto be Grayscale is very important!\nas @davidwagnerkc said [how the competition data has been provided]\n\nIn my case, \nI've been using 3 channels image without noticing it,\nscinse I use Iafoss's kernel and it loads image like\n\n    def open_rgby(path,id): \n        colors = ['red','green','blue','yellow']\n        flags = cv2.IMREAD_GRAYSCALE\n        img = [cv2.imread(os.path.join(path, id+'_'+color+'.png'), flags).astype(np.float32)/255\n               for color in colors]\n        return np.stack(img, axis=-1)\n\nVisualizing it make no difference from me..\n\nBut, from yesterday,\nI'm trying @Alex 's SingleCell Extraction.\nI found rgb-png-converted one do harm but grayscale-converted is fine.\nI attached html file.\nOnly rgb-png-converted image, detected nuclei(right image) is not clear.\nResizing is not affected the result.\n",
          "votes": 1
        },
        {
          "id": 438187,
          "postDate": "2018-12-13T08:26:25.890Z",
          "content": "<p>Hi Tomomi, </p>\n\n<p>Yes, so I convert each 3 channel red, green and blue image of a single id to grayscale using <code>.convert('L')</code>.</p>\n\n<p>Is there any other pre-processing you do? In particular there seems to be a large number of class 27 (&gt;100) which was making me wonder about the accuracy of the labels.</p>",
          "rawMarkdown": "Hi Tomomi, \n\nYes, so I convert each 3 channel red, green and blue image of a single id to grayscale using `.convert('L')`.\n\nIs there any other pre-processing you do? In particular there seems to be a large number of class 27 (&gt;100) which was making me wonder about the accuracy of the labels."
        },
        {
          "id": 438197,
          "postDate": "2018-12-13T08:49:21.400Z",
          "content": "<p>Hi there,</p>\n\n<p>I'm (also) still confused by the color channel information in the additional HPA data. When using grayscale conversion of each of the four images corresponding to a single id via convert('L'),  the problem is that the pixel statistics is drastically different compared to the original training set of the competition.</p>\n\n<p>For example, the blue channel in the original 512x512 training images on average has a pixel intensity of ~14, while the grayscaled version of the 'blue' images in the additional HPA data only has ~1.7 to ~2 (I haven't explored the full set yet). I fear that this spoils training...</p>\n\n<p>It would be simple if each image in the additional HPA data had non-zero pixels only in the \"right\" color channel (except for yellow, where things inevitably are a bit more complicated), but unfortunately that's not the case. Any ideas? </p>",
          "rawMarkdown": "Hi there,\n\nI'm (also) still confused by the color channel information in the additional HPA data. When using grayscale conversion of each of the four images corresponding to a single id via convert('L'),  the problem is that the pixel statistics is drastically different compared to the original training set of the competition.\n\nFor example, the blue channel in the original 512x512 training images on average has a pixel intensity of ~14, while the grayscaled version of the 'blue' images in the additional HPA data only has ~1.7 to ~2 (I haven't explored the full set yet). I fear that this spoils training...\n\nIt would be simple if each image in the additional HPA data had non-zero pixels only in the \"right\" color channel (except for yellow, where things inevitably are a bit more complicated), but unfortunately that's not the case. Any ideas? ",
          "votes": 3
        },
        {
          "id": 438281,
          "postDate": "2018-12-13T12:02:29.490Z",
          "content": "<p>Hi, @Mark</p>\n\n<p>About pre-processing\nI don't use any pre-processing, but recalculated mean and std\n# mean and std in of each channel in the train set\n# Iafoss original RGBY\n#stats = A([0.0804, 0.0526, 0.0547, 0.0827], [0.394, 0.321, 0.327, 0.399])\n# HPAv18 all RGBY (each image is colored RGB png image 512x512)\nstats = A([0.06734, 0.05087, 0.03266, 0.09257],[0.11997, 0.10335, 0.10124, 0.1574 ])\nI'm not confident about the stats I calculated, \nsince I tried to reproduce original-train-set stats, my result was different from others report.... </p>\n\n<p>About too match class 27\nI'll check my script right away.</p>\n\n<p>Hi, @PhysicsGuy89\n@Brian provides conversion ruby script that preserves original intensity!\n (He released 10 days ago, but I've not tried yet, I'm new to ruby...)\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206#431936\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206#431936</a></p>",
          "rawMarkdown": "Hi, @Mark\n\nAbout pre-processing\nI don't use any pre-processing, but recalculated mean and std\n\\# mean and std in of each channel in the train set\n\\# Iafoss original RGBY\n\\#stats = A([0.0804, 0.0526, 0.0547, 0.0827], [0.394, 0.321, 0.327, 0.399])\n\\# HPAv18 all RGBY (each image is colored RGB png image 512x512)\nstats = A([0.06734, 0.05087, 0.03266, 0.09257],[0.11997, 0.10335, 0.10124, 0.1574 ])\nI'm not confident about the stats I calculated, \nsince I tried to reproduce original-train-set stats, my result was different from others report.... \n\nAbout too match class 27\nI'll check my script right away.\n\nHi, @PhysicsGuy89\n@Brian provides conversion ruby script that preserves original intensity!\n (He released 10 days ago, but I've not tried yet, I'm new to ruby...)\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206#431936",
          "votes": 1
        },
        {
          "id": 438333,
          "postDate": "2018-12-13T13:47:04.933Z",
          "content": "<p>Just a quick note that you can't treat color channels in fluorescent images in the same way as standard digital images. Each color channel is actually acquired as a separate greyscale image, and researchers just stack and artificially color them to make it easier to see multiple structures of the cell at once. Standard greyscale conversion algorithms will take a weighted average of the color channels (including cv2's function), which is why you're seeing the blue intensities go down so much, since it's being multiplied by a modifier. The correct way is to just simply split the channels (and then tile them if you really need a three channel image) - so for example, in Python representing the images as a numpy array, you'd just go image[:, :, 0]... etc. </p>\n\n<p>For this reason, you should really download the yellow channel as a separate image. The RGB images have each color channel as an independent fluorescent dye, but the yellow channel will add green and blue intensities to these images that will confuse their splitting. It's all artificially colored - it's just that with four markers, you run out of independent color channels since you only have 3 max with RGB. </p>",
          "rawMarkdown": "Just a quick note that you can't treat color channels in fluorescent images in the same way as standard digital images. Each color channel is actually acquired as a separate greyscale image, and researchers just stack and artificially color them to make it easier to see multiple structures of the cell at once. Standard greyscale conversion algorithms will take a weighted average of the color channels (including cv2's function), which is why you're seeing the blue intensities go down so much, since it's being multiplied by a modifier. The correct way is to just simply split the channels (and then tile them if you really need a three channel image) - so for example, in Python representing the images as a numpy array, you'd just go image[:, :, 0]... etc. \n\nFor this reason, you should really download the yellow channel as a separate image. The RGB images have each color channel as an independent fluorescent dye, but the yellow channel will add green and blue intensities to these images that will confuse their splitting. It's all artificially colored - it's just that with four markers, you run out of independent color channels since you only have 3 max with RGB. ",
          "votes": 5
        },
        {
          "id": 438339,
          "postDate": "2018-12-13T13:59:08.430Z",
          "content": "<p>Usually color order is rgb so\nfor blue take image[:, :, 2]\nfor red take image[:, :, 0]\nfor green take image[:, :, 1]\nfor yellow you can take either image[:, :, 0] or image[:, :, 1]. Yellow is created by putting the intensity values in both r and g channel. Due to image compression (which is pretty bad for the HPA images by the way. Shame on whoever did this), r and g channel of yellow images will not be identical. I did some visual as well as numeric comparisons though and they are almost identical to the point where just using the r channel for these images should be fine (image[:, :, 0])</p>",
          "rawMarkdown": "Usually color order is rgb so\nfor blue take image[:, :, 2]\nfor red take image[:, :, 0]\nfor green take image[:, :, 1]\nfor yellow you can take either image[:, :, 0] or image[:, :, 1]. Yellow is created by putting the intensity values in both r and g channel. Due to image compression (which is pretty bad for the HPA images by the way. Shame on whoever did this), r and g channel of yellow images will not be identical. I did some visual as well as numeric comparisons though and they are almost identical to the point where just using the r channel for these images should be fine (image[:, :, 0])",
          "votes": 3
        },
        {
          "id": 438343,
          "postDate": "2018-12-13T14:06:32.797Z",
          "content": "<p>Hi Alex, thank you for this though I'm not sure I follow fully.</p>\n\n<p>We have from the protein atlas website images that are 'red', 'green', 'blue' and 'yellow' channels. Each of these 4 images themselves have 3 channels.</p>\n\n<p>The data provided as per the competition has a single grayscale image for each of red, green, blue and yellow. So some mapping has taken place to compress each competition channel image into a single grayscale image - is this just the standard weighting per: </p>\n\n<p>L = R * 299/1000 + G * 587/1000 + B * 114/1000</p>\n\n<p>Thanks again,</p>\n\n<p>Mark</p>",
          "rawMarkdown": "Hi Alex, thank you for this though I'm not sure I follow fully.\n\nWe have from the protein atlas website images that are 'red', 'green', 'blue' and 'yellow' channels. Each of these 4 images themselves have 3 channels.\n\nThe data provided as per the competition has a single grayscale image for each of red, green, blue and yellow. So some mapping has taken place to compress each competition channel image into a single grayscale image - is this just the standard weighting per: \n\nL = R * 299/1000 + G * 587/1000 + B * 114/1000\n\nThanks again,\n\nMark",
          "votes": 1
        },
        {
          "id": 438348,
          "postDate": "2018-12-13T14:20:37.677Z",
          "content": "<p>I'm not an organizer, so I can't be sure, but it's usually the other way around - the microscope will give you four greyscale images, and then you convert each of these images into a different color by only showing the intensity of the pixels in one or two channels</p>",
          "rawMarkdown": "I'm not an organizer, so I can't be sure, but it's usually the other way around - the microscope will give you four greyscale images, and then you convert each of these images into a different color by only showing the intensity of the pixels in one or two channels",
          "votes": 2
        },
        {
          "id": 438355,
          "postDate": "2018-12-13T14:32:16.253Z",
          "content": "<p>+1 to Alex Lu answer. You need to think about how the mapping was done from the filter response to rgb. Look at my comment above to see how to convert back. (image being a np.ndarray of shape (X, Y, 3))</p>",
          "rawMarkdown": "+1 to Alex Lu answer. You need to think about how the mapping was done from the filter response to rgb. Look at my comment above to see how to convert back. (image being a np.ndarray of shape (X, Y, 3))"
        },
        {
          "id": 438357,
          "postDate": "2018-12-13T14:33:27.693Z",
          "content": "<p>@Alex Lu: I fully agree, standard RGB --&gt; grayscale conversion doesn't make much sense here, as \"blue\" in the context of an HPA channel has nothing to do with the actual blue color in an RGB image.</p>\n\n<p>But as Mark is saying, this doesn't solve the problem: for the additional HPA data, we only have access to the .jpg images, which <strong>for each of the channels R, G, B, Y</strong> are three-dimensional RGB images. In other words: do you know how these jpg images were constructed from the original data (which presumably only contained a single channel of information for each of the R, G, B, Y data)? If we know this procedure, we can apply it backwards to get the actual single-channel data for each of the four channels.</p>",
          "rawMarkdown": "@Alex Lu: I fully agree, standard RGB --&gt; grayscale conversion doesn't make much sense here, as \"blue\" in the context of an HPA channel has nothing to do with the actual blue color in an RGB image.\n\nBut as Mark is saying, this doesn't solve the problem: for the additional HPA data, we only have access to the .jpg images, which **for each of the channels R, G, B, Y** are three-dimensional RGB images. In other words: do you know how these jpg images were constructed from the original data (which presumably only contained a single channel of information for each of the R, G, B, Y data)? If we know this procedure, we can apply it backwards to get the actual single-channel data for each of the four channels.",
          "votes": 1
        },
        {
          "id": 438360,
          "postDate": "2018-12-13T14:34:27.313Z",
          "content": "<p>Hi, @Mark\nI've checked class 27, total 122 samples. There's no Uncertain status in it.\nI cannot judge all the annotated labels are correct...\nand same of the sample I cannot find rod like shape...\nAttached html what I have checked.</p>",
          "rawMarkdown": "Hi, @Mark\nI've checked class 27, total 122 samples. There's no Uncertain status in it.\nI cannot judge all the annotated labels are correct...\nand same of the sample I cannot find rod like shape...\nAttached html what I have checked."
        },
        {
          "id": 438361,
          "postDate": "2018-12-13T14:36:50.403Z",
          "content": "<p>@FabianIsensee: yeah, that was my original idea as well. But then it's really confusing that e.g. the 'green' images in the additional HPA data contain non-zero pixels for R, G and B. \nDo you think that is an artefact of the jpg image compression as well?</p>",
          "rawMarkdown": "@FabianIsensee: yeah, that was my original idea as well. But then it's really confusing that e.g. the 'green' images in the additional HPA data contain non-zero pixels for R, G and B. \nDo you think that is an artefact of the jpg image compression as well?"
        },
        {
          "id": 438365,
          "postDate": "2018-12-13T14:43:22.383Z",
          "content": "<p>@PhysicsGuy89 I suspect this is the case but I don't know enough about jpg compression to be sure</p>",
          "rawMarkdown": "@PhysicsGuy89 I suspect this is the case but I don't know enough about jpg compression to be sure",
          "votes": 1
        },
        {
          "id": 438369,
          "postDate": "2018-12-13T14:47:31.827Z",
          "content": "<p>Like I said - it's just displaying the fluorescent intensities of the micrograph in one channel only. So to turn these back into greyscale images, just tile the one channel it corresponds into all channels. </p>\n\n<p>Values in other channels are unexpected and probably due to the jpg compression. But also, the intensities in the jpg HPA images have also been stretched out for display on their web-front end. Depending on how they acquired the original micrographs, the original images may not have been treated this way - for the images my lab handles, you don't do this in the original tiff files acquired from the microscope, because you lose information about the level of protein expression when you rescale the images. That being said, these are fluorescent antibodies and not GFP-tagged proteins like we work with (where we expect the fluorescence to be correlated with protein expression), so double check that their tiff files are acquired the same way. </p>",
          "rawMarkdown": "Like I said - it's just displaying the fluorescent intensities of the micrograph in one channel only. So to turn these back into greyscale images, just tile the one channel it corresponds into all channels. \n\nValues in other channels are unexpected and probably due to the jpg compression. But also, the intensities in the jpg HPA images have also been stretched out for display on their web-front end. Depending on how they acquired the original micrographs, the original images may not have been treated this way - for the images my lab handles, you don't do this in the original tiff files acquired from the microscope, because you lose information about the level of protein expression when you rescale the images. That being said, these are fluorescent antibodies and not GFP-tagged proteins like we work with (where we expect the fluorescence to be correlated with protein expression), so double check that their tiff files are acquired the same way. ",
          "votes": 2
        },
        {
          "id": 438374,
          "postDate": "2018-12-13T14:55:12.087Z",
          "content": "<p>This is getting confusing.</p>\n\n<p>I'm with <a href=\"/physicsguy89\">@physicsguy89</a> : it's not clear how they've gone from a 3 channel jpg for each channel to a single channel image for the competition data. </p>\n\n<p>@FabianIsensee are you saying that when we get a (for example) GREEN channel image from protein atlas website which has 3 channels we should just take image[:, :, 1]? And likewise for RBY extracting the corresponding channel?</p>\n\n<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a> - thank you for this though no html attached?</p>",
          "rawMarkdown": "This is getting confusing.\n\nI'm with @physicsguy89 : it's not clear how they've gone from a 3 channel jpg for each channel to a single channel image for the competition data. \n\n@FabianIsensee are you saying that when we get a (for example) GREEN channel image from protein atlas website which has 3 channels we should just take image[:, :, 1]? And likewise for RBY extracting the corresponding channel?\n\n@tomomimoriyama - thank you for this though no html attached?"
        },
        {
          "id": 438376,
          "postDate": "2018-12-13T15:00:43.797Z",
          "content": "<blockquote>\n  <p><strong>Alex Lu wrote</strong></p>\n  \n  <blockquote>\n    <p>Like I said - it's just displaying the fluorescent intensities of the micrograph in one channel only. So to turn these back into greyscale images, just tile the one channel it corresponds into all channels. </p>\n  </blockquote>\n</blockquote>\n\n<p>Hi Alex, thank you again. I'm almost there though still trying to fully parse what you said. When you say tile what do you mean? Are you saying (for example) take the 3 channel green image JPG from protein atlas and just use image[:, :, 1]? Or add up the values?</p>\n\n<p>Apologies for being slow on the uptake, I have ~0 domain knowledge here though have heard of protein and cells before. ;)</p>",
          "rawMarkdown": "\n&gt; **Alex Lu wrote**\n&gt; \n&gt; &gt; Like I said - it's just displaying the fluorescent intensities of the micrograph in one channel only. So to turn these back into greyscale images, just tile the one channel it corresponds into all channels. \n&gt; \n\nHi Alex, thank you again. I'm almost there though still trying to fully parse what you said. When you say tile what do you mean? Are you saying (for example) take the 3 channel green image JPG from protein atlas and just use image[:, :, 1]? Or add up the values?\n\nApologies for being slow on the uptake, I have ~0 domain knowledge here though have heard of protein and cells before. ;)"
        },
        {
          "id": 438378,
          "postDate": "2018-12-13T15:01:46.930Z",
          "content": "<p>@Mark Worrall this is exactly what I am implying and if I am not mistaken this is also what Alex Lu suggested</p>",
          "rawMarkdown": "@Mark Worrall this is exactly what I am implying and if I am not mistaken this is also what Alex Lu suggested"
        },
        {
          "id": 438379,
          "postDate": "2018-12-13T15:06:05.107Z",
          "content": "<p>Assuming you have the microtubule image:\ngreyscale[:, :, 0] = microtubule[:, :, 0]\ngreyscale[:, :, 1] = microtubule[:, :, 0]\ngreyscale[:, :, 2] = microtubule[:, :, 0]</p>\n\n<p>Repeat with the proper channel for nuclei/protein/ER images. </p>",
          "rawMarkdown": "Assuming you have the microtubule image:\ngreyscale[:, :, 0] = microtubule[:, :, 0]\ngreyscale[:, :, 1] = microtubule[:, :, 0]\ngreyscale[:, :, 2] = microtubule[:, :, 0]\n\nRepeat with the proper channel for nuclei/protein/ER images. ",
          "votes": 1
        },
        {
          "id": 438387,
          "postDate": "2018-12-13T15:20:12.997Z",
          "content": "<p>Hmmmmm, I'm still none the wiser, why are we repeating the values in each greyscale image? Greyscale images have a single channel.</p>",
          "rawMarkdown": "Hmmmmm, I'm still none the wiser, why are we repeating the values in each greyscale image? Greyscale images have a single channel."
        },
        {
          "id": 438390,
          "postDate": "2018-12-13T15:28:11.193Z",
          "content": "<p>Then you can do this:\ngreyscale = microtubule[:, :, 0]</p>",
          "rawMarkdown": "Then you can do this:\ngreyscale = microtubule[:, :, 0]"
        },
        {
          "id": 438393,
          "postDate": "2018-12-13T15:33:58.487Z",
          "content": "<p>Hi, @Mark\nattached file name is :Check Metadata for HPAv18.html (7.44 MB)\non the comment top.\nSince reply cell has no file-uploader...</p>",
          "rawMarkdown": "Hi, @Mark\nattached file name is :Check Metadata for HPAv18.html (7.44 MB)\non the comment top.\nSince reply cell has no file-uploader...",
          "votes": 1
        },
        {
          "id": 438405,
          "postDate": "2018-12-13T15:48:46.863Z",
          "content": "<p>How they processed the image is written this thread:\nsome questions about dataset instrumentation, image properties\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68983\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68983</a></p>\n\n<p>@Casper W[provide single slice images, non-merged. No gamma correction is made...]\nhe also said[If it still unclear, please do not hesitate to ask.]\nSo, let us ask about this problem on the thread? </p>",
          "rawMarkdown": "How they processed the image is written this thread:\nsome questions about dataset instrumentation, image properties\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68983\n\n@Casper W[provide single slice images, non-merged. No gamma correction is made...]\nhe also said[If it still unclear, please do not hesitate to ask.]\nSo, let us ask about this problem on the thread? "
        },
        {
          "id": 438460,
          "postDate": "2018-12-13T18:00:54.550Z",
          "content": "<p>Hi Alex, again :)</p>\n\n<p>Sorry to be a pain, but can we be super explicit, is the following what you mean:</p>\n\n<p>Pseudo algorithm:</p>\n\n<pre><code>def process_im(id):\n    hpa_image = np.zeros((512, 512, 3))\n    for i, c in enumerate(['red', 'green', 'blue']):\n        tmp_channel_im = download_from_hpa_site(id, c)  # returns a 3 channel JPG image \n        hpa_image[:,:,i] = tmp_channel_im[:,:,0]  # &lt;---- BIT I AM UNCLEAR ON\n    hpa_image.save()\n</code></pre>\n\n<p>Thanks in advance.</p>\n\n<p>Update: I think it should be <code>hpa_image[:,:,i] = tmp_channel_im[:,:,i]</code></p>",
          "rawMarkdown": "Hi Alex, again :)\n\nSorry to be a pain, but can we be super explicit, is the following what you mean:\n\nPseudo algorithm:\n\n    def process_im(id):\n        hpa_image = np.zeros((512, 512, 3))\n        for i, c in enumerate(['red', 'green', 'blue']):\n            tmp_channel_im = download_from_hpa_site(id, c)  # returns a 3 channel JPG image \n            hpa_image[:,:,i] = tmp_channel_im[:,:,0]  # &lt;---- BIT I AM UNCLEAR ON\n        hpa_image.save()\n\nThanks in advance.\n\nUpdate: I think it should be `hpa_image[:,:,i] = tmp_channel_im[:,:,i]`"
        },
        {
          "id": 438473,
          "postDate": "2018-12-13T18:53:37.703Z",
          "content": "<blockquote>\n  <p>Update: I think it should be hpa_image[:,:,i] = tmp_channel_im[:,:,i]</p>\n</blockquote>\n\n<p>Yes. That is what I meant. For yellow, use <code>tmp_channel_im[:,:,0]</code> (red channel) or <code>tmp_channel_im[:,:,1]</code> (green channel)</p>",
          "rawMarkdown": "&gt; Update: I think it should be hpa_image[:,:,i] = tmp_channel_im[:,:,i]\n\nYes. That is what I meant. For yellow, use `tmp_channel_im[:,:,0]` (red channel) or `tmp_channel_im[:,:,1]` (green channel)",
          "votes": 1
        },
        {
          "id": 440147,
          "postDate": "2018-12-17T06:08:20.577Z",
          "content": "<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a>,</p>\n\n<p>After 3 days of much pain I now understand why greyscale images are so important! I saw what happens if you don't use them, found the problem (after much searching down the wrong paths) and then noticed you had pointed me to the solution several days ago!</p>\n\n<p>Learning can be painful...</p>\n\n<p>tx</p>\n\n<p>Kickback</p>",
          "rawMarkdown": "@tomomimoriyama,\n\nAfter 3 days of much pain I now understand why greyscale images are so important! I saw what happens if you don't use them, found the problem (after much searching down the wrong paths) and then noticed you had pointed me to the solution several days ago!\n\nLearning can be painful...\n\ntx\n\nKickback",
          "votes": 1
        },
        {
          "id": 440881,
          "postDate": "2018-12-18T03:59:17.647Z",
          "content": "<p>I am still unsure about the correct processing of the external HPA database.\nwill this code work properly?</p>\n\n<p>img = cv2.imread(fn, cv2.IMREAD_GRAYSCALE)</p>\n\n<p>ofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"</p>\n\n<p>img = cv2.resize(img, (512,512))</p>\n\n<p>cv2.imwrite(ofn, img)</p>",
          "rawMarkdown": "I am still unsure about the correct processing of the external HPA database.\nwill this code work properly?\n\n\n img = cv2.imread(fn, cv2.IMREAD_GRAYSCALE)\n \nofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"\n\n img = cv2.resize(img, (512,512))\n\n cv2.imwrite(ofn, img)\n"
        },
        {
          "id": 441715,
          "postDate": "2018-12-19T00:23:46.287Z",
          "content": "<p>I think it will not work properly. You are reading in three channels and converting them all to a grayscale, using, I think, the L formula. You should be doing a straight conversion, channel by channel, according to the color. For yellow, use average of red and green:</p>\n\n<pre><code>        img_resized = np.array(img.resize((int(y), int(y)), Image.ANTIALIAS))   # PIL\n        img_gs = None\n        for j, c in enumerate(colors):\n            if color == c:\n                if j &lt; 3:\n                    img_gs = img_resized[:, :, j]\n                elif j == 3:\n                    img_gs = (img_resized[:, :, 1] + img_resized[:, :, 0]) / 2\n        img_gs = img_gs.astype(np.uint8)\n</code></pre>\n\n<p>One good QC is to use the images that are repeated normal test data and HPA data and compare the histograms with your HPA import - they won't be identical but should look similar.</p>",
          "rawMarkdown": "I think it will not work properly. You are reading in three channels and converting them all to a grayscale, using, I think, the L formula. You should be doing a straight conversion, channel by channel, according to the color. For yellow, use average of red and green:\n\n            img_resized = np.array(img.resize((int(y), int(y)), Image.ANTIALIAS))   # PIL\n            img_gs = None\n            for j, c in enumerate(colors):\n                if color == c:\n                    if j &lt; 3:\n                        img_gs = img_resized[:, :, j]\n                    elif j == 3:\n                        img_gs = (img_resized[:, :, 1] + img_resized[:, :, 0]) / 2\n            img_gs = img_gs.astype(np.uint8)\n\nOne good QC is to use the images that are repeated normal test data and HPA data and compare the histograms with your HPA import - they won't be identical but should look similar."
        },
        {
          "id": 441724,
          "postDate": "2018-12-19T00:54:31.800Z",
          "content": "<p>that was a good advice, to check the leak images. I see what you mean, although I will have to rearrange the code as I still save them as separate colours.</p>\n\n<p>Thank you!</p>",
          "rawMarkdown": "that was a good advice, to check the leak images. I see what you mean, although I will have to rearrange the code as I still save them as separate colours.\n\nThank you!"
        },
        {
          "id": 441770,
          "postDate": "2018-12-19T03:33:41.103Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 441773,
          "postDate": "2018-12-19T03:48:30.247Z",
          "content": "<p>When using opencv the default is BGR not RGB, so channel indexes above are incorrect. Red and blue are swapped.</p>",
          "rawMarkdown": "When using opencv the default is BGR not RGB, so channel indexes above are incorrect. Red and blue are swapped.",
          "votes": 1
        },
        {
          "id": 441783,
          "postDate": "2018-12-19T04:09:00.317Z",
          "content": "<p>Weird - I was using PIL. That's where the histograms help - in the images I looked at  the blue is distinctive.</p>",
          "rawMarkdown": "Weird - I was using PIL. That's where the histograms help - in the images I looked at  the blue is distinctive."
        },
        {
          "id": 441786,
          "postDate": "2018-12-19T04:14:09.963Z",
          "content": "<p>Thank you, fixed</p>",
          "rawMarkdown": "Thank you, fixed"
        },
        {
          "id": 441788,
          "postDate": "2018-12-19T04:19:59.513Z",
          "content": "<p>Because i am using image for each channel (don't ask me why, inertia i guess) I couldn't see this problem. I have compared, alas, green to green and they were very similar... unfortunate...</p>",
          "rawMarkdown": "Because i am using image for each channel (don't ask me why, inertia i guess) I couldn't see this problem. I have compared, alas, green to green and they were very similar... unfortunate..."
        },
        {
          "id": 441861,
          "postDate": "2018-12-19T07:15:00.017Z",
          "content": "<p>I must say my training looks about the same, so either the resnet50 doesn't care too much or I still have mistake somewhere.</p>",
          "rawMarkdown": "I must say my training looks about the same, so either the resnet50 doesn't care too much or I still have mistake somewhere."
        },
        {
          "id": 442360,
          "postDate": "2018-12-19T21:27:01.817Z",
          "content": "<p>I suppose, that blue should be img[:,:,0] and red should be img[:,:,2]. \nI iterated over same images of different channels and got following sum_along_channel/total_sum channel for \"blue\",\"green\",\"red\",\"yellow\" . So blue is mostly present in 0 channel, whereas red in 2 channel.\nPlease, correct me if I am wrong.</p>\n\n<p>[0.98039446 0.0079704  0.01163514]\n[0.10331002 0.79542022 0.10126977]\n[0.05202446 0.04258582 0.90538972]\n[0.0740666  0.46528988 0.46064352]\n&gt; <strong>FabianIsensee wrote</strong>\n&gt; \n&gt; &gt; Usually color order is rgb so\n&gt; for blue take image[:, :, 2]\n&gt; for red take image[:, :, 0]\n&gt; for green take image[:, :, 1]\n&gt; for yellow you can take either image[:, :, 0] or image[:, :, 1]. Yellow is created by putting the intensity values in both r and g channel. Due to image compression (which is pretty bad for the HPA images by the way. Shame on whoever did this), r and g channel of yellow images will not be identical. I did some visual as well as numeric comparisons though and they are almost identical to the point where just using the r channel for these images should be fine (image[:, :, 0])</p>",
          "rawMarkdown": "I suppose, that blue should be img[:,:,0] and red should be img[:,:,2]. \nI iterated over same images of different channels and got following sum_along_channel/total_sum channel for \"blue\",\"green\",\"red\",\"yellow\" . So blue is mostly present in 0 channel, whereas red in 2 channel.\nPlease, correct me if I am wrong.\n\n[0.98039446 0.0079704  0.01163514]\n[0.10331002 0.79542022 0.10126977]\n[0.05202446 0.04258582 0.90538972]\n[0.0740666  0.46528988 0.46064352]\n&gt; **FabianIsensee wrote**\n&gt; \n&gt; &gt; Usually color order is rgb so\n&gt; for blue take image[:, :, 2]\n&gt; for red take image[:, :, 0]\n&gt; for green take image[:, :, 1]\n&gt; for yellow you can take either image[:, :, 0] or image[:, :, 1]. Yellow is created by putting the intensity values in both r and g channel. Due to image compression (which is pretty bad for the HPA images by the way. Shame on whoever did this), r and g channel of yellow images will not be identical. I did some visual as well as numeric comparisons though and they are almost identical to the point where just using the r channel for these images should be fine (image[:, :, 0])"
        },
        {
          "id": 442601,
          "postDate": "2018-12-20T07:22:16.303Z",
          "content": "<p>Hi,\nnot quite sure how you get the channels ordered this way. For me all images are r - g - b so blue is in channel 2.</p>\n\n<p>In [1]: import matplotlib.pyplot as plt <br>\nIn [2]: a = plt.imread(\"9985_69_B8_2_blue.jpg\") <br>\nIn [3]: a.shape <br>\nOut[3]: (1728, 1728, 3)\nIn [4]: a[:, :, 0].mean() <br>\nOut[4]: 0.2910919817386831\nIn [5]: a[:, :, 1].mean() <br>\nOut[5]: 0.24670694819530178\nIn [6]: a[:, :, 2].mean() <br>\nOut[6]: 27.79723802940672</p>\n\n<p>In [7]: from PIL import Image <br>\nIn [8]: import numpy as np <br>\nIn [9]: a = np.array(Image.open(\"9985_69_B8_2_blue.jpg\")) <br>\nIn [10]: a[:, :, 0].mean() <br>\nOut[10]: 0.2910919817386831\nIn [11]: a[:, :, 1].mean() <br>\nOut[11]: 0.24670694819530178\nIn [12]: a[:, :, 2].mean() <br>\nOut[12]: 27.79723802940672</p>",
          "rawMarkdown": "Hi,\nnot quite sure how you get the channels ordered this way. For me all images are r - g - b so blue is in channel 2.\n\nIn [1]: import matplotlib.pyplot as plt                                         \nIn [2]: a = plt.imread(\"9985_69_B8_2_blue.jpg\")                                 \nIn [3]: a.shape                                                                 \nOut[3]: (1728, 1728, 3)\nIn [4]: a[:, :, 0].mean()                                                       \nOut[4]: 0.2910919817386831\nIn [5]: a[:, :, 1].mean()                                                       \nOut[5]: 0.24670694819530178\nIn [6]: a[:, :, 2].mean()                                                       \nOut[6]: 27.79723802940672\n\nIn [7]: from PIL import Image                                                   \nIn [8]: import numpy as np                                                      \nIn [9]: a = np.array(Image.open(\"9985_69_B8_2_blue.jpg\"))                       \nIn [10]: a[:, :, 0].mean()                                                      \nOut[10]: 0.2910919817386831\nIn [11]: a[:, :, 1].mean()                                                      \nOut[11]: 0.24670694819530178\nIn [12]: a[:, :, 2].mean()                                                      \nOut[12]: 27.79723802940672\n"
        },
        {
          "id": 442659,
          "postDate": "2018-12-20T09:41:50.687Z",
          "content": "<p>hi vaagn:\nI guess you use opencv? Just as what @amaia say: when using opencv the default is BGR not RGB.</p>",
          "rawMarkdown": "hi vaagn:\nI guess you use opencv? Just as what @amaia say: when using opencv the default is BGR not RGB."
        },
        {
          "id": 445568,
          "postDate": "2018-12-26T18:31:09.217Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        },
        {
          "id": 445706,
          "postDate": "2018-12-27T00:48:12.443Z",
          "content": "<p>@Hogger@David Silva\nWhen I use your code in  kernel,  error of  5G space limit  was reported. \n <a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a> \nBut the kernel  above downloaded more than 11G of data.\nSince my computer is not around now, I want to use kernel to download HPA data and train. \nHow can I use more than 5G space? \nIs there anyone here who can help me? </p>",
          "rawMarkdown": "@Hogger@David Silva\nWhen I use your code in  kernel,  error of  5G space limit  was reported. \n https://www.kaggle.com/artemtprv/load-external-data \nBut the kernel  above downloaded more than 11G of data.\nSince my computer is not around now, I want to use kernel to download HPA data and train. \nHow can I use more than 5G space? \nIs there anyone here who can help me? "
        },
        {
          "id": 460336,
          "postDate": "2019-01-23T13:40:58.337Z",
          "content": "<p>hi david thanks \ni used your script to download some 32 k images after i applied some filters on list. \nWhen i append this data to Training data and start training my validation loss goes very High. I  averaged the stats of HPA and competition data to do normalization. \nWhat m i  missing so my val loss is sky rocketing. ?</p>",
          "rawMarkdown": "hi david thanks \ni used your script to download some 32 k images after i applied some filters on list. \nWhen i append this data to Training data and start training my validation loss goes very High. I  averaged the stats of HPA and competition data to do normalization. \nWhat m i  missing so my val loss is sky rocketing. ?"
        },
        {
          "id": 464236,
          "postDate": "2019-01-31T12:33:31.280Z",
          "content": "<p>hi david is there any similar script to download Tiff files provided orgiginally at gcp storage bucket.\n<a href=\"https://console.cloud.google.com/storage/browser/kaggle-human-protein-atlas\">https://console.cloud.google.com/storage/browser/kaggle-human-protein-atlas</a></p>",
          "rawMarkdown": "hi david is there any similar script to download Tiff files provided orgiginally at gcp storage bucket.\nhttps://console.cloud.google.com/storage/browser/kaggle-human-protein-atlas"
        }
      ]
    },
    {
      "id": 428136,
      "postDate": "2018-11-26T20:20:35.807Z",
      "content": "<p>12760 images (800x800 RGB).\ndata: <a href=\"https://www.kaggle.com/artemtprv/external-data-for-protein-atlas\">https://www.kaggle.com/artemtprv/external-data-for-protein-atlas</a>\nnotebook:  <a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a></p>",
      "rawMarkdown": "12760 images (800x800 RGB).\ndata: https://www.kaggle.com/artemtprv/external-data-for-protein-atlas\nnotebook:  https://www.kaggle.com/artemtprv/load-external-data",
      "votes": 9,
      "replies": [
        {
          "id": 436775,
          "postDate": "2018-12-10T23:22:27.940Z",
          "content": "<p>I used your notebook to create a <a href=\"https://gist.github.com/rebryk/4320304588c55f2287540457d6830da2\">script</a> for multithreaded data downloading.</p>\n\n<p>Usage example: <br>\n<code>python external_data.py --path ~/datasets/external --n_threads 4 --batch_size 128</code></p>",
          "rawMarkdown": "I used your notebook to create a [script](https://gist.github.com/rebryk/4320304588c55f2287540457d6830da2) for multithreaded data downloading.\n\n\nUsage example: <br>\n```python external_data.py --path ~/datasets/external --n_threads 4 --batch_size 128```",
          "votes": 2
        },
        {
          "id": 445695,
          "postDate": "2018-12-27T00:21:14.850Z",
          "content": "<p>I found your data is 11.31 GB. But when I use kernel to download data, the   disk space is limited to 5G. More than 5G will make an error. </p>",
          "rawMarkdown": "I found your data is 11.31 GB. But when I use kernel to download data, the   disk space is limited to 5G. More than 5G will make an error. "
        }
      ]
    },
    {
      "id": 431345,
      "postDate": "2018-12-02T04:49:01.633Z",
      "content": "<p>HPAv18 contains \"Different ENSG id, but same image paths\"\nHere is \"noisy labels\" list</p>",
      "rawMarkdown": "HPAv18 contains \"Different ENSG id, but same image paths\"\nHere is \"noisy labels\" list",
      "votes": 5
    },
    {
      "id": 431541,
      "postDate": "2018-12-02T13:18:48.490Z",
      "content": "<p>Regarding the data leakage of the 126 images Brian mentions: is it safe to assume that they will be removed for the final evaluation?</p>",
      "rawMarkdown": "Regarding the data leakage of the 126 images Brian mentions: is it safe to assume that they will be removed for the final evaluation?",
      "votes": 6,
      "replies": [
        {
          "id": 431796,
          "postDate": "2018-12-02T23:05:06.687Z",
          "content": "<p>I agree with <a href=\"/maettes\">@maettes</a> .\nData leakage boosted my score 0.528 to 0.588.\nI'm not comfortable with my score,\nbecause it's nothing to with my models generalization ability.\nI's not fair, because not everyone access those data easily.</p>\n\n<p>I hope (HPA admin?) provide clean HPAv18 metadata, \nand make everyone can access more simple way.</p>\n\n<p>in that case, I hope metadata contains Cell type information.\nthis extra information enable us to train generative models \nthat focuses on protein location but ignores cell type. </p>",
          "rawMarkdown": "I agree with @maettes .\nData leakage boosted my score 0.528 to 0.588.\nI'm not comfortable with my score,\nbecause it's nothing to with my models generalization ability.\nI's not fair, because not everyone access those data easily.\n\nI hope (HPA admin?) provide clean HPAv18 metadata, \nand make everyone can access more simple way.\n\nin that case, I hope metadata contains Cell type information.\nthis extra information enable us to train generative models \nthat focuses on protein location but ignores cell type. ",
          "votes": 5
        },
        {
          "id": 431832,
          "postDate": "2018-12-03T01:31:27.647Z",
          "content": "<p>The 126 definitely help my score. I didn't check recently but at the time I jumped from .360 to .440</p>",
          "rawMarkdown": "The 126 definitely help my score. I didn't check recently but at the time I jumped from .360 to .440",
          "votes": 3
        },
        {
          "id": 433894,
          "postDate": "2018-12-05T16:11:07.433Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 413046,
      "postDate": "2018-10-31T07:50:46.417Z",
      "content": "<p>external data:</p>\n\n<p><a href=\"https://www.proteinatlas.org/learn/dictionary/cell\">https://www.proteinatlas.org/learn/dictionary/cell</a></p>\n\n<p><a href=\"https://www.proteinatlas.org/humancell\">https://www.proteinatlas.org/humancell</a>\n<a href=\"http://cytoconference.org/2017/Program/Image-Analysis-Challenge.aspx\">http://cytoconference.org/2017/Program/Image-Analysis-Challenge.aspx</a>\n<a href=\"https://www.allencell.org/\">https://www.allencell.org/</a></p>\n\n<p><a href=\"https://github.com/CellProfiling/pytorch_integrated_cell\">https://github.com/CellProfiling/pytorch_integrated_cell</a></p>\n\n<p><a href=\"http://hpa.scoreboard.czi.technology/challenge/1\">http://hpa.scoreboard.czi.technology/challenge/1</a></p>\n\n<p>pre-train mpodel</p>\n\n<p><a href=\"https://github.com/fyu/drn\">https://github.com/fyu/drn</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "rawMarkdown": "external data:\n\nhttps://www.proteinatlas.org/learn/dictionary/cell\n\nhttps://www.proteinatlas.org/humancell\nhttp://cytoconference.org/2017/Program/Image-Analysis-Challenge.aspx\nhttps://www.allencell.org/\n\nhttps://github.com/CellProfiling/pytorch_integrated_cell\n\n\nhttp://hpa.scoreboard.czi.technology/challenge/1\n\npre-train mpodel\n\nhttps://github.com/fyu/drn\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n",
      "votes": 6
    },
    {
      "id": 412194,
      "postDate": "2018-10-29T18:53:08.427Z",
      "content": "<p>Pretrained ResNets at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>Other models at: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "rawMarkdown": "Pretrained ResNets at: https://github.com/pytorch/vision\n\nOther models at: https://github.com/Cadene/pretrained-models.pytorch",
      "votes": 3
    },
    {
      "id": 419935,
      "postDate": "2018-11-12T19:36:38.847Z",
      "content": "<p>If there is extra data, they could just make that available and easily downloadable. Having more labeled images is a huge advantage. </p>",
      "rawMarkdown": "If there is extra data, they could just make that available and easily downloadable. Having more labeled images is a huge advantage. ",
      "votes": 4
    },
    {
      "id": 413329,
      "postDate": "2018-10-31T17:58:07.407Z",
      "content": "<p>Keras pretrained ResNet(18,34,50,101,152) models <a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a></p>",
      "rawMarkdown": "Keras pretrained ResNet(18,34,50,101,152) models https://github.com/qubvel/classification_models",
      "votes": 4,
      "replies": [
        {
          "id": 446858,
          "postDate": "2018-12-28T19:39:19.827Z",
          "content": "<p>Hello,\ndo we have any specification on input images such as size (299 * 299 * 3 ?) of normalization?</p>",
          "rawMarkdown": "Hello,\ndo we have any specification on input images such as size (299 * 299 * 3 ?) of normalization?"
        }
      ]
    },
    {
      "id": 449896,
      "postDate": "2019-01-03T23:15:59.080Z",
      "content": "<p>similar external data and pretrained models as other teams, external data from <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a> and pretrained models from <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a> <a href=\"https://github.com/osmr/imgclsmob\">https://github.com/osmr/imgclsmob</a> <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a> </p>",
      "rawMarkdown": "similar external data and pretrained models as other teams, external data from https://www.proteinatlas.org and pretrained models from https://github.com/pytorch/vision https://github.com/osmr/imgclsmob https://github.com/Cadene/pretrained-models.pytorch ",
      "votes": 1
    },
    {
      "id": 440641,
      "postDate": "2018-12-17T20:18:59.813Z",
      "content": "<p>Pretrained models from <a href=\"https://github.com/creafz/pytorch-cnn-finetune\">https://github.com/creafz/pytorch-cnn-finetune</a></p>\n\n<p>Thank you <a href=\"/creafz\">@creafz</a> for sharing great project.</p>",
      "rawMarkdown": "Pretrained models from https://github.com/creafz/pytorch-cnn-finetune\n\nThank you @creafz for sharing great project.",
      "votes": 1
    },
    {
      "id": 437252,
      "postDate": "2018-12-11T16:07:24.943Z",
      "content": "<p>External data : <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>\nMany thanks to Brian and TomomiMoriyama!!!</p>",
      "rawMarkdown": "External data : https://www.proteinatlas.org\nMany thanks to Brian and TomomiMoriyama!!!",
      "votes": 1
    },
    {
      "id": 429599,
      "postDate": "2018-11-29T03:56:25.360Z",
      "content": "<ul>\n<li><p>InceptionV3 with torchvision pretrained weights <a href=\"https://github.com/pytorch/vision/blob/master/torchvision/models/inception.py\">https://github.com/pytorch/vision/blob/master/torchvision/models/inception.py</a></p></li>\n<li><p>HPA additional dataset with props to @TomomiMoriyama for the helpful CSV files. </p></li>\n</ul>",
      "rawMarkdown": "+ InceptionV3 with torchvision pretrained weights https://github.com/pytorch/vision/blob/master/torchvision/models/inception.py\n\n+ HPA additional dataset with props to @TomomiMoriyama for the helpful CSV files. ",
      "votes": 1
    },
    {
      "id": 424588,
      "postDate": "2018-11-20T11:19:07.893Z",
      "content": "<p>using pretrained models from <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> and <a href=\"https://github.com/pytorch/vision/tree/master/torchvision/models\">https://github.com/pytorch/vision/tree/master/torchvision/models</a></p>",
      "rawMarkdown": "using pretrained models from https://github.com/fastai/fastai and https://github.com/pytorch/vision/tree/master/torchvision/models",
      "votes": 1
    },
    {
      "id": 423736,
      "postDate": "2018-11-19T00:34:24.413Z",
      "content": "<p>@Brian Not sure how interesting this is (as it's only ~1% of the test set), but the 126 test images you found on HPA v18 do have an over-representation of 16 (Cytokinetic bridge)... Maybe chance(?)<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/423736/10691/126_test.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "@Brian Not sure how interesting this is (as it's only ~1% of the test set), but the 126 test images you found on HPA v18 do have an over-representation of 16 (Cytokinetic bridge)... Maybe chance(?)![enter image description here][1]\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/423736/10691/126_test.png",
      "votes": 1,
      "replies": [
        {
          "id": 424317,
          "postDate": "2018-11-19T22:04:04.830Z",
          "content": "<p>I'm guessing it is chance. The overall distribution of the HPA images is similar to the training set. Within the HPA images that I downloaded. Anything less than a couple thousand examples will have a different distribution because of the 1000:1 class imbalance.</p>",
          "rawMarkdown": "I'm guessing it is chance. The overall distribution of the HPA images is similar to the training set. Within the HPA images that I downloaded. Anything less than a couple thousand examples will have a different distribution because of the 1000:1 class imbalance.",
          "votes": 2
        }
      ]
    },
    {
      "id": 412838,
      "postDate": "2018-10-30T21:25:49.237Z",
      "content": "<p>Pretrained models at <a href=\"http://keras.io/applications\">http://keras.io/applications</a></p>",
      "rawMarkdown": "Pretrained models at http://keras.io/applications",
      "votes": 1
    },
    {
      "id": 442592,
      "postDate": "2018-12-20T07:10:45.540Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
      "votes": 2,
      "replies": [
        {
          "id": 446367,
          "postDate": "2018-12-28T00:54:01.083Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 436838,
      "postDate": "2018-12-11T02:31:40.530Z",
      "content": "<p>BNInception with pretrained-models.pytorch weights <a href=\"https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/bninception.py\">https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/bninception.py</a></p>\n\n<p>HPA additional dataset with props to @TomomiMoriyama for the helpful CSV files.</p>",
      "rawMarkdown": "BNInception with pretrained-models.pytorch weights https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/bninception.py\n\nHPA additional dataset with props to @TomomiMoriyama for the helpful CSV files."
    },
    {
      "id": 412155,
      "postDate": "2018-10-29T17:24:43.847Z",
      "content": "<p>I am not using a pretrained model, but I am using external data. I can share this privately but do I need to share it in this thread as well?</p>",
      "rawMarkdown": "I am not using a pretrained model, but I am using external data. I can share this privately but do I need to share it in this thread as well?",
      "votes": 2,
      "replies": [
        {
          "id": 412177,
          "postDate": "2018-10-29T18:23:55.123Z",
          "content": "<p>Hello Brian - you do need to list the external data you are using in this thread.</p>",
          "rawMarkdown": "Hello Brian - you do need to list the external data you are using in this thread.",
          "votes": 1
        },
        {
          "id": 412217,
          "postDate": "2018-10-29T19:55:14.497Z",
          "content": "<p>Using v18 of protein atlas from here: <a href=\"https://www.proteinatlas.org/about/download\">https://www.proteinatlas.org/about/download</a></p>",
          "rawMarkdown": "Using v18 of protein atlas from here: https://www.proteinatlas.org/about/download",
          "votes": 5
        },
        {
          "id": 418884,
          "postDate": "2018-11-10T19:45:57.167Z",
          "content": "<p>Could you please give a hint on how to download these data? Either I am too stupid to find it or they are not publicly available?</p>\n\n<p>Edit: Sorry for bothering you. I just understood how. Man that is as cumbersome as it gets. Parsing XML files for ~12k genes to get these images... They could have made that easier...\nNow excuse me while I'm off writing a bunch of python code to download all these images... </p>",
          "rawMarkdown": "Could you please give a hint on how to download these data? Either I am too stupid to find it or they are not publicly available?\n\nEdit: Sorry for bothering you. I just understood how. Man that is as cumbersome as it gets. Parsing XML files for ~12k genes to get these images... They could have made that easier...\nNow excuse me while I'm off writing a bunch of python code to download all these images... "
        },
        {
          "id": 418905,
          "postDate": "2018-11-10T21:00:29.943Z",
          "content": "<p>Well since the cat is out of the bag...</p>\n\n<p>I downloaded the TSV file. Used the TSV file to download all the associated XML files (5GB). Concatenate all the xml files and wrote a stream parsing script to pull out the images and protein information, and then download it.</p>\n\n<p>I'm not going to make it too easy, python is not the language I know best. Even worse, I did it with the version of ruby that was install on my machine and not the latest, 1.9.</p>\n\n<p>This one downloads all xml from the tsv file:</p>\n\n<pre>require 'net/http'\n\ntext=File.open('subcellular_location.tsv').read\ntext.gsub!(/\\r\\n?/, \"\\n\")\ntext.each_line do |line|\n    split_line = line.split(\"\\t\")\n    puts \"LINE:#{split_line[0].to_s.strip}||#{split_line[3].to_s.strip}|#{split_line[4].to_s.strip}|#{split_line[5].to_s.strip}\"\n    Net::HTTP.start(\"www.proteinatlas.org\") do |http|\n        resp = http.get(\"/#{split_line[0].to_s.strip}.xml\")\n        open(\"xml/#{split_line[0].to_s.strip}.xml\", \"wb\") do |file|\n            file.write(resp.body)\n        end\n    end\nend\n</pre>\n\n<p>I used cat *.xml &gt; out.xml, which is consumed by this script:\n<a href=\"https://pastebin.com/hJDBB6D0\">https://pastebin.com/hJDBB6D0</a></p>\n\n<p><strong>Spoiler alert: 126 images of the test set can be found in the public HPA</strong></p>",
          "rawMarkdown": "Well since the cat is out of the bag...\n\nI downloaded the TSV file. Used the TSV file to download all the associated XML files (5GB). Concatenate all the xml files and wrote a stream parsing script to pull out the images and protein information, and then download it.\n\nI'm not going to make it too easy, python is not the language I know best. Even worse, I did it with the version of ruby that was install on my machine and not the latest, 1.9.\n\nThis one downloads all xml from the tsv file:\n<pre>require 'net/http'\n\ntext=File.open('subcellular_location.tsv').read\ntext.gsub!(/\\r\\n?/, \"\\n\")\ntext.each_line do |line|\n\tsplit_line = line.split(\"\\t\")\n\tputs \"LINE:#{split_line[0].to_s.strip}||#{split_line[3].to_s.strip}|#{split_line[4].to_s.strip}|#{split_line[5].to_s.strip}\"\n\tNet::HTTP.start(\"www.proteinatlas.org\") do |http|\n\t    resp = http.get(\"/#{split_line[0].to_s.strip}.xml\")\n\t    open(\"xml/#{split_line[0].to_s.strip}.xml\", \"wb\") do |file|\n\t        file.write(resp.body)\n\t    end\n\tend\nend\n</pre>\n\nI used cat *.xml &gt; out.xml, which is consumed by this script:\nhttps://pastebin.com/hJDBB6D0\n\n**Spoiler alert: 126 images of the test set can be found in the public HPA**",
          "votes": 21
        },
        {
          "id": 418909,
          "postDate": "2018-11-10T21:07:44.910Z",
          "content": "<p>Also, unless you change this to run in a multithreaded way, it will take about 2-3 days to run.</p>",
          "rawMarkdown": "Also, unless you change this to run in a multithreaded way, it will take about 2-3 days to run.",
          "votes": 2
        },
        {
          "id": 418929,
          "postDate": "2018-11-10T22:16:19.293Z",
          "content": "<p>Thank you very much! May I ask how many images are there for the minority classes?</p>",
          "rawMarkdown": "Thank you very much! May I ask how many images are there for the minority classes?\n"
        },
        {
          "id": 418943,
          "postDate": "2018-11-10T23:48:15.477Z",
          "content": "<p>I don't have an exact count off hand, but the distribution is about the same. Here is a comparison with competition data on the left:\n<img src=\"http://brians.network/images/sourcevsaugment.png\" alt=\"\"></p>",
          "rawMarkdown": "I don't have an exact count off hand, but the distribution is about the same. Here is a comparison with competition data on the left:\n![](http://brians.network/images/sourcevsaugment.png)",
          "votes": 8
        },
        {
          "id": 418959,
          "postDate": "2018-11-11T00:56:34.260Z",
          "content": "<p>Good to see that the training set is nearly doubled! The minority classes didn't get a lot more samples, which means that the label imbalance problem continues to exist.</p>",
          "rawMarkdown": "Good to see that the training set is nearly doubled! The minority classes didn't get a lot more samples, which means that the label imbalance problem continues to exist."
        },
        {
          "id": 418991,
          "postDate": "2018-11-11T03:19:45.030Z",
          "content": "<p>May I also ask if the end of the second script is missing?</p>",
          "rawMarkdown": "May I also ask if the end of the second script is missing?"
        },
        {
          "id": 419002,
          "postDate": "2018-11-11T04:09:10.747Z",
          "content": "<p>Looks like it got cut off,  I'll try to paste it again when I get back to my computer. </p>",
          "rawMarkdown": "Looks like it got cut off,  I'll try to paste it again when I get back to my computer. ",
          "votes": 1
        },
        {
          "id": 419025,
          "postDate": "2018-11-11T05:25:50.650Z",
          "content": "<p>Edited the post, it didn't like the xml tags that were being matched: <a href=\"https://pastebin.com/hJDBB6D0\">https://pastebin.com/hJDBB6D0</a></p>",
          "rawMarkdown": "Edited the post, it didn't like the xml tags that were being matched: https://pastebin.com/hJDBB6D0",
          "votes": 1
        },
        {
          "id": 419032,
          "postDate": "2018-11-11T05:42:07.050Z",
          "content": "<p>Thank you very much! I checked each xml file and the website, but it seems that there are multiple images for a specific file. And since I know nothing about ruby, may I ask what which image is the one we are looking for?</p>",
          "rawMarkdown": "Thank you very much! I checked each xml file and the website, but it seems that there are multiple images for a specific file. And since I know nothing about ruby, may I ask what which image is the one we are looking for?"
        },
        {
          "id": 419042,
          "postDate": "2018-11-11T06:01:11.147Z",
          "content": "<p>Wow, this forum doesn't like less than characters. Worked in the preview, and the whole comment disappeared. In quick summary, there are data nodes, those nodes have a location nodes containing the protein names. In the same data node there are urls that contain _red in them on their own line. Using regular expressions to match xml tags line by line instead of parsing the document as xml.</p>",
          "rawMarkdown": "Wow, this forum doesn't like less than characters. Worked in the preview, and the whole comment disappeared. In quick summary, there are data nodes, those nodes have a location nodes containing the protein names. In the same data node there are urls that contain _red in them on their own line. Using regular expressions to match xml tags line by line instead of parsing the document as xml.",
          "votes": 4
        },
        {
          "id": 420177,
          "postDate": "2018-11-13T07:50:39.407Z",
          "content": "<p>Thank you so much! How do you deal with the different verification types (uncertain, enhanced, approved)?</p>",
          "rawMarkdown": "Thank you so much! How do you deal with the different verification types (uncertain, enhanced, approved)?"
        },
        {
          "id": 420435,
          "postDate": "2018-11-13T16:17:18.377Z",
          "content": "<p>I drop any images that are labeled 'uncertain'</p>",
          "rawMarkdown": "I drop any images that are labeled 'uncertain'",
          "votes": 1
        },
        {
          "id": 420437,
          "postDate": "2018-11-13T16:18:49.633Z",
          "content": "<p>Thanks! Is that already included in your code?</p>",
          "rawMarkdown": "Thanks! Is that already included in your code?"
        },
        {
          "id": 420464,
          "postDate": "2018-11-13T16:59:16.423Z",
          "content": "<p>Not in this version. I had a version that was nicer, did multi threaded downloads and checked for uncertain. When I went looking for it, I had removed the wrong folder previously and only this older version was left :(</p>",
          "rawMarkdown": "Not in this version. I had a version that was nicer, did multi threaded downloads and checked for uncertain. When I went looking for it, I had removed the wrong folder previously and only this older version was left :(",
          "votes": 1
        },
        {
          "id": 420773,
          "postDate": "2018-11-14T05:10:17.460Z",
          "content": "<p>I assume that since the data is from the same source as the challenge data (assumption based on the names), that there is duplicate data. Did you find this to be the case? If so, how did you check for duplicate data. Thanks!</p>",
          "rawMarkdown": "I assume that since the data is from the same source as the challenge data (assumption based on the names), that there is duplicate data. Did you find this to be the case? If so, how did you check for duplicate data. Thanks!"
        },
        {
          "id": 420779,
          "postDate": "2018-11-14T05:24:31.887Z",
          "content": "<p>Yes, there is definitely duplicate data there. I used a program that is usually used for finding duplicate photos that would flag ones that were similar. Several hundred duplicates exist between the training set and the HPA images, you can also find answers to 126 of test entries in the HPA.</p>",
          "rawMarkdown": "Yes, there is definitely duplicate data there. I used a program that is usually used for finding duplicate photos that would flag ones that were similar. Several hundred duplicates exist between the training set and the HPA images, you can also find answers to 126 of test entries in the HPA.",
          "votes": 4
        },
        {
          "id": 423580,
          "postDate": "2018-11-18T15:48:14.247Z",
          "content": "<p>Do you know the class distribution for the 126 test images with labels that you found in this data? It might give us some idea whether the train/test distributions match, which could inform training strategies.</p>",
          "rawMarkdown": "Do you know the class distribution for the 126 test images with labels that you found in this data? It might give us some idea whether the train/test distributions match, which could inform training strategies."
        },
        {
          "id": 423719,
          "postDate": "2018-11-18T23:26:41.727Z",
          "content": "<p>Here is a ruby script that processes the xml file, skips 'uncertain', and downloads the images using 12 threads. Takes only a few hours: <a href=\"https://pastebin.com/zjvJH0ne\">https://pastebin.com/zjvJH0ne</a></p>\n\n<p><a href=\"/dstjhb\">@dstjhb</a> here are the values for the 126 that overlapped. I've not checked the distribution or anything:</p>\n\n<pre>'0 23'\n'14 16'\n'14 16'\n'0 16 25'\n'14 16'\n'19'\n'0 16'\n'16'\n'0 16 25'\n'5'\n'3'\n'16 25'\n'16 25'\n'0 16 25'\n'0 16'\n'4'\n'0 25'\n'5 16'\n'0 16'\n'16 25'\n'0 16 17 18'\n'2 11'\n'16 25'\n'14 16'\n'2 16'\n'0 16 17 18'\n'16 25'\n'16'\n'0 16'\n'2 16'\n'0 21'\n'5'\n'0 5'\n'0 16'\n'0 7'\n'0 21'\n'16 17 23'\n'5'\n'1'\n'17 25'\n'0 21'\n'0 16'\n'2 17'\n'17 25'\n'4'\n'17'\n'25'\n'25'\n'0 16'\n'9 10'\n'14'\n'0'\n'21 16'\n'2'\n'13'\n'4 21 26'\n'25'\n'13'\n'25'\n'4'\n'14'\n'14'\n'0'\n'21'\n'0'\n'16'\n'22 16 25'\n'21 16 19 25'\n'9 10'\n'14 16'\n'16 25'\n'2 16'\n'16'\n'7 17'\n'5'\n'15 25'\n'21 11 16'\n'12'\n'17'\n'17 25'\n'15'\n'4 21 17'\n'16 23'\n'4'\n'16'\n'17'\n'0 25'\n'0 19 25'\n'0 21'\n'16 25'\n'0 16'\n'1 2'\n'0 16 25'\n'23'\n'17 19'\n'22'\n'2'\n'16'\n'17 25'\n'0 16'\n'5'\n'0 14 18'\n'19'\n'7 25'\n'23'\n'12'\n'2 4'\n'16'\n'5'\n'0'\n'3'\n'16 17 23'\n'14 17 23'\n'0 25'\n'6 21'\n'0 16'\n'0 16 17'\n'0 16 25'\n'0 16 25'\n'14 16'\n'14 16'\n'14 16'\n'14 16 19'\n'14 16 25'\n'15 25'\n'15 25'\n</pre>",
          "rawMarkdown": "Here is a ruby script that processes the xml file, skips 'uncertain', and downloads the images using 12 threads. Takes only a few hours: https://pastebin.com/zjvJH0ne\n\n@dstjhb here are the values for the 126 that overlapped. I've not checked the distribution or anything:\n<pre>'0 23'\n'14 16'\n'14 16'\n'0 16 25'\n'14 16'\n'19'\n'0 16'\n'16'\n'0 16 25'\n'5'\n'3'\n'16 25'\n'16 25'\n'0 16 25'\n'0 16'\n'4'\n'0 25'\n'5 16'\n'0 16'\n'16 25'\n'0 16 17 18'\n'2 11'\n'16 25'\n'14 16'\n'2 16'\n'0 16 17 18'\n'16 25'\n'16'\n'0 16'\n'2 16'\n'0 21'\n'5'\n'0 5'\n'0 16'\n'0 7'\n'0 21'\n'16 17 23'\n'5'\n'1'\n'17 25'\n'0 21'\n'0 16'\n'2 17'\n'17 25'\n'4'\n'17'\n'25'\n'25'\n'0 16'\n'9 10'\n'14'\n'0'\n'21 16'\n'2'\n'13'\n'4 21 26'\n'25'\n'13'\n'25'\n'4'\n'14'\n'14'\n'0'\n'21'\n'0'\n'16'\n'22 16 25'\n'21 16 19 25'\n'9 10'\n'14 16'\n'16 25'\n'2 16'\n'16'\n'7 17'\n'5'\n'15 25'\n'21 11 16'\n'12'\n'17'\n'17 25'\n'15'\n'4 21 17'\n'16 23'\n'4'\n'16'\n'17'\n'0 25'\n'0 19 25'\n'0 21'\n'16 25'\n'0 16'\n'1 2'\n'0 16 25'\n'23'\n'17 19'\n'22'\n'2'\n'16'\n'17 25'\n'0 16'\n'5'\n'0 14 18'\n'19'\n'7 25'\n'23'\n'12'\n'2 4'\n'16'\n'5'\n'0'\n'3'\n'16 17 23'\n'14 17 23'\n'0 25'\n'6 21'\n'0 16'\n'0 16 17'\n'0 16 25'\n'0 16 25'\n'14 16'\n'14 16'\n'14 16'\n'14 16 19'\n'14 16 25'\n'15 25'\n'15 25'\n</pre>",
          "votes": 2
        },
        {
          "id": 423738,
          "postDate": "2018-11-19T00:37:53.747Z",
          "content": "<p>Thanks. I couldn't find an image upload in 'reply', but see my comment below for 126 test image distribution.</p>",
          "rawMarkdown": "Thanks. I couldn't find an image upload in 'reply', but see my comment below for 126 test image distribution.",
          "votes": 1
        },
        {
          "id": 426599,
          "postDate": "2018-11-23T14:12:44.340Z",
          "content": "<p>load external data\n<a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a></p>",
          "rawMarkdown": "load external data\nhttps://www.kaggle.com/artemtprv/load-external-data",
          "votes": 8
        },
        {
          "id": 426979,
          "postDate": "2018-11-24T09:42:21.933Z",
          "content": "<p>Am I right, that there are maximum 800x800 images available?\nWhat do you do, if you have several images for one ID, like here <a href=\"https://www.proteinatlas.org/ENSG00000002834-LASP1/antibody#ICC\">https://www.proteinatlas.org/ENSG00000002834-LASP1/antibody#ICC</a> or here <a href=\"https://www.proteinatlas.org/ENSG00000134057-CCNB1/antibody\">https://www.proteinatlas.org/ENSG00000134057-CCNB1/antibody</a> ?</p>",
          "rawMarkdown": "Am I right, that there are maximum 800x800 images available?\nWhat do you do, if you have several images for one ID, like here https://www.proteinatlas.org/ENSG00000002834-LASP1/antibody#ICC or here https://www.proteinatlas.org/ENSG00000134057-CCNB1/antibody ?"
        },
        {
          "id": 427055,
          "postDate": "2018-11-24T12:50:12.957Z",
          "content": "<p>Yes, images 800x800 RGB.\nI fixed this, now all images will load.(save like this ENSG00000001084-GCLC 0.png, ENSG00000001084-GCLC 1.png)</p>",
          "rawMarkdown": "Yes, images 800x800 RGB.\nI fixed this, now all images will load.(save like this ENSG00000001084-GCLC 0.png, ENSG00000001084-GCLC 1.png)",
          "votes": 1
        },
        {
          "id": 427539,
          "postDate": "2018-11-25T18:19:42.943Z",
          "content": "<p>Sorry to come on this a bit late... I noticed that in the HPA data there seem to be 32 classes, compared to the 28 we have in the competition. How do you deal with this mismatch?</p>",
          "rawMarkdown": "Sorry to come on this a bit late... I noticed that in the HPA data there seem to be 32 classes, compared to the 28 we have in the competition. How do you deal with this mismatch?"
        },
        {
          "id": 427623,
          "postDate": "2018-11-25T21:36:18.027Z",
          "content": "<p>I removed classes that don't exist, and removed any images that only contain classes that don't exist. You could define 'Nucleus' as the same as Nucleoplasm, but that might not be technically correct... You don't lose too many images filtering either way; so I guess see what works best for you.</p>",
          "rawMarkdown": "I removed classes that don't exist, and removed any images that only contain classes that don't exist. You could define 'Nucleus' as the same as Nucleoplasm, but that might not be technically correct... You don't lose too many images filtering either way; so I guess see what works best for you."
        },
        {
          "id": 427949,
          "postDate": "2018-11-26T13:04:22.347Z",
          "content": "<p>A word on \"uncertain images\": </p>\n\n<p>In the case of \"Uncertain\" fluorescent images: they are only \"uncertain\" because they are not corroborated by other experimental data (e.g. gene expression).</p>\n\n<p>The IF images state: \"Location not consistent with experimental gene/protein characterization data.\" in these cases. The probe localisation itself is actually correct (just the antibody may be staining the wrong thing). </p>\n\n<p>Therefore, for the purpose of simply classifying the stain location, there is no need to exclude \"uncertain\" images (in my opinion). You still have to be careful, as some classes may be visible in IHC but not fluorescence data.</p>",
          "rawMarkdown": "A word on \"uncertain images\": \n\nIn the case of \"Uncertain\" fluorescent images: they are only \"uncertain\" because they are not corroborated by other experimental data (e.g. gene expression).\n\nThe IF images state: \"Location not consistent with experimental gene/protein characterization data.\" in these cases. The probe localisation itself is actually correct (just the antibody may be staining the wrong thing). \n\nTherefore, for the purpose of simply classifying the stain location, there is no need to exclude \"uncertain\" images (in my opinion). You still have to be careful, as some classes may be visible in IHC but not fluorescence data.",
          "votes": 2
        },
        {
          "id": 428107,
          "postDate": "2018-11-26T19:09:31.570Z",
          "content": "<p>The images I downloaded were not 800x800, but closer to 2048x2048 in jpg format. I saved only files that matched to the 28 classes, I didn't notice that there were 32. I've dropped all uncertain classes anyways just out of caution, maybe 15-20% of the extra images were removed when I did this.</p>",
          "rawMarkdown": "The images I downloaded were not 800x800, but closer to 2048x2048 in jpg format. I saved only files that matched to the 28 classes, I didn't notice that there were 32. I've dropped all uncertain classes anyways just out of caution, maybe 15-20% of the extra images were removed when I did this."
        },
        {
          "id": 428685,
          "postDate": "2018-11-27T18:04:16.690Z",
          "content": "<p>Artem's webscraped ones are also either cropped or at a different magnification to original train set. Simpler to download though, as you don't have to parse the XML files for the links to the huge 2048p images.</p>",
          "rawMarkdown": "Artem's webscraped ones are also either cropped or at a different magnification to original train set. Simpler to download though, as you don't have to parse the XML files for the links to the huge 2048p images.",
          "votes": 1
        },
        {
          "id": 430093,
          "postDate": "2018-11-29T20:22:02.787Z",
          "content": "<p>Hi Brian. I understand the images are single files with 3 channels. Did you find a way to use these additional images with a 4-Channel model (that you might be using for the official dataset)?</p>",
          "rawMarkdown": "Hi Brian. I understand the images are single files with 3 channels. Did you find a way to use these additional images with a 4-Channel model (that you might be using for the official dataset)?"
        },
        {
          "id": 430133,
          "postDate": "2018-11-29T22:04:20.640Z",
          "content": "<p>I'm only using rgb from the competition data,  combined into one file to match the HPA data.</p>",
          "rawMarkdown": "I'm only using rgb from the competition data,  combined into one file to match the HPA data.",
          "votes": 2
        },
        {
          "id": 430828,
          "postDate": "2018-12-01T02:31:31.607Z",
          "content": "<blockquote>\n  <p>4-Channel  RGBY</p>\n</blockquote>\n\n<p>imageURL.replace(blue_red_green,color)</p>\n\n<p>e.g.\n<a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_blue_red_green.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_blue_red_green.jpg</a></p>\n\n<p><a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_red.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_red.jpg</a>\n<a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_green.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_green.jpg</a>\n<a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_blue.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_blue.jpg</a></p>\n\n<p><a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_yellow.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_yellow.jpg</a></p>",
          "rawMarkdown": "&gt;4-Channel  RGBY\n\nimageURL.replace(blue_red_green,color)\n\ne.g.\nhttp://v18.proteinatlas.org/images/4109/24_H11_2_blue_red_green.jpg\n\nhttp://v18.proteinatlas.org/images/4109/24_H11_2_red.jpg\nhttp://v18.proteinatlas.org/images/4109/24_H11_2_green.jpg\nhttp://v18.proteinatlas.org/images/4109/24_H11_2_blue.jpg\n\nhttp://v18.proteinatlas.org/images/4109/24_H11_2_yellow.jpg",
          "votes": 4
        },
        {
          "id": 431363,
          "postDate": "2018-12-02T05:21:23.927Z",
          "content": "<p>Thanks for this, I'm using it now to repeat my previous tests with the Y channel included. </p>",
          "rawMarkdown": "Thanks for this, I'm using it now to repeat my previous tests with the Y channel included. "
        }
      ]
    },
    {
      "id": 455529,
      "postDate": "2019-01-14T05:52:15.300Z",
      "content": "<p>@Hogger\nUse your code for multi-threaded download, download will not download in the middle, is it necessary to use VPN?</p>",
      "rawMarkdown": "@Hogger\nUse your code for multi-threaded download, download will not download in the middle, is it necessary to use VPN?"
    },
    {
      "id": 454786,
      "postDate": "2019-01-12T07:51:20.430Z",
      "content": "<p>Fastai pretrained models: <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: http://www.proteinatlas.org/"
    },
    {
      "id": 453947,
      "postDate": "2019-01-11T01:13:11.663Z",
      "content": "<p>External data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a>\npre-trained models:<a href=\"http://download.tensorflow.org/models/inception_v4_2016_09_09.tar.gz\">http://download.tensorflow.org/models/inception_v4_2016_09_09.tar.gz</a></p>",
      "rawMarkdown": "External data: http://www.proteinatlas.org/\npre-trained models:http://download.tensorflow.org/models/inception_v4_2016_09_09.tar.gz"
    },
    {
      "id": 453646,
      "postDate": "2019-01-10T14:28:58.073Z",
      "content": "<p>data：kaggle data(512×512) and hpa data (<a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a>)\npretrainedmodels: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "rawMarkdown": "data：kaggle data(512×512) and hpa data (http://www.proteinatlas.org/)\npretrainedmodels: https://github.com/Cadene/pretrained-models.pytorch"
    },
    {
      "id": 453380,
      "postDate": "2019-01-10T05:12:20.783Z",
      "content": "<p><a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "https://github.com/pytorch/vision\nExternal data: http://www.proteinatlas.org/"
    },
    {
      "id": 453042,
      "postDate": "2019-01-09T15:17:21.483Z",
      "content": "<p>As everyone else. But just in case:\n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "As everyone else. But just in case:\nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nExternal data: http://www.proteinatlas.org/"
    },
    {
      "id": 452722,
      "postDate": "2019-01-09T04:48:31.720Z",
      "content": "<p><a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n<a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\n<a href=\"https://gluon-cv.mxnet.io/\">https://gluon-cv.mxnet.io/</a></p>",
      "rawMarkdown": "https://www.proteinatlas.org/\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://keras.io/applications/\nhttps://gluon-cv.mxnet.io/"
    },
    {
      "id": 452460,
      "postDate": "2019-01-08T18:53:26.563Z",
      "content": "<p>Fastai pretrained models: <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nExternal data: <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: https://www.proteinatlas.org/"
    },
    {
      "id": 452155,
      "postDate": "2019-01-08T09:30:32.580Z",
      "content": "<p>External Data:  <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a> \nPretrained Models:  <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>    <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "rawMarkdown": "External Data:  https://www.proteinatlas.org \nPretrained Models:  https://github.com/pytorch/vision    https://github.com/Cadene/pretrained-models.pytorch"
    },
    {
      "id": 451993,
      "postDate": "2019-01-08T02:33:32.857Z",
      "content": "<p>Fastai pretrained models: <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nExternal data: <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: https://www.proteinatlas.org/"
    },
    {
      "id": 451900,
      "postDate": "2019-01-07T21:42:34.840Z",
      "content": "<p><a href=\"https://github.com/CellProfiling/FeatureExtraction\">https://github.com/CellProfiling/FeatureExtraction</a>\n<a href=\"https://github.com/CellProfiling/Loc-CAT\">https://github.com/CellProfiling/Loc-CAT</a></p>",
      "rawMarkdown": "https://github.com/CellProfiling/FeatureExtraction\nhttps://github.com/CellProfiling/Loc-CAT"
    },
    {
      "id": 451548,
      "postDate": "2019-01-07T08:54:25.097Z",
      "content": "<p>Pretrained PyramidNet <a href=\"https://github.com/dyhan0920/PyramidNet-PyTorch\">https://github.com/dyhan0920/PyramidNet-PyTorch</a></p>",
      "rawMarkdown": "Pretrained PyramidNet https://github.com/dyhan0920/PyramidNet-PyTorch"
    },
    {
      "id": 451206,
      "postDate": "2019-01-06T15:30:49.047Z",
      "content": "<p>Pretrained models from <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> and <a href=\"https://github.com/pytorch/vision/tree/master/torchvision/models\">https://github.com/pytorch/vision/tree/master/torchvision/models</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models from https://github.com/fastai/fastai and https://github.com/pytorch/vision/tree/master/torchvision/models\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 450948,
      "postDate": "2019-01-06T05:56:19.520Z",
      "content": "<p>Pretrained model\n<a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases\">https://github.com/fchollet/deep-learning-models/releases</a></p>\n\n<p>External data\n<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained model\nhttps://github.com/qubvel/classification_models\nhttps://github.com/fchollet/deep-learning-models/releases\n\nExternal data\nhttps://www.proteinatlas.org"
    },
    {
      "id": 450737,
      "postDate": "2019-01-05T15:49:47.833Z",
      "content": "<p>Apart from normal PyTorch pretrained models, pretrained models from Cadene and external data from HPA, I might use CBAM-ResNet50 from <a href=\"https://github.com/Jongchan/attention-module\">https://github.com/Jongchan/attention-module</a></p>",
      "rawMarkdown": "Apart from normal PyTorch pretrained models, pretrained models from Cadene and external data from HPA, I might use CBAM-ResNet50 from https://github.com/Jongchan/attention-module"
    },
    {
      "id": 450696,
      "postDate": "2019-01-05T14:45:13.327Z",
      "content": "<p>Pretrained models from: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data from: <a href=\"http://v18.proteinatlas.org/images\">http://v18.proteinatlas.org/images</a></p>",
      "rawMarkdown": "Pretrained models from: https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data from: http://v18.proteinatlas.org/images"
    },
    {
      "id": 450436,
      "postDate": "2019-01-05T00:12:43.790Z",
      "content": "<p>Pretrained:</p>\n\n<p><a href=\"https://github.com/pytorch/vision/\">https://github.com/pytorch/vision/</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data (planned)</p>\n\n<p><a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Pretrained:\n\nhttps://github.com/pytorch/vision/\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data (planned)\n\nhttps://www.proteinatlas.org/"
    },
    {
      "id": 450391,
      "postDate": "2019-01-04T20:52:11.367Z",
      "content": "<p>Fastai pretrained models: <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nExternal data: <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: https://www.proteinatlas.org/"
    },
    {
      "id": 450390,
      "postDate": "2019-01-04T20:51:58.380Z",
      "content": "<p>All of the below. </p>",
      "rawMarkdown": "All of the below. "
    },
    {
      "id": 450379,
      "postDate": "2019-01-04T20:19:58.660Z",
      "content": "<p><a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "https://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nExternal data: http://www.proteinatlas.org/\n",
      "replies": [
        {
          "id": 454341,
          "postDate": "2019-01-11T13:26:08.757Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 450297,
      "postDate": "2019-01-04T16:41:50.627Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a> and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nExternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision and https://github.com/Cadene/pretrained-models.pytorch\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 449913,
      "postDate": "2019-01-03T23:51:28.860Z",
      "content": "<p><a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>\n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n<a href=\"https://keras.io/applications/\">https://keras.io/applications/</a></p>",
      "rawMarkdown": "https://www.proteinatlas.org\nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://keras.io/applications/"
    },
    {
      "id": 449851,
      "postDate": "2019-01-03T21:29:41.407Z",
      "content": "<p><a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n<a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a>\n<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "https://keras.io/applications/\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://github.com/qubvel/classification_models\nhttps://www.proteinatlas.org",
      "replies": [
        {
          "id": 453327,
          "postDate": "2019-01-10T02:39:35.073Z",
          "content": "<p><a href=\"https://github.com/jacobkie/2018DSB\">https://github.com/jacobkie/2018DSB</a></p>",
          "rawMarkdown": "https://github.com/jacobkie/2018DSB"
        }
      ]
    },
    {
      "id": 449780,
      "postDate": "2019-01-03T18:20:51.990Z",
      "content": "<p>Pretrained models : \n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n<a href=\"https://keras.io/applications/\">https://keras.io/applications/</a></p>\n\n<p>External data:\n<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : \nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://keras.io/applications/\n\nExternal data:\nhttps://www.proteinatlas.org\n"
    },
    {
      "id": 449611,
      "postDate": "2019-01-03T12:33:35.867Z",
      "content": "<p>Pretrained models : \n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data:\n<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : \nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data:\nhttps://www.proteinatlas.org"
    },
    {
      "id": 449475,
      "postDate": "2019-01-03T07:47:33.943Z",
      "content": "<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>\nPretrained models:  <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>",
      "rawMarkdown": "External data: https://www.proteinatlas.org\nPretrained models:  https://github.com/pytorch/vision"
    },
    {
      "id": 449465,
      "postDate": "2019-01-03T07:24:52.427Z",
      "content": "<p>Using pretrained models from <a href=\"http://files.fast.ai/models/\">http://files.fast.ai/models/</a> and external data from <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Using pretrained models from http://files.fast.ai/models/ and external data from https://www.proteinatlas.org/"
    },
    {
      "id": 449448,
      "postDate": "2019-01-03T06:46:53.527Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 449364,
      "postDate": "2019-01-03T03:03:02.757Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 449240,
      "postDate": "2019-01-02T21:03:02.200Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 449046,
      "postDate": "2019-01-02T15:22:18.290Z",
      "content": "<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "External data: https://www.proteinatlas.org"
    },
    {
      "id": 448890,
      "postDate": "2019-01-02T10:24:50.720Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 448558,
      "postDate": "2019-01-01T13:42:43.490Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>\npretrained models: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\nexternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>. Credits to TomomiMoriyama (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860</a>) and David Silva (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386</a>)</p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org\npretrained models: https://keras.io/applications/\nexternal data: https://www.proteinatlas.org. Credits to TomomiMoriyama (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860) and David Silva (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386)"
    },
    {
      "id": 448282,
      "postDate": "2018-12-31T17:16:20.847Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 447940,
      "postDate": "2018-12-30T21:53:19.453Z",
      "content": "<p>I'm using the kernel as provided here, with some modifications:\n<a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb</a>\nSo resnet 34 and fastai. </p>\n\n<p>Along with the external data obtained using the script given by hogger here:\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>\n\n<p>Rest is some scripts written by me for processing of data, sorting and whatnot.</p>",
      "rawMarkdown": "I'm using the kernel as provided here, with some modifications:\nhttps://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\nSo resnet 34 and fastai. \n\nAlong with the external data obtained using the script given by hogger here:\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\n\nRest is some scripts written by me for processing of data, sorting and whatnot."
    },
    {
      "id": 447643,
      "postDate": "2018-12-30T08:14:16.417Z",
      "content": "<p>pretrained models: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\nexternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>. Credits to TomomiMoriyama (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860</a>) and David Silva (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386</a>)</p>",
      "rawMarkdown": "pretrained models: https://keras.io/applications/\nexternal data: https://www.proteinatlas.org. Credits to TomomiMoriyama (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860) and David Silva (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386)"
    },
    {
      "id": 447642,
      "postDate": "2018-12-30T08:14:12.130Z",
      "content": "<p>Pretrained models at: <a href=\"https://pypi.org/project/pretrainedmodels/\">https://pypi.org/project/pretrainedmodels/</a>  and  <a href=\"https://pypi.org/project/cnn-finetune/0.1.5/\">https://pypi.org/project/cnn-finetune/0.1.5/</a></p>\n\n<p>external data : <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://pypi.org/project/pretrainedmodels/  and  https://pypi.org/project/cnn-finetune/0.1.5/\n\nexternal data : https://www.proteinatlas.org"
    },
    {
      "id": 447288,
      "postDate": "2018-12-29T15:06:39.120Z",
      "content": "<p>Pretrained models:\n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data:\n<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models:\nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data:\nhttps://www.proteinatlas.org"
    },
    {
      "id": 446865,
      "postDate": "2018-12-28T19:57:04.433Z",
      "content": "<p><a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a></p>",
      "rawMarkdown": "https://github.com/qubvel/classification_models"
    },
    {
      "id": 446652,
      "postDate": "2018-12-28T12:41:57.367Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 446647,
      "postDate": "2018-12-28T12:36:00.247Z",
      "content": "<p>Pretrained models :\n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>External data: \n<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models :\nhttps://github.com/pytorch/vision\n\nExternal data: \nhttps://www.proteinatlas.org"
    },
    {
      "id": 446506,
      "postDate": "2018-12-28T07:50:11.960Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\nExternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 446492,
      "postDate": "2018-12-28T07:13:56.100Z",
      "content": "<p>Pretrained ResNets/Densenet/InceptionV3 at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\nExploring other models at: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nexternal data from <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained ResNets/Densenet/InceptionV3 at: https://github.com/pytorch/vision\nExploring other models at: https://github.com/Cadene/pretrained-models.pytorch\nexternal data from https://www.proteinatlas.org"
    },
    {
      "id": 446127,
      "postDate": "2018-12-27T15:02:36.783Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 445384,
      "postDate": "2018-12-26T10:34:12.667Z",
      "content": "<p>Pretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 445298,
      "postDate": "2018-12-26T06:55:44.227Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 445200,
      "postDate": "2018-12-25T23:26:21.883Z",
      "content": "<p>Pretrained models from <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\n<a href=\"https://pytorch.org/docs/stable/torchvision/models.html\">https://pytorch.org/docs/stable/torchvision/models.html</a>\nexternal data from: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models from https://github.com/fastai/fastai\nhttps://pytorch.org/docs/stable/torchvision/models.html\nexternal data from: https://www.proteinatlas.org"
    },
    {
      "id": 444979,
      "postDate": "2018-12-25T09:00:22.767Z",
      "content": "<p>External data from: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "External data from: https://www.proteinatlas.org"
    },
    {
      "id": 444695,
      "postDate": "2018-12-24T15:35:49.270Z",
      "content": "<p>Using pre-trained models (inceptionV4 and resnext50) from <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>  and additional external data from  <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Using pre-trained models (inceptionV4 and resnext50) from https://github.com/Cadene/pretrained-models.pytorch  and additional external data from  https://www.proteinatlas.org"
    },
    {
      "id": 444433,
      "postDate": "2018-12-24T02:55:49.387Z",
      "content": "<p>Has anyone uploaded the extra HPA RGBY images to Kaggle or other fast cloud storage to avoid traffic loads on HPA servers?</p>",
      "rawMarkdown": "Has anyone uploaded the extra HPA RGBY images to Kaggle or other fast cloud storage to avoid traffic loads on HPA servers?"
    },
    {
      "id": 444432,
      "postDate": "2018-12-24T02:51:38.017Z",
      "content": "<p>Pretrained models: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a> &amp; <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nExternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a> (as shared by @TomomiMoriyama)</p>",
      "rawMarkdown": "Pretrained models: https://github.com/pytorch/vision &amp; https://github.com/Cadene/pretrained-models.pytorch\nExternal data: https://www.proteinatlas.org (as shared by @TomomiMoriyama)"
    },
    {
      "id": 444202,
      "postDate": "2018-12-23T14:06:09.720Z",
      "content": "<p>as everyone..\nPretrained models : <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "as everyone..\nPretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 444142,
      "postDate": "2018-12-23T10:30:07.367Z",
      "content": "<p>pytorch models: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>, <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nexternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "pytorch models: https://github.com/pytorch/vision, https://github.com/Cadene/pretrained-models.pytorch\nexternal data: https://www.proteinatlas.org"
    },
    {
      "id": 443886,
      "postDate": "2018-12-22T16:55:41.823Z",
      "content": "<p>Pretrained models from <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> </p>",
      "rawMarkdown": "Pretrained models from https://github.com/fastai/fastai "
    },
    {
      "id": 443177,
      "postDate": "2018-12-21T06:46:02.577Z",
      "content": "<p>Pretrained models from <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> \nand\n<a href=\"https://github.com/pytorch/vision/\">https://github.com/pytorch/vision/</a></p>\n\n<p>External data from HPAv18</p>",
      "rawMarkdown": "Pretrained models from https://github.com/fastai/fastai \nand\nhttps://github.com/pytorch/vision/\n\nExternal data from HPAv18"
    },
    {
      "id": 442796,
      "postDate": "2018-12-20T14:14:40.243Z",
      "content": "<p>Hi everyone:\nI found there are two version of external data, which one should I use?\n1. HPA v18 share by @TomomiMoriyama\n2. 12760 images (800x800 RGB) shared by @Artem Toporov in this kernel: <a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a>\nMaybe I am wrong, maybe these two are the same data. I did not look at these two data carefully, I just found that their image names are different. I hope someone can answer me.</p>",
      "rawMarkdown": "Hi everyone:\nI found there are two version of external data, which one should I use?\n1. HPA v18 share by @TomomiMoriyama\n2. 12760 images (800x800 RGB) shared by @Artem Toporov in this kernel: https://www.kaggle.com/artemtprv/load-external-data\nMaybe I am wrong, maybe these two are the same data. I did not look at these two data carefully, I just found that their image names are different. I hope someone can answer me.",
      "replies": [
        {
          "id": 442859,
          "postDate": "2018-12-20T16:02:06.590Z",
          "content": "<p>#1 is the way to go. #2 looked like they were zoomed in or something. .They are also 800 x 800 instead of 2000+ width and height. </p>",
          "rawMarkdown": "\\#1 is the way to go. #2 looked like they were zoomed in or something. .They are also 800 x 800 instead of 2000+ width and height. "
        }
      ]
    },
    {
      "id": 442519,
      "postDate": "2018-12-20T03:53:38.693Z",
      "content": "<p>Pretrained ResNets at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\nOther models at: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nexternal data from  <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained ResNets at: https://github.com/pytorch/vision\nOther models at: https://github.com/Cadene/pretrained-models.pytorch\nexternal data from  https://www.proteinatlas.org"
    },
    {
      "id": 441687,
      "postDate": "2018-12-18T23:35:01.597Z",
      "content": "<p>I am using keras and pytorch fastai pre-trained models. </p>",
      "rawMarkdown": "I am using keras and pytorch fastai pre-trained models. "
    },
    {
      "id": 440442,
      "postDate": "2018-12-17T14:46:55.490Z",
      "content": "<ul>\n<li>pretrained pytorch models from the pretrainedmodels library</li>\n<li>additional HPA data from <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>, with the CSV files provided by @TomomiMoriyama </li>\n</ul>",
      "rawMarkdown": " - pretrained pytorch models from the pretrainedmodels library\n - additional HPA data from https://www.proteinatlas.org, with the CSV files provided by @TomomiMoriyama "
    },
    {
      "id": 440396,
      "postDate": "2018-12-17T13:55:13.043Z",
      "content": "<p>I'm using pretrained resnet pytorch model's weights via fastai library interface:\n<a href=\"https://pytorch.org/docs/stable/torchvision/models.html\">https://pytorch.org/docs/stable/torchvision/models.html</a>\n<a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nextra data:<a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "I'm using pretrained resnet pytorch model's weights via fastai library interface:\nhttps://pytorch.org/docs/stable/torchvision/models.html\nhttps://github.com/fastai/fastai\nextra data:https://www.proteinatlas.org"
    },
    {
      "id": 440270,
      "postDate": "2018-12-17T10:07:12.387Z",
      "content": "<p>External data at <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>\nPretrain Model: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a></p>",
      "rawMarkdown": "External data at https://www.proteinatlas.org\nPretrain Model: https://keras.io/applications/"
    },
    {
      "id": 440002,
      "postDate": "2018-12-16T20:55:48.930Z",
      "content": "<p>Keras pretrained models: <a href=\"http://keras.io/applications\">http://keras.io/applications</a>\nExternal data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Keras pretrained models: http://keras.io/applications\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 439290,
      "postDate": "2018-12-15T05:10:22.600Z",
      "content": "<p>Pretrained models at <a href=\"http://models.tensorpack.com/FasterRCNN/\">http://models.tensorpack.com/FasterRCNN/</a>\nExternal data at <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at http://models.tensorpack.com/FasterRCNN/\nExternal data at https://www.proteinatlas.org"
    },
    {
      "id": 438736,
      "postDate": "2018-12-14T05:27:09.140Z",
      "content": "<p>Pretrained models at <a href=\"https://github.com/chainer/chainercv\">https://github.com/chainer/chainercv</a> <br>\nExternal data at <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at https://github.com/chainer/chainercv  \nExternal data at https://www.proteinatlas.org"
    },
    {
      "id": 436927,
      "postDate": "2018-12-11T05:26:34.473Z",
      "content": "<p>Pretrained models at <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nExternal data at  <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at https://github.com/Cadene/pretrained-models.pytorch\nExternal data at  https://www.proteinatlas.org"
    },
    {
      "id": 436713,
      "postDate": "2018-12-10T19:59:54.090Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org"
    },
    {
      "id": 436115,
      "postDate": "2018-12-09T16:15:06.167Z",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and  <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data : <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "Pretrained models at: https://github.com/pytorch/vision\n\nand  https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data : https://www.proteinatlas.org"
    },
    {
      "id": 435756,
      "postDate": "2018-12-08T17:41:57.750Z",
      "content": "<p>pretrained models from \n<a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> \nand\n<a href=\"https://github.com/pytorch/vision/\">https://github.com/pytorch/vision/</a></p>\n\n<p>external data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "pretrained models from \nhttps://github.com/fastai/fastai \nand\nhttps://github.com/pytorch/vision/\n\nexternal data: https://www.proteinatlas.org"
    },
    {
      "id": 435432,
      "postDate": "2018-12-08T03:03:41.247Z",
      "content": "<p>Keras pretrained models (<a href=\"http://keras.io/applications\">http://keras.io/applications</a>)\nExternal data <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Keras pretrained models (http://keras.io/applications)\nExternal data https://www.proteinatlas.org/"
    },
    {
      "id": 435087,
      "postDate": "2018-12-07T13:27:54.343Z",
      "content": "<p>Keras pretrained models (<a href=\"http://keras.io/applications\">http://keras.io/applications</a>)\nExternal data <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Keras pretrained models (http://keras.io/applications)\nExternal data https://www.proteinatlas.org/"
    },
    {
      "id": 435057,
      "postDate": "2018-12-07T12:29:59.940Z",
      "content": "<p>Using pretrained Resnet weights from fast.ai and Keras. Other than that I am also using other pretrained weights from architectures like xception and DenseNets on keras.</p>",
      "rawMarkdown": "Using pretrained Resnet weights from fast.ai and Keras. Other than that I am also using other pretrained weights from architectures like xception and DenseNets on keras."
    },
    {
      "id": 434820,
      "postDate": "2018-12-07T02:07:43.670Z",
      "content": "<p>Keras pretrained models on ImageNet data (<a href=\"http://keras.io/applications\">http://keras.io/applications</a>)</p>",
      "rawMarkdown": "Keras pretrained models on ImageNet data (http://keras.io/applications)"
    },
    {
      "id": 433913,
      "postDate": "2018-12-05T16:40:21.123Z",
      "content": "<p>Keras pretrained models (<a href=\"http://keras.io/applications\">http://keras.io/applications</a>)\nProbably will use <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Keras pretrained models (http://keras.io/applications)\nProbably will use https://www.proteinatlas.org/\n\n"
    },
    {
      "id": 432899,
      "postDate": "2018-12-04T12:49:47.657Z",
      "content": "<p>Are kaggle or the organisers going to comment on the data leak?</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395</a></p>",
      "rawMarkdown": "Are kaggle or the organisers going to comment on the data leak?\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395",
      "replies": [
        {
          "id": 432974,
          "postDate": "2018-12-04T14:34:24.907Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 432626,
      "postDate": "2018-12-04T05:40:17.853Z",
      "content": "<p>I am looking at the site and in particular at the example url:\n<a href=\"https://www.proteinatlas.org/ENSG00000134057.xml\">https://www.proteinatlas.org/ENSG00000134057.xml</a></p>\n\n<p>Looking in this file, I see 32 mentions of _red_green_blue and two mentions of _red_green_blue_yellow.</p>\n\n<p>All the mentions appear to be 3 channel, so I don't know how to find yellow channels. Also, which of the multiple channels are correct.</p>\n\n<p>Perhaps more to the point: am I looking at the correct xml and how does one interpret it?</p>",
      "rawMarkdown": "I am looking at the site and in particular at the example url:\nhttps://www.proteinatlas.org/ENSG00000134057.xml\n\nLooking in this file, I see 32 mentions of _red_green_blue and two mentions of _red_green_blue_yellow.\n\nAll the mentions appear to be 3 channel, so I don't know how to find yellow channels. Also, which of the multiple channels are correct.\n\nPerhaps more to the point: am I looking at the correct xml and how does one interpret it?\n",
      "replies": [
        {
          "id": 432849,
          "postDate": "2018-12-04T12:02:58.420Z",
          "content": "<p>Hi, @pete</p>\n\n<p>Your are looking at correct xml.</p>\n\n<p>xml = urlopen('<a href=\"https://www.proteinatlas.org/ENSG00000134057.xml\">https://www.proteinatlas.org/ENSG00000134057.xml</a>' )\ntree = etree.parse(xml)\nimageUrls = tree.xpath('//cellExpression/<strong>subAssay</strong>/data/assayImage/image/imageUrl/text()')</p>\n\n<p>imageUrls contains multiple red_green_blue.jpg</p>\n\n<p>You can replace red_green_blue with red or green or blue or yellow</p>\n\n<p>Be aware that some samples has no yellow filters.</p>",
          "rawMarkdown": "Hi, @pete\n\nYour are looking at correct xml.\n\nxml = urlopen('https://www.proteinatlas.org/ENSG00000134057.xml' )\ntree = etree.parse(xml)\nimageUrls = tree.xpath('//cellExpression/**subAssay**/data/assayImage/image/imageUrl/text()')\n\nimageUrls contains multiple red\\_green\\_blue.jpg\n\nYou can replace red\\_green\\_blue with red or green or blue or yellow\n\nBe aware that some samples has no yellow filters.",
          "votes": 2
        },
        {
          "id": 432873,
          "postDate": "2018-12-04T12:19:29.443Z",
          "content": "<p>You should not extract\n _red_green_blue_yellow.jpg\non\n//cellExpression/<strong>customAssay</strong>\nits <strong>color mapping is different</strong> from our competitions data!</p>",
          "rawMarkdown": "You should not extract\n _red\\_green\\_blue\\_yellow.jpg\non\n//cellExpression/**customAssay**\nits **color mapping is different** from our competitions data!"
        },
        {
          "id": 433089,
          "postDate": "2018-12-04T17:19:41.240Z",
          "content": "<p>@TomomiMoriyama</p>\n\n<p>Are the color mappings for your HPAv18RBGY_wodpl.csv the same as our competition data? Thanks for sharing your work. </p>",
          "rawMarkdown": "@TomomiMoriyama\n\nAre the color mappings for your HPAv18RBGY_wodpl.csv the same as our competition data? Thanks for sharing your work. "
        },
        {
          "id": 433116,
          "postDate": "2018-12-04T18:06:59.557Z",
          "content": "<p>Hi, @David Wagner</p>\n\n<p>Yes.  The list extracted from <strong>subAssay</strong>.\n<a href=\"https://www.proteinatlas.org/about/help#17\">https://www.proteinatlas.org/about/help#17</a>\nIt reads \"The image link for subcellular images resides in the element for each antibody in the subAssay/data/cellLine and <strong>subAssay</strong>/data/assayImage/imageUrl elements. The image in the element is named after which channels that are toggled on and may include <strong>nucleus (blue), microtubuli (red), antibody (green) and endoplasmatic reticulum (yellow)</strong>, in this order. \"</p>\n\n<p>You cannot tell which sample yellow filter exists or not from xml:</p>\n\n<pre><code>&lt;image imageType=\"sampleImage\"&gt;&lt;channel color=\"blue\"&gt;nucleus&lt;/channel&gt;&lt;channel color=\"red\"&gt;microtubules&lt;/channel&gt;&lt;channel color=\"green\"&gt;antibody&lt;/channel&gt;&lt;imageUrl&gt;\n</code></pre>\n\n<p><a href=\"https://v18.proteinatlas.org/images/115/672_E2_1_yellow.jpg\">https://v18.proteinatlas.org/images/115/672_E2_1_yellow.jpg</a></p>\n\n<p>If your model only use RGB, please try HPAv18*<em>RGB</em>*_wodpl.csv(74,606 sample)\ninstead of  HPAv18*<em>RBGY</em>*_wodpl.csv(75,040 sample) which contains more sample</p>",
          "rawMarkdown": "Hi, @David Wagner\n\nYes.  The list extracted from **subAssay**.\nhttps://www.proteinatlas.org/about/help#17\nIt reads \"The image link for subcellular images resides in the element for each antibody in the subAssay/data/cellLine and **subAssay**/data/assayImage/imageUrl elements. The image in the element is named after which channels that are toggled on and may include **nucleus (blue), microtubuli (red), antibody (green) and endoplasmatic reticulum (yellow)**, in this order. \"\n\nYou cannot tell which sample yellow filter exists or not from xml:\n\n    <img>",
          "votes": 1
        },
        {
          "id": 433219,
          "postDate": "2018-12-04T21:25:34.490Z",
          "content": "<p>Thanks! I did not notice the subAssay/customAssay difference. </p>",
          "rawMarkdown": "Thanks! I did not notice the subAssay/customAssay difference. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 431868,
      "postDate": "2018-12-03T03:31:08.630Z",
      "content": "<p>pretrained models from \n<a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> \nand\n<a href=\"https://github.com/pytorch/vision/\">https://github.com/pytorch/vision/</a></p>\n\n<p>Hi Martin\nCan I use GAN to augment the train data??</p>",
      "rawMarkdown": "\npretrained models from \nhttps://github.com/fastai/fastai \nand\nhttps://github.com/pytorch/vision/\n\nHi Martin\nCan I use GAN to augment the train data??"
    },
    {
      "id": 430414,
      "postDate": "2018-11-30T11:05:07.473Z",
      "content": "<p>Pretrained:\n- <a href=\"https://github.com/pytorch/vision/\">https://github.com/pytorch/vision/</a>\n- <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>Plan to add external data:\n- <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "rawMarkdown": "Pretrained:\n- https://github.com/pytorch/vision/\n- https://github.com/Cadene/pretrained-models.pytorch\n\nPlan to add external data:\n- https://www.proteinatlas.org/"
    },
    {
      "id": 429650,
      "postDate": "2018-11-29T06:03:43.110Z",
      "content": "<p>I'm using keras MobileNet: <a href=\"https://keras.io/applications/#mobilenet\">https://keras.io/applications/#mobilenet</a>\nkernel: <a href=\"https://www.kaggle.com/amneves/keras-proteins\">https://www.kaggle.com/amneves/keras-proteins</a></p>",
      "rawMarkdown": "I'm using keras MobileNet: https://keras.io/applications/#mobilenet\nkernel: https://www.kaggle.com/amneves/keras-proteins"
    },
    {
      "id": 429322,
      "postDate": "2018-11-28T17:24:25.533Z",
      "content": "<p>external data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "rawMarkdown": "external data: https://www.proteinatlas.org"
    },
    {
      "id": 427131,
      "postDate": "2018-11-24T17:16:02.290Z",
      "content": "<p>Pretrained ResNets_34 at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>",
      "rawMarkdown": "Pretrained ResNets_34 at: https://github.com/pytorch/vision"
    },
    {
      "id": 426702,
      "postDate": "2018-11-23T17:12:37.893Z",
      "content": "<p>I am new to this, need some direction. There are 4 channels (each with different information - blue,red,yellow with organelle location information, and green with location of proteins) . The goal as I understand is to identify an image with the location of the protein.  If I was to combine the images and then train, can I also combine the images in the test set to validate and find accuracy of the model. In other words can the model be built that will only predict on images where all channels are combined into a single image ?</p>",
      "rawMarkdown": "I am new to this, need some direction. There are 4 channels (each with different information - blue,red,yellow with organelle location information, and green with location of proteins) . The goal as I understand is to identify an image with the location of the protein.  If I was to combine the images and then train, can I also combine the images in the test set to validate and find accuracy of the model. In other words can the model be built that will only predict on images where all channels are combined into a single image ?"
    },
    {
      "id": 423701,
      "postDate": "2018-11-18T22:27:43.280Z",
      "content": "<p>Keras pretrained models on ImageNet data (<a href=\"http://keras.io/applications\">http://keras.io/applications</a>)</p>",
      "rawMarkdown": "Keras pretrained models on ImageNet data (http://keras.io/applications)"
    },
    {
      "id": 423549,
      "postDate": "2018-11-18T14:16:01.217Z",
      "content": "<p>Inception_v4 and weights... <a href=\"https://github.com/kentsommer/keras-inceptionV4/releases\">https://github.com/kentsommer/keras-inceptionV4/releases</a></p>",
      "rawMarkdown": "Inception_v4 and weights... https://github.com/kentsommer/keras-inceptionV4/releases"
    },
    {
      "id": 420810,
      "postDate": "2018-11-14T06:45:59.630Z",
      "content": "<p>Im using the keras pretrained resnet50 weights.</p>",
      "rawMarkdown": "Im using the keras pretrained resnet50 weights."
    },
    {
      "id": 420174,
      "postDate": "2018-11-13T07:45:27.937Z",
      "content": "<p>I'm using pretrained resnet pytorch model's weights via fastai library interface:\n<a href=\"https://pytorch.org/docs/stable/torchvision/models.html\">https://pytorch.org/docs/stable/torchvision/models.html</a>\n<a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a></p>",
      "rawMarkdown": "I'm using pretrained resnet pytorch model's weights via fastai library interface:\nhttps://pytorch.org/docs/stable/torchvision/models.html\nhttps://github.com/fastai/fastai\n"
    },
    {
      "id": 418998,
      "postDate": "2018-11-11T03:55:35.993Z",
      "content": "<p>Pre-trained ResNext models at: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "rawMarkdown": "Pre-trained ResNext models at: https://github.com/Cadene/pretrained-models.pytorch"
    },
    {
      "id": 415361,
      "postDate": "2018-11-05T01:53:26.583Z",
      "content": "<p>Pretrained models:\n<a href=\"https://github.com/pytorch/vision/\">https://github.com/pytorch/vision/</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "rawMarkdown": "Pretrained models:\nhttps://github.com/pytorch/vision/\nhttps://github.com/Cadene/pretrained-models.pytorch"
    },
    {
      "id": 414427,
      "postDate": "2018-11-02T18:46:53.720Z",
      "content": "<p>I am using pretrained weights from <a href=\"https://github.com/fchollet/deep-learning-models/releases\">https://github.com/fchollet/deep-learning-models/releases</a> as well</p>",
      "rawMarkdown": "I am using pretrained weights from https://github.com/fchollet/deep-learning-models/releases as well"
    },
    {
      "id": 413339,
      "postDate": "2018-10-31T18:38:14.980Z",
      "content": "<p>I am using some pre-trained keras model weights from:\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/download/v0.7/inception_resnet_v2_weights_tf_dim_ordering_tf_kernels_notop.h5\">https://github.com/fchollet/deep-learning-models/releases/download/v0.7/inception_resnet_v2_weights_tf_dim_ordering_tf_kernels_notop.h5</a>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/download/v0.5/inception_v3_weights_tf_dim_ordering_tf_kernels_notop.h5\">https://github.com/fchollet/deep-learning-models/releases/download/v0.5/inception_v3_weights_tf_dim_ordering_tf_kernels_notop.h5</a>\n<a href=\"https://github.com/titu1994/Keras-NASNet/releases/download/v1.2/NASNet-large-no-top.h5\">https://github.com/titu1994/Keras-NASNet/releases/download/v1.2/NASNet-large-no-top.h5</a>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/download/v0.2/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5\">https://github.com/fchollet/deep-learning-models/releases/download/v0.2/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5</a>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/download/v0.4/xception_weights_tf_dim_ordering_tf_kernels_notop.h5\">https://github.com/fchollet/deep-learning-models/releases/download/v0.4/xception_weights_tf_dim_ordering_tf_kernels_notop.h5</a>\n<a href=\"https://github.com/keras-team/keras-applications/releases/download/densenet/densenet121_weights_tf_dim_ordering_tf_kernels_notop.h5\">https://github.com/keras-team/keras-applications/releases/download/densenet/densenet121_weights_tf_dim_ordering_tf_kernels_notop.h5</a></p>",
      "rawMarkdown": "I am using some pre-trained keras model weights from:\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.7/inception_resnet_v2_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.5/inception_v3_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/titu1994/Keras-NASNet/releases/download/v1.2/NASNet-large-no-top.h5\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.2/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.4/xception_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/keras-team/keras-applications/releases/download/densenet/densenet121_weights_tf_dim_ordering_tf_kernels_notop.h5"
    },
    {
      "id": 452551,
      "postDate": "2019-01-08T22:02:39.437Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 448616,
      "postDate": "2019-01-01T17:07:38.717Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 442282,
      "postDate": "2018-12-19T18:20:09.643Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true,
      "replies": [
        {
          "id": 443261,
          "postDate": "2018-12-21T10:08:26.550Z",
          "content": "<p>Thanks for sharing, Jarvis!\nThe archive with intensities (8Gb, 512*512 JPGs) contains 264012 files, that correspond to 66003 identifiers. However, the CSV contains 70327 identifiers. </p>\n\n<p>For those, who are confused by that - there are 4315 duplicating identifiers in CSV, some of them, apparently, repeat more than once.</p>",
          "rawMarkdown": "Thanks for sharing, Jarvis!\nThe archive with intensities (8Gb, 512*512 JPGs) contains 264012 files, that correspond to 66003 identifiers. However, the CSV contains 70327 identifiers. \n\nFor those, who are confused by that - there are 4315 duplicating identifiers in CSV, some of them, apparently, repeat more than once.",
          "votes": 3
        },
        {
          "id": 445011,
          "postDate": "2018-12-25T11:02:04.310Z",
          "content": "<p>And it looks like there is some missing channels too :   </p>\n\n<pre><code>There is only the red channel for 780_F9_1\nOnly the blue and green for 780_F9_2  \n1544_F3_4 is missing yellow   \n161_G2_2 is missing green and yellow  \n</code></pre>\n\n<p>And many more. I am missing something ?  </p>\n\n<p>By the way big thanks for sharing this dataset and kernel !  </p>",
          "rawMarkdown": "And it looks like there is some missing channels too :   \n  \n\n    There is only the red channel for 780_F9_1\n    Only the blue and green for 780_F9_2  \n    1544_F3_4 is missing yellow   \n    161_G2_2 is missing green and yellow  \n\nAnd many more. I am missing something ?  \n  \nBy the way big thanks for sharing this dataset and kernel !  ",
          "votes": 1
        },
        {
          "id": 446689,
          "postDate": "2018-12-28T14:11:42.877Z",
          "content": "<p>Have you found  a way to resolve  it?</p>",
          "rawMarkdown": "Have you found  a way to resolve  it?\n\n",
          "votes": 1
        },
        {
          "id": 449638,
          "postDate": "2019-01-03T13:33:21.210Z",
          "content": "<p>I removed all not unique ids and ids that correspond to groups of images with missing channels. I don't know the results yet, training is in progress. Dataset: <a href=\"https://www.kaggle.com/kgeorge/human-protein-atlas-additional-images\">https://www.kaggle.com/kgeorge/human-protein-atlas-additional-images</a></p>",
          "rawMarkdown": "I removed all not unique ids and ids that correspond to groups of images with missing channels. I don't know the results yet, training is in progress. Dataset: https://www.kaggle.com/kgeorge/human-protein-atlas-additional-images",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 430860,
      "author_name": "TomomiMoriyama",
      "author_url": "",
      "post_date": "2018-12-01T04:28:06.257000",
      "content": "<p>Hi! @FabSchreiber </p>\n\n<p>Here's list</p>\n\n<p>!! still contains duplicated images\n　RGB images(HPAv18.csv) :  77,878 sample NG\n　RGB images 77,864 sample  ...deleted id duplication,still contains duplicated images\n　RGBY images  77,430 sample\n　RGBY withoutUncertain:73,881 sample</p>\n\n<p>old csv list contains\n GeneID_ Dir_ImageURL Target(28 class)</p>\n\n<p>\"<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)</p>\n\n<p><strong>Modified:</strong> ..without duplicated images (Gene information lost,labels merged)\n　RGB_wodpl 75,040 sample\n　RGBY_wodpl  74,606 sample\n　RGBY withoutUncertain_wodpl:71,437 sample</p>\n\n<p>new csv list contains\n　Dir_ImageURL Target(28 class)</p>\n\n<p><strong>AddFiles:</strong> @Chase_the_Trane adviced me to open how I make those csv files.\n　 1. Parce XML and Download HPAv18 Image.html\n--&gt;total downloaded image 60GB\n　  2.  NoYellow512.txt\n--&gt;resize img 512x512 png 70GB, and I realized some sample has no yellow filter.\n　  3. Make Metadata for HPAv18 Image.html\n--&gt;duplicate image exist\n　 4. Clean HPAv18 dataset.html...How I merged the labels.</p>\n\n<p><strong>AddFiles2:</strong> add CellLine information \n*_withCellLine.csv</p>",
      "votes": 39,
      "replies": [
        {
          "id": 431235,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-01T22:53:04.120000",
          "content": "<p>I noticed a few duplicates here. Different ENSG id, but same image paths.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 431287,
          "author_name": "Russ Wolfinger",
          "author_url": "",
          "post_date": "2018-12-02T01:50:10.870000",
          "content": "<p>Yes, I'm seeing only 75,040 unique paths.  For one extreme example, it looks like path 10580_1610_C1_1 is repeated 22 times and has four different label assignments.   So we have dups + some label noise.   It would be interesting if an expert would examine the associated Ensembl gene sets to determine how much of a cause for concern this might be.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 431291,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-02T01:58:11.767000",
          "content": "<p>&gt;Different ENSG id, but same image paths.\nYou are right. There are 4552 such duplicate rows in one file.\nI should have checked before sharering the list, sorry.</p>\n\n<p>I 'm now trying to remove duplicates, \nand found same image but different labels, \nso not simply remove duplicate rows but merge labels.</p>\n\n<p>e.g.\nENSG00000081853\nCytosol (GO:0005829);Nucleoli (GO:0005730);Nucleus (GO:0005634);Plasma membrane (GO:0005886);Vesicles (GO:0043231)</p>\n\n<p>ENSG00000240764\nCytosol (GO:0005829);Nucleoplasm (GO:0005654);Nucleus (GO:0005634);Plasma membrane (GO:0005886);Vesicles (GO:0043231)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 431367,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-02T05:30:06.990000",
          "content": "<p>I only noticed because my latest download script checks and skips any files that exist already. For now I am leaving the noise in, having multiple rows with the same image id but different labels. I think it might help with overfitting.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 431484,
          "author_name": "Russ Wolfinger",
          "author_url": "",
          "post_date": "2018-12-02T11:22:13.980000",
          "content": "<p>Thanks Tomomi and Brian.  BTW, when typing a word that contains more than one underscore, use a backslash in front of each so markup does not convert the intervening text to italics.  I think this trick may work for other special characters as well, e..g \\&lt;;  these can alternatively be specified with an ampersand.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 431523,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-02T12:35:07.550000",
          "content": "<p>Thank you @Russ. \\trick worked!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432922,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2018-12-04T13:26:21.433000",
          "content": "<p>Thanks for your script.\nNow I am running it to download the HPAv18 data,</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 436319,
      "author_name": "TomomiMoriyama",
      "author_url": "",
      "post_date": "2018-12-10T05:52:40.370000",
      "content": "<p>Those who read this comment, please help me.\nI want to use some open source code and need some advice about the usage...</p>\n\n<p>One month a go, <a href=\"/hengck23\">@hengck23</a> introduced this paper</p>\n\n<p>Learning unsupervised feature representations for single cell microscopy images with paired cell inpainting\nAlex Lu, Oren Z Kraus, Sam Cooper, Alan M Moses</p>\n\n<p>on [ideas and discussion] thread\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69955\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69955</a></p>\n\n<p>I found it very attractive, but source code was not released at the time.</p>\n\n<p>One week ago, source code has finally released!\n<a href=\"https://github.com/alexxijielu/paired_cell_inpainting#human-model\">https://github.com/alexxijielu/paired_cell_inpainting#human-model</a></p>\n\n<p>It contains \n- Cleaner HPA data downloading scripts.\n- Extracting single cell.\n- pre-trained weights \n- and more</p>\n\n<p>Their task setting(unsupervised) is different from us(supervised classification),\nbut can be used with some extra code.</p>\n\n<p>They released the code under [GPL-2.0]</p>\n\n<p>My question is, \n\"Is it OK to use this code and pre-trained weight as a baseline model?\"\n if it's against the rule \"Can I re-implement from scratch and imitate the idea?\"\n if it's within the rule \"Should I ask the author of this code about the usage in Github/Issues section beforehand?\"\nOr they might also participated in this competition already!?</p>",
      "votes": 13,
      "replies": [
        {
          "id": 436726,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2018-12-10T20:36:46.483000",
          "content": "<p>Thank you for sharing. My understanding is as long as the algorithm described in the paper is not patented we are okay to use the implementation (software) to compete whatever the license given is. Let's wait for the organizers to rule this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436740,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-10T21:24:03.863000",
          "content": "<p>Hi! Original author of this preprint and code here! I was wondering where the sudden traffic on my GitHub came from!</p>\n\n<p>I'd be delighted if anyone ended up using my work to pre-train a model for this competition, and you're free to use my code if the Kaggle organizers rule this to be okay. It would be a really good way to validate the utility of our method, so I totally encourage it. We were hoping to put together a team from my lab, but as the competition deadline inches closer and we're all busy with end-of-semester academic stuff, we aren't sure if we can put together a fully polished result. </p>\n\n<p>Off the top of my head, here's some of the problems and limitations you'd need to solve:\n - Our method is designed to extract single-cell features, and you'd need to find a way to aggregate these into a full-image prediction. You could just take an average of the features and train a classifier on that, but that seems a bit unrefined - I'd love to see what kind of an end-to-end approach might be possible for this (a RCNN maybe?)\n - One major limitation is with phenotypes that don't penetrate well (e.g. mitotic spindle, etc.) In my preliminary exploration of the feature space, I think the model ends up learning to ignore phenotypes where maybe only one out of fifty cells expresses the phenotype. It sounds like fine-tuning might solve this problem though.\n - Preprocessing: to input single-cell crops into this method, you need to, of course, get single cell crops. The current method we have on the GitHub is very ad-hoc - we just run the human nuclei channel through an Otsu filter. I'm sure a better segmentation method would be more robust to this.</p>\n\n<p>Let me know if you have any questions/comments and I'll do my best to address them! </p>",
          "votes": 23,
          "replies": []
        },
        {
          "id": 436751,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2018-12-10T21:38:42.827000",
          "content": "<p>Alex, thanks. You should join and set a strong baseline for us</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 436788,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-11T00:06:04.130000",
          "content": "<p>Thank you!!!!!! Alex.\nThank you for allowing everyone to use your great model and sharing many ideas.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436794,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-11T00:34:13.843000",
          "content": "<p>No problem. :) If you end up using my pre-trained weights, use the one named \"filtered_model_weights.h5\", those are way better than the default because we filtered out all variable proteins in training it instead of treating the problem as a fully unsupervised one. I wish we had weights for stronger models to offer, but AlexNet was all we experimented with for now. In any case, I think it should be better than ImageNet pretrained weights, which I think most people are using, just because the weights have learned HPA-specific channel correlations. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 436837,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-11T02:26:59.413000",
          "content": "<p>Great code, thanks for sharing. I've integrated it into my processing and generated over 1M single cell images to feed into a network:</p>\n\n<p><img src=\"http://brians.network/images/hpa_cells.png\" alt=\"Image\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437012,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-11T08:49:02.370000",
          "content": "<p>I have experimented with single cell images and was not overly successful. Maybe that was because I did not aggregate the scores properly for an entire image, but honestly I think that it was unsuccessful because we have only image-level annotations. I looked at some images manually and noticed (especially for the very rare classes) that the structure of interest (such as cell joints or some other small classes) must be present in only one cell within the entire image for the entire image to get this class. Given that there are always multiple cells in an image, sometimes even &gt;20, you would need to reformulate the whole  task as a <strong>multiple instance learning</strong> problem.\nAlso I would be interested in understanding how you (@Alex Lu) managed to separate cells. Cells are of very different size and often cells are touching with their borders so if you crop around the nucleus of a cell you can end up with multiple cells in your crop. (But honestly I think this is much less of a problem than the multiple instance learning thing).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437171,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-11T13:36:16.013000",
          "content": "<p>We simply segmented the nuclei channel and extracted a fixed crop around the cell. You're right that it does sometime cause multiple cells in the same crop, but that generally doesn't seem to pose an issue for learning the feature representation - you can check out in the paper how we used the model generatively, because there's an example of where we inpaint a protein from a crop with two cells onto a single cell, I think.</p>\n\n<p>In general, the method we use seems to learn a feature representation robust to single cell variability, including morphology, cell line, etc. I think this will be important to the competition - knowing the domain, the organizers may try to confound your models by giving them unseen cell lines in the test stage, because this would really emphasize generalizable models. Imagine going from bone cells to skin cells or brain cells - they look very different, and you want to make sure your model hasn't overfit to the morphologies in your training set.</p>\n\n<p>You're right that the challenge is going to be aggregating these crops, though! But I don't think the use of single cell crops is fundamentally a bad idea - if you think about it, it's basically a pseudo-attention mechanism, that directs your network to the relevant parts of the image. MIL sounds like a good idea for solving some of the issues with poorly penetrating classes. </p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 432870,
      "author_name": "TomomiMoriyama",
      "author_url": "",
      "post_date": "2018-12-04T12:13:17.520000",
      "content": "<p>Hi, @pete\nXML parsing is time consuming.\nAlternatively, you can download from csv list.</p>\n\n<p>(How I processed the list was posted few comments ago.\nMy code is far from efficient and safe, there's no Error handrring, though )</p>\n\n<p><strong>Note 12/13:</strong>\n*<em>Please use <a href=\"/davidwagnerkc\">@davidwagnerkc</a> 's downloading script !</em>*\nNot gray-scaled image hurts me:11b_Segment_and_Crop.html\nToo many class 27?:Check metadata for HPAv18.html</p>",
      "votes": 14,
      "replies": [
        {
          "id": 436787,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-11T00:04:50.230000",
          "content": "<p>Hi <a href=\"/tomomimoriyama\">@tomomimoriyama</a>, thanks for sharing this great scripts. After getting HPAv18 datasets, I found that the downloaded images are all having 3 channels. For example, <code>img_HPAv18_some_name_blue.jpg</code> has 3 channels and all channels have some values. Would you tell how you convert these 3-channel \"RGB\" image to a grayscale one? <code>0.299 R + 0.587 G + 0.114 B</code> is what I am doing now but some images ended up with very low visibility, not sure if I am doing something wrong.</p>\n\n<p>The following are two images, the first one is <code>11_203_h3_1_yellow</code> to grayscale, the second one is <code>11_203_h3_1_blue</code> to grayscale. Before converting to grayscale the blue one looks good, but after the conversion, it becomes very dark.</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/14886380/49769533-c762b500-fc9d-11e8-9570-fbcfaf4ca370.png\" alt=\"yellow\">\n<img src=\"https://user-images.githubusercontent.com/14886380/49769530-c762b500-fc9d-11e8-8580-e881e1605458.png\" alt=\"blue\"></p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 436934,
          "author_name": "kickback",
          "author_url": "",
          "post_date": "2018-12-11T05:42:47.477000",
          "content": "<p>Hi <a href=\"/tomomimoriyama\">@tomomimoriyama</a>, </p>\n\n<p>Your scripts are working fine for me but the download time is very long. I have downloaded about 18k of the 74k  images in about 11 hours, maybe there are many kagglers accessing the protein atlas.</p>\n\n<p>Thanks</p>\n\n<p>Kickback</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436948,
          "author_name": "David Wagner",
          "author_url": "",
          "post_date": "2018-12-11T06:24:13.087000",
          "content": "<p>I used <a href=\"/tomomimoriyama\">@tomomimoriyama</a>'s csv file to download all RGBY images and resize to (512, 512) in under 8 hours if I remember correctly. Here is the code from my notebook, I haven't tested it, but it worked before in notebook form. Maybe map only <code>df[:1000].Id</code> to make sure it works and get a time estimate. I ran this on a MBP with 4 cores. </p>\n\n<p>Change where csv is read from and the <code>hpa_dir</code> you want to write to. </p>\n\n<p>EDIT: Thanks to @Mark Worrall I am actually downloading RGBY images now, was missing B channel before!\n```\nimport pandas as pd\nimport requests\nfrom multiprocessing import Pool\nfrom pathlib import Path\nfrom PIL import Image\nfrom io import BytesIO</p>\n\n<p>df = pd.read_csv('../HPAv18RBGY_wodpl.csv')</p>\n\n<p>def download(id_):\n    try:\n        hpa_dir = Path('../HPAv18')\n        hpa_dir.mkdir(parents=True, exist_ok=True)\n        image_dir, _, image_id = id_.partition('_')\n        url = f'<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>{image_dir}/{image_id}_blue_red_green_yellow.jpg'\n        r = requests.get(url)\n        image = Image.open(BytesIO(r.content)).resize((512, 512), Image.LANCZOS)\n        image.save(hpa_dir / f'{id_}.png', format='png')\n    except:\n        print(f'{id_} broke...')</p>\n\n<p>p = Pool()\np.map(download, df.Id)\n```</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 437011,
          "author_name": "Hogger",
          "author_url": "",
          "post_date": "2018-12-11T08:48:55.647000",
          "content": "<p>A multi-processing version may works better for chinese guys who has a gfw :)</p>\n\n<pre><code>import os\nfrom multiprocessing.pool import Pool\nfrom tqdm import tqdm\nimport requests\nimport pandas as pd\ndef download(pid, sp, ep):\n    colors = ['red', 'green', 'blue', 'yellow']\n    DIR = \"/home/czy/dataset/human_protain/moredata/\"\n    v18_url = 'http://v18.proteinatlas.org/images/'\n    imgList = pd.read_csv(\"/home/czy/dataset/human_protain/moredata/HPAv18RBGY_wodpl.csv\")\n    for i in tqdm(imgList['Id'][sp:ep], postfix=pid):  # [:5] means downloard only first 5 samples, if it works, please remove it\n        img = i.split('_')\n        for color in colors:\n            img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n            img_name = i + \"_\" + color + \".jpg\"\n            img_url = v18_url + img_path\n            r = requests.get(img_url, allow_redirects=True)\n            open(DIR + img_name, 'wb').write(r.content)\n\ndef run_proc(name, sp, ep):\n    print('Run child process %s (%s) sp:%d ep: %d' % (name, os.getpid(), sp, ep))\n    download(name, sp, ep)\n    print('Run child process %s done' % (name))\n\nif __name__ == \"__main__\":\n    print('Parent process %s.' % os.getpid())\n    img_list = pd.read_csv(\"/home/czy/dataset/human_protain/moredata/HPAv18RBGY_wodpl.csv\")['Id']\n    list_len = len(img_list)\n    process_num = 100\n    p = Pool(process_num)\n    for i in range(process_num):\n        p.apply_async(run_proc, args=(str(i), int(i * list_len / process_num), int((i + 1) * list_len / process_num)))\n    print('Waiting for all subprocesses done...')\n    p.close()\n    p.join()\n    print('All subprocesses done.')\n</code></pre>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 437054,
          "author_name": "GhMa",
          "author_url": "",
          "post_date": "2018-12-11T09:59:09.780000",
          "content": "<p>Thank u. for sharing. \ntiny remind, it`s protein not protain.\nbtw, how long could u download all the image files using ur script ?  I am suffering for my slow speed, is it because the EDU NET?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437119,
          "author_name": "Hogger",
          "author_url": "",
          "post_date": "2018-12-11T11:57:46.903000",
          "content": "<p>thanks for your reminding. i'm also in the EDU NET, i use 100 threads which takes me about 4 hours to download all the 74k+ images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437120,
          "author_name": "Vaagn Minasian",
          "author_url": "",
          "post_date": "2018-12-11T11:58:23.203000",
          "content": "<p>I also faced with the issue of 3 channels in .jpg vs 1 in .png files. Moreover, if I convert .jpg images to grayscale and then train my model with extended dataset - I got very low score. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 437122,
          "author_name": "Hogger",
          "author_url": "",
          "post_date": "2018-12-11T12:01:33.090000",
          "content": "<p>and I'm in the IPv6 net</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437167,
          "author_name": "wchzh",
          "author_url": "",
          "post_date": "2018-12-11T13:29:41.613000",
          "content": "<p>hello，@Hogger, I run your code to download the HPA data，but it failed. So I change process_num from 100 to 5, but get following result:  could you help me? thank you.</p>\n\n<p>Parent process 6352.\nWaiting for all subprocesses done...\nRun child process 0 (6356) sp:0 ep: 14921\nRun child process 1 (6357) sp:14921 ep: 29842\nRun child process 2 (6358) sp:29842 ep: 44763\nRun child process 3 (6359) sp:44763 ep: 59684\nRun child process 4 (6360) sp:59684 ep: 74606\n  0%|                                                                                    | 0/14921 [00:00\n\n</p><p>All subprocesses done.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437203,
          "author_name": "David Wagner",
          "author_url": "",
          "post_date": "2018-12-11T14:37:30.520000",
          "content": "<p>@ManyFoldCV\n<code>\nfrom PIL import Image\nImage.open('img_HPAv18_some_name_blue.jpg').convert('L')\n</code>\nGive that a shot. The object returned has a <code>.save()</code> method and also can be converted to <code>ndarray</code> with <code>np.array(pil_image)</code>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 437244,
          "author_name": "Hogger",
          "author_url": "",
          "post_date": "2018-12-11T15:53:05.280000",
          "content": "<p>these are the normal output. if it stuck at 0. Try to access a single image by your internet browser,  if you even can't access the image by your browser, maybe its your internet that makes u can't access it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437318,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-11T17:54:03.577000",
          "content": "<p><a href=\"/davidwagnerkc\">@davidwagnerkc</a>\nMuch appreciated!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437386,
          "author_name": "David Silva",
          "author_url": "",
          "post_date": "2018-12-11T20:16:31.047000",
          "content": "<p>I have made a few small modifications to @Hogger code:\n- images are resized to a specified size: <code>image_size</code>\n- images are saved as grayscale\n- images are saved as PNG</p>\n\n<p>```\nimport os\nimport errno\nfrom multiprocessing.pool import Pool\nfrom tqdm import tqdm\nimport requests\nimport pandas as pd\nfrom PIL import Image</p>\n\n<p>def download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path</p>\n\n<pre><code>        # Get the raw response from the url\n        r = requests.get(img_url, allow_redirects=True, stream=True)\n        r.raw.decode_content = True\n\n        # Use PIL to resize the image and to convert it to L\n        # (8-bit pixels, black and white)\n        im = Image.open(r.raw)\n        im = im.resize(image_size, Image.LANCZOS).convert('L')\n        im.save(os.path.join(save_dir, img_name), 'PNG')\n</code></pre>\n\n<p>if <strong>name</strong> == '<strong>main</strong>':\n    # Parameters\n    process_num = 24\n    image_size = (512, 512)\n    url = '<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>'\n    csv_path =  \"path/to/csv/HPAv18RBGY_wodpl.csv\"\n    save_dir = \"where/the/images/are/saved\"</p>\n\n<pre><code># Create the directory to save the images in case it doesn't exist\ntry:\n    os.makedirs(save_dir)\nexcept OSError as exc:\n    if exc.errno != errno.EEXIST:\n        raise\n    pass\n\nprint('Parent process %s.' % os.getpid())\nimg_list = pd.read_csv(csv_path)['Id']\nlist_len = len(img_list)\np = Pool(process_num)\nfor i in range(process_num):\n    start = int(i * list_len / process_num)\n    end = int((i + 1) * list_len / process_num)\n    process_images = img_list[start:end]\n    p.apply_async(\n        download, args=(str(i), process_images, url, save_dir, image_size)\n    )\nprint('Waiting for all subprocesses done...')\np.close()\np.join()\nprint('All subprocesses done.')\n</code></pre>\n\n<p>```</p>\n\n<p>The script from @Hogger downloads 100 images in about 16 seconds; this script increases that time to 27 seconds, not that bad of a trade for some convenience.  On my computer tqdm says it should take around 4 hours total. I'm also CPU bottlenecked because I'm training at the same time, so a good internet connection and a free CPU can probably decrease this time significantly.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 437516,
          "author_name": "Hogger",
          "author_url": "",
          "post_date": "2018-12-12T04:25:54.587000",
          "content": "<p>great job!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437563,
          "author_name": "kickback",
          "author_url": "",
          "post_date": "2018-12-12T06:04:36.213000",
          "content": "<p>I am now running Hogger's code with David's tweaks, download is much faster but file sizes are also much larger than with <a href=\"/tomomimoriyama\">@tomomimoriyama</a> code, look like total down load will be 250-260 GB. I commented out </p>\n\n<p>im = im.resize(image_size, Image.LANCZOS).convert('L')</p>\n\n<p>because I don't want gray scale images. I am curious as to why David Silva wants gray scale vs. color images.</p>\n\n<p>Kickback</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437586,
          "author_name": "David Wagner",
          "author_url": "",
          "post_date": "2018-12-12T06:44:53.357000",
          "content": "<p>@kickback All 74k RGBY images takes 30GB disk space on my machine being resized to 512 x 512 RGB PNGs. Silva's code saves as grayscale because he is saving four separate images for each of the channel, in a similar manner to how the competition data has been provided.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437730,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-12T11:43:04.167000",
          "content": "<p>Hi <a href=\"/davidwagnerkc\">@davidwagnerkc</a>,</p>\n\n<p>Thanks for your script above though it's a little confusing following this thread. Does your script download RGB images as the url has <code>'red_green_yellow'</code> in the name?</p>\n\n<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437831,
          "author_name": "David Wagner",
          "author_url": "",
          "post_date": "2018-12-12T15:30:12.987000",
          "content": "<p>@Mark Worrall That might be a bug! Haha. I must have just forgot to add <em>blue</em> wherever that is supposed to go in the color order. I'll fix it and edit the post. Thanks!</p>\n\n<p>EDIT: I updated the code above if anybody is actually using it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437848,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-12T16:03:08.400000",
          "content": "<p>No worries, I fixed it up to get each channel color individually anyway. Thank you for this!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438010,
          "author_name": "kickback",
          "author_url": "",
          "post_date": "2018-12-12T23:51:55.793000",
          "content": "<p>I ran the code with the one line commented out and after 16 hours I have downloaded 292k items that consume 728 GB of space. I have no idea what I really have but some of the images look pretty interesting!</p>\n\n<p>Time for some resizing.......</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438012,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-12T23:56:58.443000",
          "content": "<p>@ManyFoldCV\nThank you for your information that downloaded image has 3 channels.\nI apology for those who affected by my mistake.I'm very sorry.</p>\n\n<p><a href=\"/davidwagnerkc\">@davidwagnerkc</a>,@Hogger\nThank you for your great code!</p>\n\n<p>@Mark Worrall, @kickback\nto be Grayscale is very important!\nas <a href=\"/davidwagnerkc\">@davidwagnerkc</a> said [how the competition data has been provided]</p>\n\n<p>In my case, \nI've been using 3 channels image without noticing it,\nscinse I use Iafoss's kernel and it loads image like</p>\n\n<pre><code>def open_rgby(path,id): \n    colors = ['red','green','blue','yellow']\n    flags = cv2.IMREAD_GRAYSCALE\n    img = [cv2.imread(os.path.join(path, id+'_'+color+'.png'), flags).astype(np.float32)/255\n           for color in colors]\n    return np.stack(img, axis=-1)\n</code></pre>\n\n<p>Visualizing it make no difference from me..</p>\n\n<p>But, from yesterday,\nI'm trying @Alex 's SingleCell Extraction.\nI found rgb-png-converted one do harm but grayscale-converted is fine.\nI attached html file.\nOnly rgb-png-converted image, detected nuclei(right image) is not clear.\nResizing is not affected the result.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438187,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-13T08:26:25.890000",
          "content": "<p>Hi Tomomi, </p>\n\n<p>Yes, so I convert each 3 channel red, green and blue image of a single id to grayscale using <code>.convert('L')</code>.</p>\n\n<p>Is there any other pre-processing you do? In particular there seems to be a large number of class 27 (&gt;100) which was making me wonder about the accuracy of the labels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438197,
          "author_name": "PhysicsGuy89",
          "author_url": "",
          "post_date": "2018-12-13T08:49:21.400000",
          "content": "<p>Hi there,</p>\n\n<p>I'm (also) still confused by the color channel information in the additional HPA data. When using grayscale conversion of each of the four images corresponding to a single id via convert('L'),  the problem is that the pixel statistics is drastically different compared to the original training set of the competition.</p>\n\n<p>For example, the blue channel in the original 512x512 training images on average has a pixel intensity of ~14, while the grayscaled version of the 'blue' images in the additional HPA data only has ~1.7 to ~2 (I haven't explored the full set yet). I fear that this spoils training...</p>\n\n<p>It would be simple if each image in the additional HPA data had non-zero pixels only in the \"right\" color channel (except for yellow, where things inevitably are a bit more complicated), but unfortunately that's not the case. Any ideas? </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 438281,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-13T12:02:29.490000",
          "content": "<p>Hi, @Mark</p>\n\n<p>About pre-processing\nI don't use any pre-processing, but recalculated mean and std\n# mean and std in of each channel in the train set\n# Iafoss original RGBY\n#stats = A([0.0804, 0.0526, 0.0547, 0.0827], [0.394, 0.321, 0.327, 0.399])\n# HPAv18 all RGBY (each image is colored RGB png image 512x512)\nstats = A([0.06734, 0.05087, 0.03266, 0.09257],[0.11997, 0.10335, 0.10124, 0.1574 ])\nI'm not confident about the stats I calculated, \nsince I tried to reproduce original-train-set stats, my result was different from others report.... </p>\n\n<p>About too match class 27\nI'll check my script right away.</p>\n\n<p>Hi, @PhysicsGuy89\n@Brian provides conversion ruby script that preserves original intensity!\n (He released 10 days ago, but I've not tried yet, I'm new to ruby...)\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206#431936\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206#431936</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438333,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-13T13:47:04.933000",
          "content": "<p>Just a quick note that you can't treat color channels in fluorescent images in the same way as standard digital images. Each color channel is actually acquired as a separate greyscale image, and researchers just stack and artificially color them to make it easier to see multiple structures of the cell at once. Standard greyscale conversion algorithms will take a weighted average of the color channels (including cv2's function), which is why you're seeing the blue intensities go down so much, since it's being multiplied by a modifier. The correct way is to just simply split the channels (and then tile them if you really need a three channel image) - so for example, in Python representing the images as a numpy array, you'd just go image[:, :, 0]... etc. </p>\n\n<p>For this reason, you should really download the yellow channel as a separate image. The RGB images have each color channel as an independent fluorescent dye, but the yellow channel will add green and blue intensities to these images that will confuse their splitting. It's all artificially colored - it's just that with four markers, you run out of independent color channels since you only have 3 max with RGB. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 438339,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-13T13:59:08.430000",
          "content": "<p>Usually color order is rgb so\nfor blue take image[:, :, 2]\nfor red take image[:, :, 0]\nfor green take image[:, :, 1]\nfor yellow you can take either image[:, :, 0] or image[:, :, 1]. Yellow is created by putting the intensity values in both r and g channel. Due to image compression (which is pretty bad for the HPA images by the way. Shame on whoever did this), r and g channel of yellow images will not be identical. I did some visual as well as numeric comparisons though and they are almost identical to the point where just using the r channel for these images should be fine (image[:, :, 0])</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 438343,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-13T14:06:32.797000",
          "content": "<p>Hi Alex, thank you for this though I'm not sure I follow fully.</p>\n\n<p>We have from the protein atlas website images that are 'red', 'green', 'blue' and 'yellow' channels. Each of these 4 images themselves have 3 channels.</p>\n\n<p>The data provided as per the competition has a single grayscale image for each of red, green, blue and yellow. So some mapping has taken place to compress each competition channel image into a single grayscale image - is this just the standard weighting per: </p>\n\n<p>L = R * 299/1000 + G * 587/1000 + B * 114/1000</p>\n\n<p>Thanks again,</p>\n\n<p>Mark</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438348,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-13T14:20:37.677000",
          "content": "<p>I'm not an organizer, so I can't be sure, but it's usually the other way around - the microscope will give you four greyscale images, and then you convert each of these images into a different color by only showing the intensity of the pixels in one or two channels</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 438355,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-13T14:32:16.253000",
          "content": "<p>+1 to Alex Lu answer. You need to think about how the mapping was done from the filter response to rgb. Look at my comment above to see how to convert back. (image being a np.ndarray of shape (X, Y, 3))</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438357,
          "author_name": "PhysicsGuy89",
          "author_url": "",
          "post_date": "2018-12-13T14:33:27.693000",
          "content": "<p>@Alex Lu: I fully agree, standard RGB --&gt; grayscale conversion doesn't make much sense here, as \"blue\" in the context of an HPA channel has nothing to do with the actual blue color in an RGB image.</p>\n\n<p>But as Mark is saying, this doesn't solve the problem: for the additional HPA data, we only have access to the .jpg images, which <strong>for each of the channels R, G, B, Y</strong> are three-dimensional RGB images. In other words: do you know how these jpg images were constructed from the original data (which presumably only contained a single channel of information for each of the R, G, B, Y data)? If we know this procedure, we can apply it backwards to get the actual single-channel data for each of the four channels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438360,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-13T14:34:27.313000",
          "content": "<p>Hi, @Mark\nI've checked class 27, total 122 samples. There's no Uncertain status in it.\nI cannot judge all the annotated labels are correct...\nand same of the sample I cannot find rod like shape...\nAttached html what I have checked.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438361,
          "author_name": "PhysicsGuy89",
          "author_url": "",
          "post_date": "2018-12-13T14:36:50.403000",
          "content": "<p>@FabianIsensee: yeah, that was my original idea as well. But then it's really confusing that e.g. the 'green' images in the additional HPA data contain non-zero pixels for R, G and B. \nDo you think that is an artefact of the jpg image compression as well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438365,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-13T14:43:22.383000",
          "content": "<p>@PhysicsGuy89 I suspect this is the case but I don't know enough about jpg compression to be sure</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438369,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-13T14:47:31.827000",
          "content": "<p>Like I said - it's just displaying the fluorescent intensities of the micrograph in one channel only. So to turn these back into greyscale images, just tile the one channel it corresponds into all channels. </p>\n\n<p>Values in other channels are unexpected and probably due to the jpg compression. But also, the intensities in the jpg HPA images have also been stretched out for display on their web-front end. Depending on how they acquired the original micrographs, the original images may not have been treated this way - for the images my lab handles, you don't do this in the original tiff files acquired from the microscope, because you lose information about the level of protein expression when you rescale the images. That being said, these are fluorescent antibodies and not GFP-tagged proteins like we work with (where we expect the fluorescence to be correlated with protein expression), so double check that their tiff files are acquired the same way. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 438374,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-13T14:55:12.087000",
          "content": "<p>This is getting confusing.</p>\n\n<p>I'm with <a href=\"/physicsguy89\">@physicsguy89</a> : it's not clear how they've gone from a 3 channel jpg for each channel to a single channel image for the competition data. </p>\n\n<p>@FabianIsensee are you saying that when we get a (for example) GREEN channel image from protein atlas website which has 3 channels we should just take image[:, :, 1]? And likewise for RBY extracting the corresponding channel?</p>\n\n<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a> - thank you for this though no html attached?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438376,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-13T15:00:43.797000",
          "content": "<blockquote>\n  <p><strong>Alex Lu wrote</strong></p>\n  \n  <blockquote>\n    <p>Like I said - it's just displaying the fluorescent intensities of the micrograph in one channel only. So to turn these back into greyscale images, just tile the one channel it corresponds into all channels. </p>\n  </blockquote>\n</blockquote>\n\n<p>Hi Alex, thank you again. I'm almost there though still trying to fully parse what you said. When you say tile what do you mean? Are you saying (for example) take the 3 channel green image JPG from protein atlas and just use image[:, :, 1]? Or add up the values?</p>\n\n<p>Apologies for being slow on the uptake, I have ~0 domain knowledge here though have heard of protein and cells before. ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438378,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-13T15:01:46.930000",
          "content": "<p>@Mark Worrall this is exactly what I am implying and if I am not mistaken this is also what Alex Lu suggested</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438379,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-13T15:06:05.107000",
          "content": "<p>Assuming you have the microtubule image:\ngreyscale[:, :, 0] = microtubule[:, :, 0]\ngreyscale[:, :, 1] = microtubule[:, :, 0]\ngreyscale[:, :, 2] = microtubule[:, :, 0]</p>\n\n<p>Repeat with the proper channel for nuclei/protein/ER images. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438387,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-13T15:20:12.997000",
          "content": "<p>Hmmmmm, I'm still none the wiser, why are we repeating the values in each greyscale image? Greyscale images have a single channel.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438390,
          "author_name": "Alex Lu",
          "author_url": "",
          "post_date": "2018-12-13T15:28:11.193000",
          "content": "<p>Then you can do this:\ngreyscale = microtubule[:, :, 0]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438393,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-13T15:33:58.487000",
          "content": "<p>Hi, @Mark\nattached file name is :Check Metadata for HPAv18.html (7.44 MB)\non the comment top.\nSince reply cell has no file-uploader...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438405,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-13T15:48:46.863000",
          "content": "<p>How they processed the image is written this thread:\nsome questions about dataset instrumentation, image properties\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68983\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68983</a></p>\n\n<p>@Casper W[provide single slice images, non-merged. No gamma correction is made...]\nhe also said[If it still unclear, please do not hesitate to ask.]\nSo, let us ask about this problem on the thread? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438460,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-13T18:00:54.550000",
          "content": "<p>Hi Alex, again :)</p>\n\n<p>Sorry to be a pain, but can we be super explicit, is the following what you mean:</p>\n\n<p>Pseudo algorithm:</p>\n\n<pre><code>def process_im(id):\n    hpa_image = np.zeros((512, 512, 3))\n    for i, c in enumerate(['red', 'green', 'blue']):\n        tmp_channel_im = download_from_hpa_site(id, c)  # returns a 3 channel JPG image \n        hpa_image[:,:,i] = tmp_channel_im[:,:,0]  # &lt;---- BIT I AM UNCLEAR ON\n    hpa_image.save()\n</code></pre>\n\n<p>Thanks in advance.</p>\n\n<p>Update: I think it should be <code>hpa_image[:,:,i] = tmp_channel_im[:,:,i]</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438473,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-13T18:53:37.703000",
          "content": "<blockquote>\n  <p>Update: I think it should be hpa_image[:,:,i] = tmp_channel_im[:,:,i]</p>\n</blockquote>\n\n<p>Yes. That is what I meant. For yellow, use <code>tmp_channel_im[:,:,0]</code> (red channel) or <code>tmp_channel_im[:,:,1]</code> (green channel)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440147,
          "author_name": "kickback",
          "author_url": "",
          "post_date": "2018-12-17T06:08:20.577000",
          "content": "<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a>,</p>\n\n<p>After 3 days of much pain I now understand why greyscale images are so important! I saw what happens if you don't use them, found the problem (after much searching down the wrong paths) and then noticed you had pointed me to the solution several days ago!</p>\n\n<p>Learning can be painful...</p>\n\n<p>tx</p>\n\n<p>Kickback</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440881,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-12-18T03:59:17.647000",
          "content": "<p>I am still unsure about the correct processing of the external HPA database.\nwill this code work properly?</p>\n\n<p>img = cv2.imread(fn, cv2.IMREAD_GRAYSCALE)</p>\n\n<p>ofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"</p>\n\n<p>img = cv2.resize(img, (512,512))</p>\n\n<p>cv2.imwrite(ofn, img)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441715,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-19T00:23:46.287000",
          "content": "<p>I think it will not work properly. You are reading in three channels and converting them all to a grayscale, using, I think, the L formula. You should be doing a straight conversion, channel by channel, according to the color. For yellow, use average of red and green:</p>\n\n<pre><code>        img_resized = np.array(img.resize((int(y), int(y)), Image.ANTIALIAS))   # PIL\n        img_gs = None\n        for j, c in enumerate(colors):\n            if color == c:\n                if j &lt; 3:\n                    img_gs = img_resized[:, :, j]\n                elif j == 3:\n                    img_gs = (img_resized[:, :, 1] + img_resized[:, :, 0]) / 2\n        img_gs = img_gs.astype(np.uint8)\n</code></pre>\n\n<p>One good QC is to use the images that are repeated normal test data and HPA data and compare the histograms with your HPA import - they won't be identical but should look similar.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441724,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-12-19T00:54:31.800000",
          "content": "<p>that was a good advice, to check the leak images. I see what you mean, although I will have to rearrange the code as I still save them as separate colours.</p>\n\n<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441770,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-19T03:33:41.103000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441773,
          "author_name": "amaia",
          "author_url": "",
          "post_date": "2018-12-19T03:48:30.247000",
          "content": "<p>When using opencv the default is BGR not RGB, so channel indexes above are incorrect. Red and blue are swapped.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441783,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-19T04:09:00.317000",
          "content": "<p>Weird - I was using PIL. That's where the histograms help - in the images I looked at  the blue is distinctive.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441786,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-12-19T04:14:09.963000",
          "content": "<p>Thank you, fixed</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441788,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-12-19T04:19:59.513000",
          "content": "<p>Because i am using image for each channel (don't ask me why, inertia i guess) I couldn't see this problem. I have compared, alas, green to green and they were very similar... unfortunate...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441861,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-12-19T07:15:00.017000",
          "content": "<p>I must say my training looks about the same, so either the resnet50 doesn't care too much or I still have mistake somewhere.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442360,
          "author_name": "Vaagn Minasian",
          "author_url": "",
          "post_date": "2018-12-19T21:27:01.817000",
          "content": "<p>I suppose, that blue should be img[:,:,0] and red should be img[:,:,2]. \nI iterated over same images of different channels and got following sum_along_channel/total_sum channel for \"blue\",\"green\",\"red\",\"yellow\" . So blue is mostly present in 0 channel, whereas red in 2 channel.\nPlease, correct me if I am wrong.</p>\n\n<p>[0.98039446 0.0079704  0.01163514]\n[0.10331002 0.79542022 0.10126977]\n[0.05202446 0.04258582 0.90538972]\n[0.0740666  0.46528988 0.46064352]\n&gt; <strong>FabianIsensee wrote</strong>\n&gt; \n&gt; &gt; Usually color order is rgb so\n&gt; for blue take image[:, :, 2]\n&gt; for red take image[:, :, 0]\n&gt; for green take image[:, :, 1]\n&gt; for yellow you can take either image[:, :, 0] or image[:, :, 1]. Yellow is created by putting the intensity values in both r and g channel. Due to image compression (which is pretty bad for the HPA images by the way. Shame on whoever did this), r and g channel of yellow images will not be identical. I did some visual as well as numeric comparisons though and they are almost identical to the point where just using the r channel for these images should be fine (image[:, :, 0])</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442601,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-12-20T07:22:16.303000",
          "content": "<p>Hi,\nnot quite sure how you get the channels ordered this way. For me all images are r - g - b so blue is in channel 2.</p>\n\n<p>In [1]: import matplotlib.pyplot as plt <br>\nIn [2]: a = plt.imread(\"9985_69_B8_2_blue.jpg\") <br>\nIn [3]: a.shape <br>\nOut[3]: (1728, 1728, 3)\nIn [4]: a[:, :, 0].mean() <br>\nOut[4]: 0.2910919817386831\nIn [5]: a[:, :, 1].mean() <br>\nOut[5]: 0.24670694819530178\nIn [6]: a[:, :, 2].mean() <br>\nOut[6]: 27.79723802940672</p>\n\n<p>In [7]: from PIL import Image <br>\nIn [8]: import numpy as np <br>\nIn [9]: a = np.array(Image.open(\"9985_69_B8_2_blue.jpg\")) <br>\nIn [10]: a[:, :, 0].mean() <br>\nOut[10]: 0.2910919817386831\nIn [11]: a[:, :, 1].mean() <br>\nOut[11]: 0.24670694819530178\nIn [12]: a[:, :, 2].mean() <br>\nOut[12]: 27.79723802940672</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442659,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2018-12-20T09:41:50.687000",
          "content": "<p>hi vaagn:\nI guess you use opencv? Just as what @amaia say: when using opencv the default is BGR not RGB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445568,
          "author_name": "roush",
          "author_url": "",
          "post_date": "2018-12-26T18:31:09.217000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445706,
          "author_name": "zymale",
          "author_url": "",
          "post_date": "2018-12-27T00:48:12.443000",
          "content": "<p>@Hogger@David Silva\nWhen I use your code in  kernel,  error of  5G space limit  was reported. \n <a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a> \nBut the kernel  above downloaded more than 11G of data.\nSince my computer is not around now, I want to use kernel to download HPA data and train. \nHow can I use more than 5G space? \nIs there anyone here who can help me? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 460336,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-01-23T13:40:58.337000",
          "content": "<p>hi david thanks \ni used your script to download some 32 k images after i applied some filters on list. \nWhen i append this data to Training data and start training my validation loss goes very High. I  averaged the stats of HPA and competition data to do normalization. \nWhat m i  missing so my val loss is sky rocketing. ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 464236,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-01-31T12:33:31.280000",
          "content": "<p>hi david is there any similar script to download Tiff files provided orgiginally at gcp storage bucket.\n<a href=\"https://console.cloud.google.com/storage/browser/kaggle-human-protein-atlas\">https://console.cloud.google.com/storage/browser/kaggle-human-protein-atlas</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 428136,
      "author_name": "Artem Toporov",
      "author_url": "",
      "post_date": "2018-11-26T20:20:35.807000",
      "content": "<p>12760 images (800x800 RGB).\ndata: <a href=\"https://www.kaggle.com/artemtprv/external-data-for-protein-atlas\">https://www.kaggle.com/artemtprv/external-data-for-protein-atlas</a>\nnotebook:  <a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a></p>",
      "votes": 9,
      "replies": [
        {
          "id": 436775,
          "author_name": "Yurii Rebryk",
          "author_url": "",
          "post_date": "2018-12-10T23:22:27.940000",
          "content": "<p>I used your notebook to create a <a href=\"https://gist.github.com/rebryk/4320304588c55f2287540457d6830da2\">script</a> for multithreaded data downloading.</p>\n\n<p>Usage example: <br>\n<code>python external_data.py --path ~/datasets/external --n_threads 4 --batch_size 128</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 445695,
          "author_name": "zymale",
          "author_url": "",
          "post_date": "2018-12-27T00:21:14.850000",
          "content": "<p>I found your data is 11.31 GB. But when I use kernel to download data, the   disk space is limited to 5G. More than 5G will make an error. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 431345,
      "author_name": "TomomiMoriyama",
      "author_url": "",
      "post_date": "2018-12-02T04:49:01.633000",
      "content": "<p>HPAv18 contains \"Different ENSG id, but same image paths\"\nHere is \"noisy labels\" list</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 431541,
      "author_name": "maettes",
      "author_url": "",
      "post_date": "2018-12-02T13:18:48.490000",
      "content": "<p>Regarding the data leakage of the 126 images Brian mentions: is it safe to assume that they will be removed for the final evaluation?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 431796,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-02T23:05:06.687000",
          "content": "<p>I agree with <a href=\"/maettes\">@maettes</a> .\nData leakage boosted my score 0.528 to 0.588.\nI'm not comfortable with my score,\nbecause it's nothing to with my models generalization ability.\nI's not fair, because not everyone access those data easily.</p>\n\n<p>I hope (HPA admin?) provide clean HPAv18 metadata, \nand make everyone can access more simple way.</p>\n\n<p>in that case, I hope metadata contains Cell type information.\nthis extra information enable us to train generative models \nthat focuses on protein location but ignores cell type. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 431832,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-03T01:31:27.647000",
          "content": "<p>The 126 definitely help my score. I didn't check recently but at the time I jumped from .360 to .440</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 433894,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T16:11:07.433000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 413046,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-31T07:50:46.417000",
      "content": "<p>external data:</p>\n\n<p><a href=\"https://www.proteinatlas.org/learn/dictionary/cell\">https://www.proteinatlas.org/learn/dictionary/cell</a></p>\n\n<p><a href=\"https://www.proteinatlas.org/humancell\">https://www.proteinatlas.org/humancell</a>\n<a href=\"http://cytoconference.org/2017/Program/Image-Analysis-Challenge.aspx\">http://cytoconference.org/2017/Program/Image-Analysis-Challenge.aspx</a>\n<a href=\"https://www.allencell.org/\">https://www.allencell.org/</a></p>\n\n<p><a href=\"https://github.com/CellProfiling/pytorch_integrated_cell\">https://github.com/CellProfiling/pytorch_integrated_cell</a></p>\n\n<p><a href=\"http://hpa.scoreboard.czi.technology/challenge/1\">http://hpa.scoreboard.czi.technology/challenge/1</a></p>\n\n<p>pre-train mpodel</p>\n\n<p><a href=\"https://github.com/fyu/drn\">https://github.com/fyu/drn</a></p>\n\n<p><a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 412194,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2018-10-29T18:53:08.427000",
      "content": "<p>Pretrained ResNets at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>Other models at: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 419935,
      "author_name": "Cape",
      "author_url": "",
      "post_date": "2018-11-12T19:36:38.847000",
      "content": "<p>If there is extra data, they could just make that available and easily downloadable. Having more labeled images is a huge advantage. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 413329,
      "author_name": "Pavel Iakubovskii (qubvel)",
      "author_url": "",
      "post_date": "2018-10-31T17:58:07.407000",
      "content": "<p>Keras pretrained ResNet(18,34,50,101,152) models <a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 446858,
          "author_name": "souraj",
          "author_url": "",
          "post_date": "2018-12-28T19:39:19.827000",
          "content": "<p>Hello,\ndo we have any specification on input images such as size (299 * 299 * 3 ?) of normalization?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 449896,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2019-01-03T23:15:59.080000",
      "content": "<p>similar external data and pretrained models as other teams, external data from <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a> and pretrained models from <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a> <a href=\"https://github.com/osmr/imgclsmob\">https://github.com/osmr/imgclsmob</a> <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 440641,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2018-12-17T20:18:59.813000",
      "content": "<p>Pretrained models from <a href=\"https://github.com/creafz/pytorch-cnn-finetune\">https://github.com/creafz/pytorch-cnn-finetune</a></p>\n\n<p>Thank you <a href=\"/creafz\">@creafz</a> for sharing great project.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 437252,
      "author_name": "FabianIsensee",
      "author_url": "",
      "post_date": "2018-12-11T16:07:24.943000",
      "content": "<p>External data : <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a>\nMany thanks to Brian and TomomiMoriyama!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 429599,
      "author_name": "David Wagner",
      "author_url": "",
      "post_date": "2018-11-29T03:56:25.360000",
      "content": "<ul>\n<li><p>InceptionV3 with torchvision pretrained weights <a href=\"https://github.com/pytorch/vision/blob/master/torchvision/models/inception.py\">https://github.com/pytorch/vision/blob/master/torchvision/models/inception.py</a></p></li>\n<li><p>HPA additional dataset with props to @TomomiMoriyama for the helpful CSV files. </p></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 424588,
      "author_name": "Christoph Neuner",
      "author_url": "",
      "post_date": "2018-11-20T11:19:07.893000",
      "content": "<p>using pretrained models from <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a> and <a href=\"https://github.com/pytorch/vision/tree/master/torchvision/models\">https://github.com/pytorch/vision/tree/master/torchvision/models</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423736,
      "author_name": "DStjhb",
      "author_url": "",
      "post_date": "2018-11-19T00:34:24.413000",
      "content": "<p>@Brian Not sure how interesting this is (as it's only ~1% of the test set), but the 126 test images you found on HPA v18 do have an over-representation of 16 (Cytokinetic bridge)... Maybe chance(?)<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/423736/10691/126_test.png\" alt=\"enter image description here\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 424317,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-19T22:04:04.830000",
          "content": "<p>I'm guessing it is chance. The overall distribution of the HPA images is similar to the training set. Within the HPA images that I downloaded. Anything less than a couple thousand examples will have a different distribution because of the 1000:1 class imbalance.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 412838,
      "author_name": "Aleksey Alekseev",
      "author_url": "",
      "post_date": "2018-10-30T21:25:49.237000",
      "content": "<p>Pretrained models at <a href=\"http://keras.io/applications\">http://keras.io/applications</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 442592,
      "author_name": "Kulbear",
      "author_url": "",
      "post_date": "2018-12-20T07:10:45.540000",
      "content": "<p>Pretrained models at: <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a></p>\n\n<p>and <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>\n\n<p>External data: <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 446367,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-28T00:54:01.083000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436838,
      "author_name": "LbyG",
      "author_url": "",
      "post_date": "2018-12-11T02:31:40.530000",
      "content": "<p>BNInception with pretrained-models.pytorch weights <a href=\"https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/bninception.py\">https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/bninception.py</a></p>\n\n<p>HPA additional dataset with props to @TomomiMoriyama for the helpful CSV files.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 412155,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-10-29T17:24:43.847000",
      "content": "<p>I am not using a pretrained model, but I am using external data. I can share this privately but do I need to share it in this thread as well?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 412177,
          "author_name": "Maggie",
          "author_url": "",
          "post_date": "2018-10-29T18:23:55.123000",
          "content": "<p>Hello Brian - you do need to list the external data you are using in this thread.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 412217,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-10-29T19:55:14.497000",
          "content": "<p>Using v18 of protein atlas from here: <a href=\"https://www.proteinatlas.org/about/download\">https://www.proteinatlas.org/about/download</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 418884,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-11-10T19:45:57.167000",
          "content": "<p>Could you please give a hint on how to download these data? Either I am too stupid to find it or they are not publicly available?</p>\n\n<p>Edit: Sorry for bothering you. I just understood how. Man that is as cumbersome as it gets. Parsing XML files for ~12k genes to get these images... They could have made that easier...\nNow excuse me while I'm off writing a bunch of python code to download all these images... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418905,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-10T21:00:29.943000",
          "content": "<p>Well since the cat is out of the bag...</p>\n\n<p>I downloaded the TSV file. Used the TSV file to download all the associated XML files (5GB). Concatenate all the xml files and wrote a stream parsing script to pull out the images and protein information, and then download it.</p>\n\n<p>I'm not going to make it too easy, python is not the language I know best. Even worse, I did it with the version of ruby that was install on my machine and not the latest, 1.9.</p>\n\n<p>This one downloads all xml from the tsv file:</p>\n\n<pre>require 'net/http'\n\ntext=File.open('subcellular_location.tsv').read\ntext.gsub!(/\\r\\n?/, \"\\n\")\ntext.each_line do |line|\n    split_line = line.split(\"\\t\")\n    puts \"LINE:#{split_line[0].to_s.strip}||#{split_line[3].to_s.strip}|#{split_line[4].to_s.strip}|#{split_line[5].to_s.strip}\"\n    Net::HTTP.start(\"www.proteinatlas.org\") do |http|\n        resp = http.get(\"/#{split_line[0].to_s.strip}.xml\")\n        open(\"xml/#{split_line[0].to_s.strip}.xml\", \"wb\") do |file|\n            file.write(resp.body)\n        end\n    end\nend\n</pre>\n\n<p>I used cat *.xml &gt; out.xml, which is consumed by this script:\n<a href=\"https://pastebin.com/hJDBB6D0\">https://pastebin.com/hJDBB6D0</a></p>\n\n<p><strong>Spoiler alert: 126 images of the test set can be found in the public HPA</strong></p>",
          "votes": 21,
          "replies": []
        },
        {
          "id": 418909,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-10T21:07:44.910000",
          "content": "<p>Also, unless you change this to run in a multithreaded way, it will take about 2-3 days to run.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 418929,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2018-11-10T22:16:19.293000",
          "content": "<p>Thank you very much! May I ask how many images are there for the minority classes?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418943,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-10T23:48:15.477000",
          "content": "<p>I don't have an exact count off hand, but the distribution is about the same. Here is a comparison with competition data on the left:\n<img src=\"http://brians.network/images/sourcevsaugment.png\" alt=\"\"></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 418959,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2018-11-11T00:56:34.260000",
          "content": "<p>Good to see that the training set is nearly doubled! The minority classes didn't get a lot more samples, which means that the label imbalance problem continues to exist.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418991,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2018-11-11T03:19:45.030000",
          "content": "<p>May I also ask if the end of the second script is missing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419002,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-11T04:09:10.747000",
          "content": "<p>Looks like it got cut off,  I'll try to paste it again when I get back to my computer. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 419025,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-11T05:25:50.650000",
          "content": "<p>Edited the post, it didn't like the xml tags that were being matched: <a href=\"https://pastebin.com/hJDBB6D0\">https://pastebin.com/hJDBB6D0</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 419032,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2018-11-11T05:42:07.050000",
          "content": "<p>Thank you very much! I checked each xml file and the website, but it seems that there are multiple images for a specific file. And since I know nothing about ruby, may I ask what which image is the one we are looking for?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419042,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-11T06:01:11.147000",
          "content": "<p>Wow, this forum doesn't like less than characters. Worked in the preview, and the whole comment disappeared. In quick summary, there are data nodes, those nodes have a location nodes containing the protein names. In the same data node there are urls that contain _red in them on their own line. Using regular expressions to match xml tags line by line instead of parsing the document as xml.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 420177,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-11-13T07:50:39.407000",
          "content": "<p>Thank you so much! How do you deal with the different verification types (uncertain, enhanced, approved)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420435,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-13T16:17:18.377000",
          "content": "<p>I drop any images that are labeled 'uncertain'</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 420437,
          "author_name": "FabianIsensee",
          "author_url": "",
          "post_date": "2018-11-13T16:18:49.633000",
          "content": "<p>Thanks! Is that already included in your code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420464,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-13T16:59:16.423000",
          "content": "<p>Not in this version. I had a version that was nicer, did multi threaded downloads and checked for uncertain. When I went looking for it, I had removed the wrong folder previously and only this older version was left :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 420773,
          "author_name": "Dennis Subachev",
          "author_url": "",
          "post_date": "2018-11-14T05:10:17.460000",
          "content": "<p>I assume that since the data is from the same source as the challenge data (assumption based on the names), that there is duplicate data. Did you find this to be the case? If so, how did you check for duplicate data. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420779,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-14T05:24:31.887000",
          "content": "<p>Yes, there is definitely duplicate data there. I used a program that is usually used for finding duplicate photos that would flag ones that were similar. Several hundred duplicates exist between the training set and the HPA images, you can also find answers to 126 of test entries in the HPA.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 423580,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-11-18T15:48:14.247000",
          "content": "<p>Do you know the class distribution for the 126 test images with labels that you found in this data? It might give us some idea whether the train/test distributions match, which could inform training strategies.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423719,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-18T23:26:41.727000",
          "content": "<p>Here is a ruby script that processes the xml file, skips 'uncertain', and downloads the images using 12 threads. Takes only a few hours: <a href=\"https://pastebin.com/zjvJH0ne\">https://pastebin.com/zjvJH0ne</a></p>\n\n<p><a href=\"/dstjhb\">@dstjhb</a> here are the values for the 126 that overlapped. I've not checked the distribution or anything:</p>\n\n<pre>'0 23'\n'14 16'\n'14 16'\n'0 16 25'\n'14 16'\n'19'\n'0 16'\n'16'\n'0 16 25'\n'5'\n'3'\n'16 25'\n'16 25'\n'0 16 25'\n'0 16'\n'4'\n'0 25'\n'5 16'\n'0 16'\n'16 25'\n'0 16 17 18'\n'2 11'\n'16 25'\n'14 16'\n'2 16'\n'0 16 17 18'\n'16 25'\n'16'\n'0 16'\n'2 16'\n'0 21'\n'5'\n'0 5'\n'0 16'\n'0 7'\n'0 21'\n'16 17 23'\n'5'\n'1'\n'17 25'\n'0 21'\n'0 16'\n'2 17'\n'17 25'\n'4'\n'17'\n'25'\n'25'\n'0 16'\n'9 10'\n'14'\n'0'\n'21 16'\n'2'\n'13'\n'4 21 26'\n'25'\n'13'\n'25'\n'4'\n'14'\n'14'\n'0'\n'21'\n'0'\n'16'\n'22 16 25'\n'21 16 19 25'\n'9 10'\n'14 16'\n'16 25'\n'2 16'\n'16'\n'7 17'\n'5'\n'15 25'\n'21 11 16'\n'12'\n'17'\n'17 25'\n'15'\n'4 21 17'\n'16 23'\n'4'\n'16'\n'17'\n'0 25'\n'0 19 25'\n'0 21'\n'16 25'\n'0 16'\n'1 2'\n'0 16 25'\n'23'\n'17 19'\n'22'\n'2'\n'16'\n'17 25'\n'0 16'\n'5'\n'0 14 18'\n'19'\n'7 25'\n'23'\n'12'\n'2 4'\n'16'\n'5'\n'0'\n'3'\n'16 17 23'\n'14 17 23'\n'0 25'\n'6 21'\n'0 16'\n'0 16 17'\n'0 16 25'\n'0 16 25'\n'14 16'\n'14 16'\n'14 16'\n'14 16 19'\n'14 16 25'\n'15 25'\n'15 25'\n</pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423738,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-11-19T00:37:53.747000",
          "content": "<p>Thanks. I couldn't find an image upload in 'reply', but see my comment below for 126 test image distribution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426599,
          "author_name": "Artem Toporov",
          "author_url": "",
          "post_date": "2018-11-23T14:12:44.340000",
          "content": "<p>load external data\n<a href=\"https://www.kaggle.com/artemtprv/load-external-data\">https://www.kaggle.com/artemtprv/load-external-data</a></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 426979,
          "author_name": "Maksim Rodin",
          "author_url": "",
          "post_date": "2018-11-24T09:42:21.933000",
          "content": "<p>Am I right, that there are maximum 800x800 images available?\nWhat do you do, if you have several images for one ID, like here <a href=\"https://www.proteinatlas.org/ENSG00000002834-LASP1/antibody#ICC\">https://www.proteinatlas.org/ENSG00000002834-LASP1/antibody#ICC</a> or here <a href=\"https://www.proteinatlas.org/ENSG00000134057-CCNB1/antibody\">https://www.proteinatlas.org/ENSG00000134057-CCNB1/antibody</a> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427055,
          "author_name": "Artem Toporov",
          "author_url": "",
          "post_date": "2018-11-24T12:50:12.957000",
          "content": "<p>Yes, images 800x800 RGB.\nI fixed this, now all images will load.(save like this ENSG00000001084-GCLC 0.png, ENSG00000001084-GCLC 1.png)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 427539,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-11-25T18:19:42.943000",
          "content": "<p>Sorry to come on this a bit late... I noticed that in the HPA data there seem to be 32 classes, compared to the 28 we have in the competition. How do you deal with this mismatch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427623,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-11-25T21:36:18.027000",
          "content": "<p>I removed classes that don't exist, and removed any images that only contain classes that don't exist. You could define 'Nucleus' as the same as Nucleoplasm, but that might not be technically correct... You don't lose too many images filtering either way; so I guess see what works best for you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427949,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-11-26T13:04:22.347000",
          "content": "<p>A word on \"uncertain images\": </p>\n\n<p>In the case of \"Uncertain\" fluorescent images: they are only \"uncertain\" because they are not corroborated by other experimental data (e.g. gene expression).</p>\n\n<p>The IF images state: \"Location not consistent with experimental gene/protein characterization data.\" in these cases. The probe localisation itself is actually correct (just the antibody may be staining the wrong thing). </p>\n\n<p>Therefore, for the purpose of simply classifying the stain location, there is no need to exclude \"uncertain\" images (in my opinion). You still have to be careful, as some classes may be visible in IHC but not fluorescence data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 428107,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-26T19:09:31.570000",
          "content": "<p>The images I downloaded were not 800x800, but closer to 2048x2048 in jpg format. I saved only files that matched to the 28 classes, I didn't notice that there were 32. I've dropped all uncertain classes anyways just out of caution, maybe 15-20% of the extra images were removed when I did this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428685,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-11-27T18:04:16.690000",
          "content": "<p>Artem's webscraped ones are also either cropped or at a different magnification to original train set. Simpler to download though, as you don't have to parse the XML files for the links to the huge 2048p images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 430093,
          "author_name": "FabSchreiber",
          "author_url": "",
          "post_date": "2018-11-29T20:22:02.787000",
          "content": "<p>Hi Brian. I understand the images are single files with 3 channels. Did you find a way to use these additional images with a 4-Channel model (that you might be using for the official dataset)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430133,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-29T22:04:20.640000",
          "content": "<p>I'm only using rgb from the competition data,  combined into one file to match the HPA data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 430828,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-01T02:31:31.607000",
          "content": "<blockquote>\n  <p>4-Channel  RGBY</p>\n</blockquote>\n\n<p>imageURL.replace(blue_red_green,color)</p>\n\n<p>e.g.\n<a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_blue_red_green.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_blue_red_green.jpg</a></p>\n\n<p><a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_red.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_red.jpg</a>\n<a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_green.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_green.jpg</a>\n<a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_blue.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_blue.jpg</a></p>\n\n<p><a href=\"http://v18.proteinatlas.org/images/4109/24_H11_2_yellow.jpg\">http://v18.proteinatlas.org/images/4109/24_H11_2_yellow.jpg</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 431363,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-02T05:21:23.927000",
          "content": "<p>Thanks for this, I'm using it now to repeat my previous tests with the Y channel included. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 455529,
      "author_name": "steveagry",
      "author_url": "",
      "post_date": "2019-01-14T05:52:15.300000",
      "content": "<p>@Hogger\nUse your code for multi-threaded download, download will not download in the middle, is it necessary to use VPN?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454786,
      "author_name": "GuGu",
      "author_url": "",
      "post_date": "2019-01-12T07:51:20.430000",
      "content": "<p>Fastai pretrained models: <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 453947,
      "author_name": "K.Duan",
      "author_url": "",
      "post_date": "2019-01-11T01:13:11.663000",
      "content": "<p>External data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a>\npre-trained models:<a href=\"http://download.tensorflow.org/models/inception_v4_2016_09_09.tar.gz\">http://download.tensorflow.org/models/inception_v4_2016_09_09.tar.gz</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 453646,
      "author_name": "Shaohan Hu",
      "author_url": "",
      "post_date": "2019-01-10T14:28:58.073000",
      "content": "<p>data：kaggle data(512×512) and hpa data (<a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a>)\npretrainedmodels: <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 453380,
      "author_name": "Phil Butcher",
      "author_url": "",
      "post_date": "2019-01-10T05:12:20.783000",
      "content": "<p><a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 453042,
      "author_name": "FlYM",
      "author_url": "",
      "post_date": "2019-01-09T15:17:21.483000",
      "content": "<p>As everyone else. But just in case:\n<a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nExternal data: <a href=\"http://www.proteinatlas.org/\">http://www.proteinatlas.org/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 452722,
      "author_name": "tascj",
      "author_url": "",
      "post_date": "2019-01-09T04:48:31.720000",
      "content": "<p><a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n<a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\n<a href=\"https://gluon-cv.mxnet.io/\">https://gluon-cv.mxnet.io/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 452460,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2019-01-08T18:53:26.563000",
      "content": "<p>Fastai pretrained models: <a href=\"https://github.com/fastai/fastai\">https://github.com/fastai/fastai</a>\nExternal data: <a href=\"https://www.proteinatlas.org/\">https://www.proteinatlas.org/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 452155,
      "author_name": "seigato",
      "author_url": "",
      "post_date": "2019-01-08T09:30:32.580000",
      "content": "<p>External Data:  <a href=\"https://www.proteinatlas.org\">https://www.proteinatlas.org</a> \nPretrained Models:  <a href=\"https://github.com/pytorch/vision\">https://github.com/pytorch/vision</a>    <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451993,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-08T02:33:32.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451900,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-07T21:42:34.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451548,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-07T08:54:25.097000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451206,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-06T15:30:49.047000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450948,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-06T05:56:19.520000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450737,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-05T15:49:47.833000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450696,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-05T14:45:13.327000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450436,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-05T00:12:43.790000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450391,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T20:52:11.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450390,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T20:51:58.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450379,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T20:19:58.660000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 454341,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-11T13:26:08.757000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 450297,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T16:41:50.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449913,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T23:51:28.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449851,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T21:29:41.407000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 453327,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-10T02:39:35.073000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 449780,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T18:20:51.990000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449611,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T12:33:35.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449475,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T07:47:33.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449465,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T07:24:52.427000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449448,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T06:46:53.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449364,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-03T03:03:02.757000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449240,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-02T21:03:02.200000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449046,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-02T15:22:18.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448890,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-02T10:24:50.720000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448558,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-01T13:42:43.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448282,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-31T17:16:20.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 447940,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-30T21:53:19.453000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 447643,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-30T08:14:16.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 447642,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-30T08:14:12.130000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 447288,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-29T15:06:39.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446865,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T19:57:04.433000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446652,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T12:41:57.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446647,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T12:36:00.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446506,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T07:50:11.960000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446492,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T07:13:56.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446127,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-27T15:02:36.783000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445384,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-26T10:34:12.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445298,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-26T06:55:44.227000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445200,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-25T23:26:21.883000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444979,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-25T09:00:22.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444695,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-24T15:35:49.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444433,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-24T02:55:49.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444432,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-24T02:51:38.017000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444202,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-23T14:06:09.720000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444142,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-23T10:30:07.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 443886,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-22T16:55:41.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 443177,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-21T06:46:02.577000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 442796,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-20T14:14:40.243000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 442859,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-20T16:02:06.590000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442519,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-20T03:53:38.693000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441687,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-18T23:35:01.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 440442,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-17T14:46:55.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 440396,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-17T13:55:13.043000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 440270,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-17T10:07:12.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 440002,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-16T20:55:48.930000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 439290,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-15T05:10:22.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 438736,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-14T05:27:09.140000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436927,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-11T05:26:34.473000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436713,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-10T19:59:54.090000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436115,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-09T16:15:06.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435756,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T17:41:57.750000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435432,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T03:03:41.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435087,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T13:27:54.343000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435057,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T12:29:59.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 434820,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T02:07:43.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433913,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-05T16:40:21.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 432899,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-04T12:49:47.657000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 432974,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-04T14:34:24.907000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 432626,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-04T05:40:17.853000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 432849,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-04T12:02:58.420000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 432873,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-04T12:19:29.443000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433089,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-04T17:19:41.240000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433116,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-04T18:06:59.557000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433219,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-04T21:25:34.490000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 431868,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-03T03:31:08.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 430414,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-30T11:05:07.473000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 429650,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-29T06:03:43.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 429322,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-28T17:24:25.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 427131,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-24T17:16:02.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426702,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-23T17:12:37.893000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423701,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-18T22:27:43.280000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423549,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-18T14:16:01.217000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 420810,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-14T06:45:59.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 420174,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-13T07:45:27.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418998,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-11T03:55:35.993000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415361,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-05T01:53:26.583000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 414427,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-02T18:46:53.720000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 413339,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-31T18:38:14.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 452551,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-08T22:02:39.437000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448616,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-01T17:07:38.717000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 442282,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-19T18:20:09.643000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 443261,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-21T10:08:26.550000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 445011,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-25T11:02:04.310000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446689,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-28T14:11:42.877000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 449638,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-03T13:33:21.210000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "430860": "Hi! @FabSchreiber \n\nHere's list\n\n!! still contains duplicated images\n　RGB images(HPAv18.csv) :  77,878 sample NG\n　RGB images 77,864 sample  ...deleted id duplication,still contains duplicated images\n　RGBY images  77,430 sample\n　RGBY withoutUncertain:73,881 sample\n\nold csv list contains\n GeneID_ Dir_ImageURL Target(28 class)\n\n\"http://v18.proteinatlas.org/images/\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)\n\n**Modified:** ..without duplicated images (Gene information lost,labels merged)\n　RGB\\_wodpl 75,040 sample\n　RGBY\\_wodpl  74,606 sample\n　RGBY withoutUncertain_wodpl:71,437 sample\n\nnew csv list contains\n　Dir_ImageURL Target(28 class)\n\n**AddFiles:** @Chase\\_the\\_Trane adviced me to open how I make those csv files.\n　 1. Parce XML and Download HPAv18 Image.html\n--&gt;total downloaded image 60GB\n　  2.  NoYellow512.txt\n--&gt;resize img 512x512 png 70GB, and I realized some sample has no yellow filter.\n　  3. Make Metadata for HPAv18 Image.html\n--&gt;duplicate image exist\n　 4. Clean HPAv18 dataset.html...How I merged the labels.\n\n**AddFiles2:** add CellLine information \n*\\_withCellLine.csv",
    "412120": "Please post which pre-trained model(s) you are using in this thread.\n\n&gt; EXTERNAL DATA\nYou may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. Pre-trained models may be used to construct the algorithms. Please specify which pre-trained model(s) you are using via specified discussion post.",
    "436319": "Those who read this comment, please help me.\nI want to use some open source code and need some advice about the usage...\n\n\nOne month a go, @hengck23 introduced this paper\n\nLearning unsupervised feature representations for single cell microscopy images with paired cell inpainting\nAlex Lu, Oren Z Kraus, Sam Cooper, Alan M Moses\n\non [ideas and discussion] thread\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69955\n\nI found it very attractive, but source code was not released at the time.\n\nOne week ago, source code has finally released!\nhttps://github.com/alexxijielu/paired_cell_inpainting#human-model\n\nIt contains \n- Cleaner HPA data downloading scripts.\n- Extracting single cell.\n- pre-trained weights \n- and more\n\nTheir task setting(unsupervised) is different from us(supervised classification),\nbut can be used with some extra code.\n\nThey released the code under [GPL-2.0]\n\nMy question is, \n\"Is it OK to use this code and pre-trained weight as a baseline model?\"\n if it's against the rule \"Can I re-implement from scratch and imitate the idea?\"\n if it's within the rule \"Should I ask the author of this code about the usage in Github/Issues section beforehand?\"\nOr they might also participated in this competition already!?",
    "432870": "Hi, @pete\nXML parsing is time consuming.\nAlternatively, you can download from csv list.\n\n(How I processed the list was posted few comments ago.\nMy code is far from efficient and safe, there's no Error handrring, though )\n\n**Note 12/13:**\n**Please use @davidwagnerkc 's downloading script !**\nNot gray-scaled image hurts me:11b_Segment_and_Crop.html\nToo many class 27?:Check metadata for HPAv18.html",
    "428136": "12760 images (800x800 RGB).\ndata: https://www.kaggle.com/artemtprv/external-data-for-protein-atlas\nnotebook:  https://www.kaggle.com/artemtprv/load-external-data",
    "431345": "HPAv18 contains \"Different ENSG id, but same image paths\"\nHere is \"noisy labels\" list",
    "431541": "Regarding the data leakage of the 126 images Brian mentions: is it safe to assume that they will be removed for the final evaluation?",
    "413046": "external data:\n\nhttps://www.proteinatlas.org/learn/dictionary/cell\n\nhttps://www.proteinatlas.org/humancell\nhttp://cytoconference.org/2017/Program/Image-Analysis-Challenge.aspx\nhttps://www.allencell.org/\n\nhttps://github.com/CellProfiling/pytorch_integrated_cell\n\n\nhttp://hpa.scoreboard.czi.technology/challenge/1\n\npre-train mpodel\n\nhttps://github.com/fyu/drn\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n",
    "412194": "Pretrained ResNets at: https://github.com/pytorch/vision\n\nOther models at: https://github.com/Cadene/pretrained-models.pytorch",
    "419935": "If there is extra data, they could just make that available and easily downloadable. Having more labeled images is a huge advantage. ",
    "413329": "Keras pretrained ResNet(18,34,50,101,152) models https://github.com/qubvel/classification_models",
    "449896": "similar external data and pretrained models as other teams, external data from https://www.proteinatlas.org and pretrained models from https://github.com/pytorch/vision https://github.com/osmr/imgclsmob https://github.com/Cadene/pretrained-models.pytorch ",
    "440641": "Pretrained models from https://github.com/creafz/pytorch-cnn-finetune\n\nThank you @creafz for sharing great project.",
    "437252": "External data : https://www.proteinatlas.org\nMany thanks to Brian and TomomiMoriyama!!!",
    "429599": "+ InceptionV3 with torchvision pretrained weights https://github.com/pytorch/vision/blob/master/torchvision/models/inception.py\n\n+ HPA additional dataset with props to @TomomiMoriyama for the helpful CSV files. ",
    "424588": "using pretrained models from https://github.com/fastai/fastai and https://github.com/pytorch/vision/tree/master/torchvision/models",
    "423736": "@Brian Not sure how interesting this is (as it's only ~1% of the test set), but the 126 test images you found on HPA v18 do have an over-representation of 16 (Cytokinetic bridge)... Maybe chance(?)![enter image description here][1]\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/423736/10691/126_test.png",
    "412838": "Pretrained models at http://keras.io/applications",
    "442592": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "436838": "BNInception with pretrained-models.pytorch weights https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/bninception.py\n\nHPA additional dataset with props to @TomomiMoriyama for the helpful CSV files.",
    "412155": "I am not using a pretrained model, but I am using external data. I can share this privately but do I need to share it in this thread as well?",
    "455529": "@Hogger\nUse your code for multi-threaded download, download will not download in the middle, is it necessary to use VPN?",
    "454786": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: http://www.proteinatlas.org/",
    "453947": "External data: http://www.proteinatlas.org/\npre-trained models:http://download.tensorflow.org/models/inception_v4_2016_09_09.tar.gz",
    "453646": "data：kaggle data(512×512) and hpa data (http://www.proteinatlas.org/)\npretrainedmodels: https://github.com/Cadene/pretrained-models.pytorch",
    "453380": "https://github.com/pytorch/vision\nExternal data: http://www.proteinatlas.org/",
    "453042": "As everyone else. But just in case:\nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nExternal data: http://www.proteinatlas.org/",
    "452722": "https://www.proteinatlas.org/\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://keras.io/applications/\nhttps://gluon-cv.mxnet.io/",
    "452460": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: https://www.proteinatlas.org/",
    "452155": "External Data:  https://www.proteinatlas.org \nPretrained Models:  https://github.com/pytorch/vision    https://github.com/Cadene/pretrained-models.pytorch",
    "451993": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: https://www.proteinatlas.org/",
    "451900": "https://github.com/CellProfiling/FeatureExtraction\nhttps://github.com/CellProfiling/Loc-CAT",
    "451548": "Pretrained PyramidNet https://github.com/dyhan0920/PyramidNet-PyTorch",
    "451206": "Pretrained models from https://github.com/fastai/fastai and https://github.com/pytorch/vision/tree/master/torchvision/models\n\nExternal data: https://www.proteinatlas.org",
    "450948": "Pretrained model\nhttps://github.com/qubvel/classification_models\nhttps://github.com/fchollet/deep-learning-models/releases\n\nExternal data\nhttps://www.proteinatlas.org",
    "450737": "Apart from normal PyTorch pretrained models, pretrained models from Cadene and external data from HPA, I might use CBAM-ResNet50 from https://github.com/Jongchan/attention-module",
    "450696": "Pretrained models from: https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data from: http://v18.proteinatlas.org/images",
    "450436": "Pretrained:\n\nhttps://github.com/pytorch/vision/\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data (planned)\n\nhttps://www.proteinatlas.org/",
    "450391": "Fastai pretrained models: https://github.com/fastai/fastai\nExternal data: https://www.proteinatlas.org/",
    "450390": "All of the below. ",
    "450379": "https://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nExternal data: http://www.proteinatlas.org/\n",
    "450297": "Pretrained models at: https://github.com/pytorch/vision and https://github.com/Cadene/pretrained-models.pytorch\nExternal data: https://www.proteinatlas.org",
    "449913": "https://www.proteinatlas.org\nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://keras.io/applications/",
    "449851": "https://keras.io/applications/\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://github.com/qubvel/classification_models\nhttps://www.proteinatlas.org",
    "449780": "Pretrained models : \nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\nhttps://keras.io/applications/\n\nExternal data:\nhttps://www.proteinatlas.org\n",
    "449611": "Pretrained models : \nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data:\nhttps://www.proteinatlas.org",
    "449475": "External data: https://www.proteinatlas.org\nPretrained models:  https://github.com/pytorch/vision",
    "449465": "Using pretrained models from http://files.fast.ai/models/ and external data from https://www.proteinatlas.org/",
    "449448": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "449364": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "449240": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "449046": "External data: https://www.proteinatlas.org",
    "448890": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "448558": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org\npretrained models: https://keras.io/applications/\nexternal data: https://www.proteinatlas.org. Credits to TomomiMoriyama (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860) and David Silva (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386)",
    "448282": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "447940": "I'm using the kernel as provided here, with some modifications:\nhttps://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\nSo resnet 34 and fastai. \n\nAlong with the external data obtained using the script given by hogger here:\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\n\nRest is some scripts written by me for processing of data, sorting and whatnot.",
    "447643": "pretrained models: https://keras.io/applications/\nexternal data: https://www.proteinatlas.org. Credits to TomomiMoriyama (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860) and David Silva (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#437386)",
    "447642": "Pretrained models at: https://pypi.org/project/pretrainedmodels/  and  https://pypi.org/project/cnn-finetune/0.1.5/\n\nexternal data : https://www.proteinatlas.org",
    "447288": "Pretrained models:\nhttps://github.com/pytorch/vision\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data:\nhttps://www.proteinatlas.org",
    "446865": "https://github.com/qubvel/classification_models",
    "446652": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "446647": "Pretrained models :\nhttps://github.com/pytorch/vision\n\nExternal data: \nhttps://www.proteinatlas.org",
    "446506": "Pretrained models : https://github.com/pytorch/vision\nExternal data: https://www.proteinatlas.org",
    "446492": "Pretrained ResNets/Densenet/InceptionV3 at: https://github.com/pytorch/vision\nExploring other models at: https://github.com/Cadene/pretrained-models.pytorch\nexternal data from https://www.proteinatlas.org",
    "446127": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "445384": "Pretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "445298": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "445200": "Pretrained models from https://github.com/fastai/fastai\nhttps://pytorch.org/docs/stable/torchvision/models.html\nexternal data from: https://www.proteinatlas.org",
    "444979": "External data from: https://www.proteinatlas.org",
    "444695": "Using pre-trained models (inceptionV4 and resnext50) from https://github.com/Cadene/pretrained-models.pytorch  and additional external data from  https://www.proteinatlas.org",
    "444433": "Has anyone uploaded the extra HPA RGBY images to Kaggle or other fast cloud storage to avoid traffic loads on HPA servers?",
    "444432": "Pretrained models: https://github.com/pytorch/vision &amp; https://github.com/Cadene/pretrained-models.pytorch\nExternal data: https://www.proteinatlas.org (as shared by @TomomiMoriyama)",
    "444202": "as everyone..\nPretrained models : https://github.com/pytorch/vision\n\nhttps://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "444142": "pytorch models: https://github.com/pytorch/vision, https://github.com/Cadene/pretrained-models.pytorch\nexternal data: https://www.proteinatlas.org",
    "443886": "Pretrained models from https://github.com/fastai/fastai ",
    "443177": "Pretrained models from https://github.com/fastai/fastai \nand\nhttps://github.com/pytorch/vision/\n\nExternal data from HPAv18",
    "442796": "Hi everyone:\nI found there are two version of external data, which one should I use?\n1. HPA v18 share by @TomomiMoriyama\n2. 12760 images (800x800 RGB) shared by @Artem Toporov in this kernel: https://www.kaggle.com/artemtprv/load-external-data\nMaybe I am wrong, maybe these two are the same data. I did not look at these two data carefully, I just found that their image names are different. I hope someone can answer me.",
    "442519": "Pretrained ResNets at: https://github.com/pytorch/vision\nOther models at: https://github.com/Cadene/pretrained-models.pytorch\nexternal data from  https://www.proteinatlas.org",
    "441687": "I am using keras and pytorch fastai pre-trained models. ",
    "440442": " - pretrained pytorch models from the pretrainedmodels library\n - additional HPA data from https://www.proteinatlas.org, with the CSV files provided by @TomomiMoriyama ",
    "440396": "I'm using pretrained resnet pytorch model's weights via fastai library interface:\nhttps://pytorch.org/docs/stable/torchvision/models.html\nhttps://github.com/fastai/fastai\nextra data:https://www.proteinatlas.org",
    "440270": "External data at https://www.proteinatlas.org\nPretrain Model: https://keras.io/applications/",
    "440002": "Keras pretrained models: http://keras.io/applications\nExternal data: https://www.proteinatlas.org",
    "439290": "Pretrained models at http://models.tensorpack.com/FasterRCNN/\nExternal data at https://www.proteinatlas.org",
    "438736": "Pretrained models at https://github.com/chainer/chainercv  \nExternal data at https://www.proteinatlas.org",
    "436927": "Pretrained models at https://github.com/Cadene/pretrained-models.pytorch\nExternal data at  https://www.proteinatlas.org",
    "436713": "Pretrained models at: https://github.com/pytorch/vision\n\nand https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data: https://www.proteinatlas.org",
    "436115": "Pretrained models at: https://github.com/pytorch/vision\n\nand  https://github.com/Cadene/pretrained-models.pytorch\n\nExternal data : https://www.proteinatlas.org",
    "435756": "pretrained models from \nhttps://github.com/fastai/fastai \nand\nhttps://github.com/pytorch/vision/\n\nexternal data: https://www.proteinatlas.org",
    "435432": "Keras pretrained models (http://keras.io/applications)\nExternal data https://www.proteinatlas.org/",
    "435087": "Keras pretrained models (http://keras.io/applications)\nExternal data https://www.proteinatlas.org/",
    "435057": "Using pretrained Resnet weights from fast.ai and Keras. Other than that I am also using other pretrained weights from architectures like xception and DenseNets on keras.",
    "434820": "Keras pretrained models on ImageNet data (http://keras.io/applications)",
    "433913": "Keras pretrained models (http://keras.io/applications)\nProbably will use https://www.proteinatlas.org/\n\n",
    "432899": "Are kaggle or the organisers going to comment on the data leak?\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73395",
    "432626": "I am looking at the site and in particular at the example url:\nhttps://www.proteinatlas.org/ENSG00000134057.xml\n\nLooking in this file, I see 32 mentions of _red_green_blue and two mentions of _red_green_blue_yellow.\n\nAll the mentions appear to be 3 channel, so I don't know how to find yellow channels. Also, which of the multiple channels are correct.\n\nPerhaps more to the point: am I looking at the correct xml and how does one interpret it?\n",
    "431868": "\npretrained models from \nhttps://github.com/fastai/fastai \nand\nhttps://github.com/pytorch/vision/\n\nHi Martin\nCan I use GAN to augment the train data??",
    "430414": "Pretrained:\n- https://github.com/pytorch/vision/\n- https://github.com/Cadene/pretrained-models.pytorch\n\nPlan to add external data:\n- https://www.proteinatlas.org/",
    "429650": "I'm using keras MobileNet: https://keras.io/applications/#mobilenet\nkernel: https://www.kaggle.com/amneves/keras-proteins",
    "429322": "external data: https://www.proteinatlas.org",
    "427131": "Pretrained ResNets_34 at: https://github.com/pytorch/vision",
    "426702": "I am new to this, need some direction. There are 4 channels (each with different information - blue,red,yellow with organelle location information, and green with location of proteins) . The goal as I understand is to identify an image with the location of the protein.  If I was to combine the images and then train, can I also combine the images in the test set to validate and find accuracy of the model. In other words can the model be built that will only predict on images where all channels are combined into a single image ?",
    "423701": "Keras pretrained models on ImageNet data (http://keras.io/applications)",
    "423549": "Inception_v4 and weights... https://github.com/kentsommer/keras-inceptionV4/releases",
    "420810": "Im using the keras pretrained resnet50 weights.",
    "420174": "I'm using pretrained resnet pytorch model's weights via fastai library interface:\nhttps://pytorch.org/docs/stable/torchvision/models.html\nhttps://github.com/fastai/fastai\n",
    "418998": "Pre-trained ResNext models at: https://github.com/Cadene/pretrained-models.pytorch",
    "415361": "Pretrained models:\nhttps://github.com/pytorch/vision/\nhttps://github.com/Cadene/pretrained-models.pytorch",
    "414427": "I am using pretrained weights from https://github.com/fchollet/deep-learning-models/releases as well",
    "413339": "I am using some pre-trained keras model weights from:\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.7/inception_resnet_v2_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.5/inception_v3_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/titu1994/Keras-NASNet/releases/download/v1.2/NASNet-large-no-top.h5\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.2/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/fchollet/deep-learning-models/releases/download/v0.4/xception_weights_tf_dim_ordering_tf_kernels_notop.h5\nhttps://github.com/keras-team/keras-applications/releases/download/densenet/densenet121_weights_tf_dim_ordering_tf_kernels_notop.h5",
    "452551": "",
    "448616": "",
    "442282": ""
  }
}