{
  "id": 70206,
  "title": "How I process the tif images",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70206",
  "author_name": "",
  "post_date": "2018-11-01T00:20:27.816691800Z",
  "votes": 14,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Starting to explore the tif files, they are quite large. Uncompressed the training set is over 600GB. To shrink this down I am converting them to RGB files stored as png. Here is how I am doing it. Locally it looks like this is going to take about 15 hours, so far it looks like it will decrease the size significantly. Especially so because I am using 3 channels instead of 4.</p>\n\n<pre># use convert to merge images to rgb\nfrom tqdm import tqdm_notebook\nfrom pathlib import Path\nimport os\n\ninput_path = Path(\"../input/human_protein_atlas\")\ntrain_csv = input_path / \"train.csv\"\n\ninput_path_large = Path(\"J:/hpa_images\")\ntrain_dir = input_path_large / \"train_full_size\"\ntrain_combined_dir = input_path_fast / \"train_combined_rgb\"\n\ndata = pd.read_csv(str(train_csv))\nids = data['Id'].tolist()\n\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(train_dir / img_id)\n    red_img = (img_path + \"_red.tif\")\n    yellow_img = (img_path + \"_yellow.tif\")\n    blue_img = (img_path + \"_blue.tif\")\n    green_img = (img_path + \"_green.tif\")\n\n    out_img = str(train_combined_dir / img_id) + \"_rgb.png\"\n\n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s  -set colorspace RGB -combine %s\" % (red_img,green_img,blue_img,out_img) \n    #print(cmd)\n    os.system(cmd)\n</pre>",
  "messages": [
    {
      "id": "413432",
      "postDate": "11/01/2018 00:20:27",
      "content": "<p>Starting to explore the tif files, they are quite large. Uncompressed the training set is over 600GB. To shrink this down I am converting them to RGB files stored as png. Here is how I am doing it. Locally it looks like this is going to take about 15 hours, so far it looks like it will decrease the size significantly. Especially so because I am using 3 channels instead of 4.</p>\n\n<pre># use convert to merge images to rgb\nfrom tqdm import tqdm_notebook\nfrom pathlib import Path\nimport os\n\ninput_path = Path(\"../input/human_protein_atlas\")\ntrain_csv = input_path / \"train.csv\"\n\ninput_path_large = Path(\"J:/hpa_images\")\ntrain_dir = input_path_large / \"train_full_size\"\ntrain_combined_dir = input_path_fast / \"train_combined_rgb\"\n\ndata = pd.read_csv(str(train_csv))\nids = data['Id'].tolist()\n\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(train_dir / img_id)\n    red_img = (img_path + \"_red.tif\")\n    yellow_img = (img_path + \"_yellow.tif\")\n    blue_img = (img_path + \"_blue.tif\")\n    green_img = (img_path + \"_green.tif\")\n\n    out_img = str(train_combined_dir / img_id) + \"_rgb.png\"\n\n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s  -set colorspace RGB -combine %s\" % (red_img,green_img,blue_img,out_img) \n    #print(cmd)\n    os.system(cmd)\n</pre>",
      "rawMarkdown": "Starting to explore the tif files, they are quite large. Uncompressed the training set is over 600GB. To shrink this down I am converting them to RGB files stored as png. Here is how I am doing it. Locally it looks like this is going to take about 15 hours, so far it looks like it will decrease the size significantly. Especially so because I am using 3 channels instead of 4.\n\n<pre># use convert to merge images to rgb\nfrom tqdm import tqdm_notebook\nfrom pathlib import Path\nimport os\n\ninput_path = Path(\"../input/human_protein_atlas\")\ntrain_csv = input_path / \"train.csv\"\n\ninput_path_large = Path(\"J:/hpa_images\")\ntrain_dir = input_path_large / \"train_full_size\"\ntrain_combined_dir = input_path_fast / \"train_combined_rgb\"\n\ndata = pd.read_csv(str(train_csv))\nids = data['Id'].tolist()\n\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(train_dir / img_id)\n    red_img = (img_path + \"_red.tif\")\n    yellow_img = (img_path + \"_yellow.tif\")\n    blue_img = (img_path + \"_blue.tif\")\n    green_img = (img_path + \"_green.tif\")\n    \n    out_img = str(train_combined_dir / img_id) + \"_rgb.png\"\n    \n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s  -set colorspace RGB -combine %s\" % (red_img,green_img,blue_img,out_img) \n    #print(cmd)\n    os.system(cmd)\n</pre>",
      "votes": null
    },
    {
      "id": "413471",
      "postDate": "11/01/2018 02:21:04",
      "content": "<p>Actually, the images can be directly read from archive using zipfile library, so no need to extract them first. As an example, the following function, used in another competition, reads images from archive 1 by 1, resizes them, and writes into another archive:</p>\n\n<pre><code>def resize(filename, sz):\n    with zipfile.ZipFile(os.path.join(PATH, f'{filename}.zip'), 'r') as archive, \\\n      zipfile.ZipFile(os.path.join(PATH, filename + str(sz) + '.zip'), 'w') as archive_out:\n        for name in archive.namelist():\n            img = Image.open(io.BytesIO(archive.read(name)))\n            output = io.BytesIO()\n            img.resize((sz,sz)).save(output, format='png')\n            archive_out.writestr(name, output.getvalue())\n</code></pre>",
      "rawMarkdown": "Actually, the images can be directly read from archive using zipfile library, so no need to extract them first. As an example, the following function, used in another competition, reads images from archive 1 by 1, resizes them, and writes into another archive:\n\n    def resize(filename, sz):\n        with zipfile.ZipFile(os.path.join(PATH, f'{filename}.zip'), 'r') as archive, \\\n          zipfile.ZipFile(os.path.join(PATH, filename + str(sz) + '.zip'), 'w') as archive_out:\n            for name in archive.namelist():\n                img = Image.open(io.BytesIO(archive.read(name)))\n                output = io.BytesIO()\n                img.resize((sz,sz)).save(output, format='png')\n                archive_out.writestr(name, output.getvalue())",
      "votes": null
    },
    {
      "id": "413515",
      "postDate": "11/01/2018 04:06:34",
      "content": "<p>Assuming that you have ImageMagic's <code>convert</code> installed, a single line will do the same trick on Linux. Just go to the appropriate directory and paste:</p>\n\n<p><code>find . -name \"*_blue.tif\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.tif//g' | xargs -i convert \"{}\"_red.tif \"{}\"_green.tif \"{}\"_blue.tif -set colorspace RGB -combine \"{}\"_rgb.tif</code></p>\n\n<p>Obviously, the same can be done with .png files:</p>\n\n<p><code>find . -name \"*_blue.png\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.png//g' | xargs -i convert \"{}\"_red.png \"{}\"_green.png \"{}\"_blue.png -set colorspace RGB -combine \"{}\"_rgb.png</code></p>",
      "rawMarkdown": "Assuming that you have ImageMagic's `convert` installed, a single line will do the same trick on Linux. Just go to the appropriate directory and paste:\n\n`find . -name \"*_blue.tif\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.tif//g' | xargs -i convert \"{}\"_red.tif \"{}\"_green.tif \"{}\"_blue.tif -set colorspace RGB -combine \"{}\"_rgb.tif`\n\nObviously, the same can be done with .png files:\n\n`find . -name \"*_blue.png\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.png//g' | xargs -i convert \"{}\"_red.png \"{}\"_green.png \"{}\"_blue.png -set colorspace RGB -combine \"{}\"_rgb.png`",
      "votes": null
    },
    {
      "id": "413521",
      "postDate": "11/01/2018 04:18:41",
      "content": "<p>I ended up switching away from linux to my gaming computer because of the graphics card. This probably will work under git bash or cygwin too.</p>",
      "rawMarkdown": "I ended up switching away from linux to my gaming computer because of the graphics card. This probably will work under git bash or cygwin too.",
      "votes": null
    },
    {
      "id": "413523",
      "postDate": "11/01/2018 04:22:13",
      "content": "<p>I might be able to do something like this. I'm using Keras currently, it seemed easier to combine them first externally. </p>",
      "rawMarkdown": "I might be able to do something like this. I'm using Keras currently, it seemed easier to combine them first externally.",
      "votes": null
    },
    {
      "id": "413938",
      "postDate": "11/01/2018 20:28:01",
      "content": "<p>Overnight it finished and it looks like it has worked well.</p>\n\n<ul>\n<li>Before: 618GB, RGBY in separate tif</li>\n<li>After: 155GB, RGB png files. Y discarded.</li>\n</ul>",
      "rawMarkdown": "Overnight it finished and it looks like it has worked well.\n\n* Before: 618GB, RGBY in separate tif\n* After: 155GB, RGB png files. Y discarded.",
      "votes": null
    },
    {
      "id": "414004",
      "postDate": "11/02/2018 00:06:22",
      "content": "<p>Using 44GB memory: 3108 validation images, 1024x1024x3 float32(because keras is broke at float16). </p>",
      "rawMarkdown": "Using 44GB memory: 3108 validation images, 1024x1024x3 float32(because keras is broke at float16).",
      "votes": null
    },
    {
      "id": "414070",
      "postDate": "11/02/2018 03:37:23",
      "content": "<p>Limit reached, I don't think my drive can go any faster</p>",
      "rawMarkdown": "Limit reached, I don't think my drive can go any faster",
      "votes": null
    },
    {
      "id": "416677",
      "postDate": "11/07/2018 04:28:28",
      "content": "<p>Full size TIF data is 7z compressed...</p>",
      "rawMarkdown": "Full size TIF data is 7z compressed...",
      "votes": null
    },
    {
      "id": "417916",
      "postDate": "11/09/2018 01:19:29",
      "content": "<p>I did it in a parallelized fashion:</p>\n\n<pre><code>def save_to_dir(i):\n    if(i%1000==0): print(\"loaded sample {}\".format(i))\n    imid=df['Id'][i]\n    r = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_red.tif\"))\n    g = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_green.tif\"))\n    b = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_blue.tif\"))\n    y = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_yellow.tif\"))\n\n    image = np.dstack((r,g,b,y))\n    np.save(ROOT+\"train_full_size_np/\"+imid,image)\nnum_cores = 8\nParallel(n_jobs=num_cores, prefer=\"threads\")(delayed(save_to_dir)(i) for i in range(len(df)))\n</code></pre>",
      "rawMarkdown": "I did it in a parallelized fashion:\n\n    def save_to_dir(i):\n        if(i%1000==0): print(\"loaded sample {}\".format(i))\n        imid=df['Id'][i]\n        r = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_red.tif\"))\n        g = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_green.tif\"))\n        b = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_blue.tif\"))\n        y = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_yellow.tif\"))\n\n        image = np.dstack((r,g,b,y))\n        np.save(ROOT+\"train_full_size_np/\"+imid,image)\n    num_cores = 8\n    Parallel(n_jobs=num_cores, prefer=\"threads\")(delayed(save_to_dir)(i) for i in range(len(df)))",
      "votes": null
    },
    {
      "id": "417990",
      "postDate": "11/09/2018 04:47:52",
      "content": "<p>I tried to do something using the multiprocessing library but it didn't play along with Jupyter. I ended up just letting it run. For now I've just used mogrify and convert to make a few sets at different resolutions.</p>",
      "rawMarkdown": "I tried to do something using the multiprocessing library but it didn't play along with Jupyter. I ended up just letting it run. For now I've just used mogrify and convert to make a few sets at different resolutions.",
      "votes": null
    },
    {
      "id": "418750",
      "postDate": "11/10/2018 14:56:39",
      "content": "<p>Try concurrent.futures import ThreadPoolExecutor\nIt works fine with Jupyter</p>",
      "rawMarkdown": "Try concurrent.futures import ThreadPoolExecutor\nIt works fine with Jupyter",
      "votes": null
    },
    {
      "id": "418907",
      "postDate": "11/10/2018 21:05:47",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "425652",
      "postDate": "11/21/2018 22:51:29",
      "content": "<p>Awesome!</p>",
      "rawMarkdown": "Awesome!",
      "votes": null
    },
    {
      "id": "427702",
      "postDate": "11/26/2018 01:26:39",
      "content": "<p>Hi,TomomiMoriyama,how did you get such a high score, is it convenient to say something?</p>",
      "rawMarkdown": "Hi,TomomiMoriyama,how did you get such a high score, is it convenient to say something?",
      "votes": null
    },
    {
      "id": "430856",
      "postDate": "12/01/2018 04:07:22",
      "content": "<p>Hi! willer, I'm not using TIF :( yet.</p>\n\n<p>Model:\n　Iafoss's [pretrained ResNet34 with RGBY (0.460 public LB)]\n　　arch = resnet34 changed to resnet50\n　　sz = 256 changed to 512\nDown Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image</p>\n\n<p>Stratification:\n　　Trent's [Multilabel Stratification Python Package]</p>\n\n<p>Data:\n　　512 RGBY png image\nExternal Data:\n　HumanProteinAtras v18[Official pre-trained models and external data thread]\n　　including Uncertain\n→got LB 0.528: My model's True Score\n　　contains similar image with test\n　　found 259 matching image by Tilii's phash [A list of identical and near-identical images]\n→with this leakage , got LB 0.588, so it's a cheat and I feel sorry for that.</p>",
      "rawMarkdown": "Hi! willer, I'm not using TIF :( yet.\n\nModel:\n　Iafoss's [pretrained ResNet34 with RGBY (0.460 public LB)]\n　　arch = resnet34 changed to resnet50\n　　sz = 256 changed to 512\nDown Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image\n\nStratification:\n　　Trent's [Multilabel Stratification Python Package]\n\nData:\n　　512 RGBY png image\nExternal Data:\n　HumanProteinAtras v18[Official pre-trained models and external data thread]\n　　including Uncertain\n→got LB 0.528: My model's True Score\n　　contains similar image with test\n　　found 259 matching image by Tilii's phash [A list of identical and near-identical images]\n→with this leakage , got LB 0.588, so it's a cheat and I feel sorry for that.",
      "votes": null
    },
    {
      "id": "431236",
      "postDate": "12/01/2018 22:57:55",
      "content": "<p>That's pretty much how I got my score but with my own model. I need to work better on my leakage, I only used 126 images to boost things. My model is only 1/10th the size of resnet though.</p>",
      "rawMarkdown": "That's pretty much how I got my score but with my own model. I need to work better on my leakage, I only used 126 images to boost things. My model is only 1/10th the size of resnet though.",
      "votes": null
    },
    {
      "id": "431259",
      "postDate": "12/02/2018 00:32:43",
      "content": "<p>Hi, @Brian\nI'm almost new to kaggle competition,\nand I read carefully all the discussion thread.\nI just followed and tried techniques on the commet(most of it is yours)...\nI'm scared to see my score ranked 1st, and thought I did something wrong or break the rules...\nIs it ok to use similarity data?</p>",
      "rawMarkdown": "Hi, @Brian\nI'm almost new to kaggle competition,\nand I read carefully all the discussion thread.\nI just followed and tried techniques on the commet(most of it is yours)...\nI'm scared to see my score ranked 1st, and thought I did something wrong or break the rules...\nIs it ok to use similarity data?",
      "votes": null
    },
    {
      "id": "431268",
      "postDate": "12/02/2018 00:51:40",
      "content": "<p>This is only my second contest, from what I understand as long as you list it in the thread and the license allows you to use it, then it is ok. You found more images in the HPA than I did, I am currently downloading them with yellow from the list you posted on the other thread.</p>",
      "rawMarkdown": "This is only my second contest, from what I understand as long as you list it in the thread and the license allows you to use it, then it is ok. You found more images in the HPA than I did, I am currently downloading them with yellow from the list you posted on the other thread.",
      "votes": null
    },
    {
      "id": "431297",
      "postDate": "12/02/2018 02:38:31",
      "content": "<p>Thank you @Brian\nI feel safe to hear its within the rule. Thank you.</p>",
      "rawMarkdown": "Thank you @Brian\nI feel safe to hear its within the rule. Thank you.",
      "votes": null
    },
    {
      "id": "431314",
      "postDate": "12/02/2018 03:22:35",
      "content": "<p>Phew, I thought I was doing something really wrong.... Almost gave up on this competition few weeks back. Thank you guys very much for the information!</p>",
      "rawMarkdown": "Phew, I thought I was doing something really wrong.... Almost gave up on this competition few weeks back. Thank you guys very much for the information!",
      "votes": null
    },
    {
      "id": "431490",
      "postDate": "12/02/2018 11:36:39",
      "content": "<p>@TomomiMoriyama: the rules are related to the use of external data, they do not address directly the leakage topic\". What is not yet clear to me is why organizers in this competition do not enforce (yet) people using the HPA external data to make it available to everyone. Until now, I though this was the rule: you can use external data as long as it is easily accessible to everybody. Easily means: there is a file with structured data, weights, whatever (no scraping, parsing involved).\nWhen I tried in a previous competition to make use of external data by scraping public webpages, when I wrote my approach in the \"external data thread\" the organizers replied that it was allowed only if I provided the final .csv to everyone through a public Kaggle dataset and so I did. This is the example I am referring to:\n<a href=\"https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia\">https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia</a></p>\n\n<p>Hope it helps</p>",
      "rawMarkdown": "TomomiMoriyama: the rules are related to the use of external data, they do not address directly the leakage topic\". What is not yet clear to me is why organizers in this competition do not enforce (yet) people using the HPA external data to make it available to everyone. Until now, I though this was the rule: you can use external data as long as it is easily accessible to everybody. Easily means: there is a file with structured data, weights, whatever (no scraping, parsing involved).\nWhen I tried in a previous competition to make use of external data by scraping public webpages, when I wrote my approach in the \"external data thread\" the organizers replied that it was allowed only if I provided the final .csv to everyone through a public Kaggle dataset and so I did. This is the example I am referring to:\nhttps://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia\n\nHope it helps",
      "votes": null
    },
    {
      "id": "431531",
      "postDate": "12/02/2018 12:43:34",
      "content": "<p>Thank you @Chase the Trane .\n I'll upload how I scraped public webpage and processed data on  \"external data thread\" .\nFinal dataset-size is 60GB, so it's difficult reproduce on kaggle kernels. </p>",
      "rawMarkdown": "Thank you @Chase the Trane .\n I'll upload how I scraped public webpage and processed data on  \"external data thread\" .\nFinal dataset-size is 60GB, so it's difficult reproduce on kaggle kernels.",
      "votes": null
    },
    {
      "id": "431560",
      "postDate": "12/02/2018 13:47:21",
      "content": "<p>Maybe you can contact directly the organizers about that. They have the final word.\nGood luck and thanks for your sharing!</p>",
      "rawMarkdown": "Maybe you can contact directly the organizers about that. They have the final word.\nGood luck and thanks for your sharing!",
      "votes": null
    },
    {
      "id": "431936",
      "postDate": "12/03/2018 05:54:14",
      "content": "<p>Here is a faster way in ruby. This combines the png/jpg images into 512x512. I use this for converting the separate files consistently. Contest data is PNG grayscale, but HPA is JPG RGB with the colors you select in the URL. This script will take both formats and output singular RGBA png files with the same intensity.</p>\n\n<p><a href=\"https://pastebin.com/v10a1Ckq\">https://pastebin.com/v10a1Ckq</a></p>",
      "rawMarkdown": "Here is a faster way in ruby. This combines the png/jpg images into 512x512. I use this for converting the separate files consistently. Contest data is PNG grayscale, but HPA is JPG RGB with the colors you select in the URL. This script will take both formats and output singular RGBA png files with the same intensity.\n\nhttps://pastebin.com/v10a1Ckq",
      "votes": null
    },
    {
      "id": "432441",
      "postDate": "12/03/2018 21:52:16",
      "content": "<p>Thanks for sharing your external data and explaining what you have done. I think it's not cheating if everybody knows about it. The similarity is only 2.5%, so my point of view  is that it's not that significant. I have only a third of the data you shared (the one that Brain shared) and I thought that everybody had it. Oh my, I believed that those 0.58 were training with 2048x2048 images, lol. Let's see how the competition evolves from now on ;). </p>",
      "rawMarkdown": "Thanks for sharing your external data and explaining what you have done. I think it's not cheating if everybody knows about it. The similarity is only 2.5%, so my point of view  is that it's not that significant. I have only a third of the data you shared (the one that Brain shared) and I thought that everybody had it. Oh my, I believed that those 0.58 were training with 2048x2048 images, lol. Let's see how the competition evolves from now on ;).",
      "votes": null
    },
    {
      "id": "432445",
      "postDate": "12/03/2018 22:04:41",
      "content": "<p>I've gone through again with the yellow information included and better image processing. My new list is up to 216.  We will see how it goes when I can do some more submissions.</p>",
      "rawMarkdown": "I've gone through again with the yellow information included and better image processing. My new list is up to 216.  We will see how it goes when I can do some more submissions.",
      "votes": null
    },
    {
      "id": "439047",
      "postDate": "12/14/2018 16:37:20",
      "content": "<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a>,</p>\n\n<p>When you downsample, are you trying to remove single label image ? Or to keep them ? What is the idea behind this ? Is it better for your pretrained model to get images with only 1 label ? Or as many as possible ? Does someone has any explanation on this ? </p>\n\n<p>Thanks,</p>\n\n<p>Down Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image</p>",
      "rawMarkdown": "tomomimoriyama,\n\nWhen you downsample, are you trying to remove single label image ? Or to keep them ? What is the idea behind this ? Is it better for your pretrained model to get images with only 1 label ? Or as many as possible ? Does someone has any explanation on this ? \n\nThanks,\n\nDown Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image",
      "votes": null
    },
    {
      "id": "439234",
      "postDate": "12/15/2018 00:50:11",
      "content": "<p>Hi, <a href=\"/areveillon\">@areveillon</a></p>\n\n<p>I just wanted to <strong>cope with class-imbalance</strong>.\nand didn't know proper way at that time.</p>\n\n<p>There are more <strong>sophisticated way</strong>!:\n- class-weight:<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a>\n- oversampling:<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374</a></p>\n\n<p>I faithfuly followed @Brian 's ideas on the discussion \nand tried the idea one by one, integrating it into Iafoss's kernel. \nHe once mentioned about <strong>dawnsampling</strong>,\nand my implementation ended up like this:\nI reduced only one-class-labeld sample, \nbecuse it is more easy to implement.\nThere is no theoretical foundation. </p>\n\n<p>#LB 0.502 ;) without-downsampling use all-train-HPAv18\n#saji = [1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1]\n#LB 0.528 :)\n#saji = [0.6,0.9,0.7,0.8,0.7,0.8,0.9,0.7,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.7,1,0.6,1,0.6,1,1]\n#LB 0.498 :(\n#saji = [0.2,0.9,0.5,0.6,0.5,0.6,0.9,0.5,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.5,1,0.5,1,0.5,1,1]\n#LB 0.047 :(\n#saji = [0.1,0.7,0.4,0.5,0.4,0.5,0.7,0.4,1,1,1,0.7,1,1,0.8,1,1,1,1,1,1,0.4,1,0.3,1,0.2,1,1]</p>\n\n<p>def getTrainDataset_RandomSample(data):</p>\n\n<pre><code>paths = []\nlabels = []\n\nfor name, lbl in zip(data['Id'], data['Target'].str.split(' ')):\n    y = np.zeros(28)\n    for key in lbl:\n        y[int(key)] = 1\n    if len(lbl) == 1:\n        s = np.random.random_sample()\n        if s &amp;gt; saji[int(lbl[0])]:\n            continue\n    paths.append(name)\n    labels.append(y)\n\nreturn np.array(paths), np.array(labels)\n</code></pre>\n\n<p>Your idea is great!\nI'll try and feed-back the result \npretrain only one-class (there is no class 15 sample though)\nfine-tune with multi-labeled sample.\nI think no one has mentioned about <strong>Curriculum Learning</strong>.</p>\n\n<p>12/17 Update the <strong>result</strong>:\nTrained on Only 1 label image: LB 0.494\nthen fine-tune with all data: LB 0.025 ... :(</p>",
      "rawMarkdown": "Hi, @areveillon\n\nI just wanted to **cope with class-imbalance**.\nand didn't know proper way at that time.\n\nThere are more **sophisticated way**!:\n- class-weight:https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\n- oversampling:https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374\n\nI faithfuly followed @Brian 's ideas on the discussion \nand tried the idea one by one, integrating it into Iafoss's kernel. \nHe once mentioned about **dawnsampling**,\nand my implementation ended up like this:\nI reduced only one-class-labeld sample, \nbecuse it is more easy to implement.\nThere is no theoretical foundation. \n\n\\#LB 0.502 ;) without-downsampling use all-train-HPAv18\n\\#saji = [1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1]\n\\#LB 0.528 :)\n\\#saji = [0.6,0.9,0.7,0.8,0.7,0.8,0.9,0.7,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.7,1,0.6,1,0.6,1,1]\n\\#LB 0.498 :(\n\\#saji = [0.2,0.9,0.5,0.6,0.5,0.6,0.9,0.5,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.5,1,0.5,1,0.5,1,1]\n\\#LB 0.047 :(\n\\#saji = [0.1,0.7,0.4,0.5,0.4,0.5,0.7,0.4,1,1,1,0.7,1,1,0.8,1,1,1,1,1,1,0.4,1,0.3,1,0.2,1,1]\n\ndef getTrainDataset_RandomSample(data):\n    \n    paths = []\n    labels = []\n    \n    for name, lbl in zip(data['Id'], data['Target'].str.split(' ')):\n        y = np.zeros(28)\n        for key in lbl:\n            y[int(key)] = 1\n        if len(lbl) == 1:\n            s = np.random.random_sample()\n            if s &gt; saji[int(lbl[0])]:\n                continue\n        paths.append(name)\n        labels.append(y)\n\n    return np.array(paths), np.array(labels)\n\nYour idea is great!\nI'll try and feed-back the result \npretrain only one-class (there is no class 15 sample though)\nfine-tune with multi-labeled sample.\nI think no one has mentioned about **Curriculum Learning**.\n\n12/17 Update the **result**:\nTrained on Only 1 label image: LB 0.494\nthen fine-tune with all data: LB 0.025 ... :(",
      "votes": null
    },
    {
      "id": "439249",
      "postDate": "12/15/2018 02:19:38",
      "content": "<p>Thanks a lot ! I was just wondering if it has an impact to keep only 1 label image versus \"many labels\" images... Thanks for taking the time to share your experiments !</p>",
      "rawMarkdown": "Thanks a lot ! I was just wondering if it has an impact to keep only 1 label image versus \"many labels\" images... Thanks for taking the time to share your experiments !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 413471,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "11/01/2018 02:21:04",
      "content": "<p>Actually, the images can be directly read from archive using zipfile library, so no need to extract them first. As an example, the following function, used in another competition, reads images from archive 1 by 1, resizes them, and writes into another archive:</p>\n\n<pre><code>def resize(filename, sz):\n    with zipfile.ZipFile(os.path.join(PATH, f'{filename}.zip'), 'r') as archive, \\\n      zipfile.ZipFile(os.path.join(PATH, filename + str(sz) + '.zip'), 'w') as archive_out:\n        for name in archive.namelist():\n            img = Image.open(io.BytesIO(archive.read(name)))\n            output = io.BytesIO()\n            img.resize((sz,sz)).save(output, format='png')\n            archive_out.writestr(name, output.getvalue())\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 413523,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/01/2018 04:22:13",
          "content": "<p>I might be able to do something like this. I'm using Keras currently, it seemed easier to combine them first externally. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 416677,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "11/07/2018 04:28:28",
          "content": "<p>Full size TIF data is 7z compressed...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 425652,
          "author_name": "felipekitamura",
          "author_url": "",
          "post_date": "11/21/2018 22:51:29",
          "content": "<p>Awesome!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 427702,
          "author_name": "willor",
          "author_url": "",
          "post_date": "11/26/2018 01:26:39",
          "content": "<p>Hi,TomomiMoriyama,how did you get such a high score, is it convenient to say something?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 430856,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/01/2018 04:07:22",
          "content": "<p>Hi! willer, I'm not using TIF :( yet.</p>\n\n<p>Model:\n　Iafoss's [pretrained ResNet34 with RGBY (0.460 public LB)]\n　　arch = resnet34 changed to resnet50\n　　sz = 256 changed to 512\nDown Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image</p>\n\n<p>Stratification:\n　　Trent's [Multilabel Stratification Python Package]</p>\n\n<p>Data:\n　　512 RGBY png image\nExternal Data:\n　HumanProteinAtras v18[Official pre-trained models and external data thread]\n　　including Uncertain\n→got LB 0.528: My model's True Score\n　　contains similar image with test\n　　found 259 matching image by Tilii's phash [A list of identical and near-identical images]\n→with this leakage , got LB 0.588, so it's a cheat and I feel sorry for that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431236,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/01/2018 22:57:55",
          "content": "<p>That's pretty much how I got my score but with my own model. I need to work better on my leakage, I only used 126 images to boost things. My model is only 1/10th the size of resnet though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431259,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/02/2018 00:32:43",
          "content": "<p>Hi, @Brian\nI'm almost new to kaggle competition,\nand I read carefully all the discussion thread.\nI just followed and tried techniques on the commet(most of it is yours)...\nI'm scared to see my score ranked 1st, and thought I did something wrong or break the rules...\nIs it ok to use similarity data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431268,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/02/2018 00:51:40",
          "content": "<p>This is only my second contest, from what I understand as long as you list it in the thread and the license allows you to use it, then it is ok. You found more images in the HPA than I did, I am currently downloading them with yellow from the list you posted on the other thread.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431297,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/02/2018 02:38:31",
          "content": "<p>Thank you @Brian\nI feel safe to hear its within the rule. Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431314,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "12/02/2018 03:22:35",
          "content": "<p>Phew, I thought I was doing something really wrong.... Almost gave up on this competition few weeks back. Thank you guys very much for the information!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431490,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "12/02/2018 11:36:39",
          "content": "<p>@TomomiMoriyama: the rules are related to the use of external data, they do not address directly the leakage topic\". What is not yet clear to me is why organizers in this competition do not enforce (yet) people using the HPA external data to make it available to everyone. Until now, I though this was the rule: you can use external data as long as it is easily accessible to everybody. Easily means: there is a file with structured data, weights, whatever (no scraping, parsing involved).\nWhen I tried in a previous competition to make use of external data by scraping public webpages, when I wrote my approach in the \"external data thread\" the organizers replied that it was allowed only if I provided the final .csv to everyone through a public Kaggle dataset and so I did. This is the example I am referring to:\n<a href=\"https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia\">https://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia</a></p>\n\n<p>Hope it helps</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431531,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/02/2018 12:43:34",
          "content": "<p>Thank you @Chase the Trane .\n I'll upload how I scraped public webpage and processed data on  \"external data thread\" .\nFinal dataset-size is 60GB, so it's difficult reproduce on kaggle kernels. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431560,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "12/02/2018 13:47:21",
          "content": "<p>Maybe you can contact directly the organizers about that. They have the final word.\nGood luck and thanks for your sharing!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432441,
          "author_name": "arnaurm",
          "author_url": "",
          "post_date": "12/03/2018 21:52:16",
          "content": "<p>Thanks for sharing your external data and explaining what you have done. I think it's not cheating if everybody knows about it. The similarity is only 2.5%, so my point of view  is that it's not that significant. I have only a third of the data you shared (the one that Brain shared) and I thought that everybody had it. Oh my, I believed that those 0.58 were training with 2048x2048 images, lol. Let's see how the competition evolves from now on ;). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432445,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/03/2018 22:04:41",
          "content": "<p>I've gone through again with the yellow information included and better image processing. My new list is up to 216.  We will see how it goes when I can do some more submissions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439047,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "12/14/2018 16:37:20",
          "content": "<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a>,</p>\n\n<p>When you downsample, are you trying to remove single label image ? Or to keep them ? What is the idea behind this ? Is it better for your pretrained model to get images with only 1 label ? Or as many as possible ? Does someone has any explanation on this ? </p>\n\n<p>Thanks,</p>\n\n<p>Down Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439234,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/15/2018 00:50:11",
          "content": "<p>Hi, <a href=\"/areveillon\">@areveillon</a></p>\n\n<p>I just wanted to <strong>cope with class-imbalance</strong>.\nand didn't know proper way at that time.</p>\n\n<p>There are more <strong>sophisticated way</strong>!:\n- class-weight:<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a>\n- oversampling:<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374</a></p>\n\n<p>I faithfuly followed @Brian 's ideas on the discussion \nand tried the idea one by one, integrating it into Iafoss's kernel. \nHe once mentioned about <strong>dawnsampling</strong>,\nand my implementation ended up like this:\nI reduced only one-class-labeld sample, \nbecuse it is more easy to implement.\nThere is no theoretical foundation. </p>\n\n<p>#LB 0.502 ;) without-downsampling use all-train-HPAv18\n#saji = [1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1]\n#LB 0.528 :)\n#saji = [0.6,0.9,0.7,0.8,0.7,0.8,0.9,0.7,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.7,1,0.6,1,0.6,1,1]\n#LB 0.498 :(\n#saji = [0.2,0.9,0.5,0.6,0.5,0.6,0.9,0.5,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.5,1,0.5,1,0.5,1,1]\n#LB 0.047 :(\n#saji = [0.1,0.7,0.4,0.5,0.4,0.5,0.7,0.4,1,1,1,0.7,1,1,0.8,1,1,1,1,1,1,0.4,1,0.3,1,0.2,1,1]</p>\n\n<p>def getTrainDataset_RandomSample(data):</p>\n\n<pre><code>paths = []\nlabels = []\n\nfor name, lbl in zip(data['Id'], data['Target'].str.split(' ')):\n    y = np.zeros(28)\n    for key in lbl:\n        y[int(key)] = 1\n    if len(lbl) == 1:\n        s = np.random.random_sample()\n        if s &amp;gt; saji[int(lbl[0])]:\n            continue\n    paths.append(name)\n    labels.append(y)\n\nreturn np.array(paths), np.array(labels)\n</code></pre>\n\n<p>Your idea is great!\nI'll try and feed-back the result \npretrain only one-class (there is no class 15 sample though)\nfine-tune with multi-labeled sample.\nI think no one has mentioned about <strong>Curriculum Learning</strong>.</p>\n\n<p>12/17 Update the <strong>result</strong>:\nTrained on Only 1 label image: LB 0.494\nthen fine-tune with all data: LB 0.025 ... :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439249,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "12/15/2018 02:19:38",
          "content": "<p>Thanks a lot ! I was just wondering if it has an impact to keep only 1 label image versus \"many labels\" images... Thanks for taking the time to share your experiments !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 413515,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "11/01/2018 04:06:34",
      "content": "<p>Assuming that you have ImageMagic's <code>convert</code> installed, a single line will do the same trick on Linux. Just go to the appropriate directory and paste:</p>\n\n<p><code>find . -name \"*_blue.tif\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.tif//g' | xargs -i convert \"{}\"_red.tif \"{}\"_green.tif \"{}\"_blue.tif -set colorspace RGB -combine \"{}\"_rgb.tif</code></p>\n\n<p>Obviously, the same can be done with .png files:</p>\n\n<p><code>find . -name \"*_blue.png\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.png//g' | xargs -i convert \"{}\"_red.png \"{}\"_green.png \"{}\"_blue.png -set colorspace RGB -combine \"{}\"_rgb.png</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 413521,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/01/2018 04:18:41",
          "content": "<p>I ended up switching away from linux to my gaming computer because of the graphics card. This probably will work under git bash or cygwin too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 413938,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/01/2018 20:28:01",
      "content": "<p>Overnight it finished and it looks like it has worked well.</p>\n\n<ul>\n<li>Before: 618GB, RGBY in separate tif</li>\n<li>After: 155GB, RGB png files. Y discarded.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414004,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/02/2018 00:06:22",
      "content": "<p>Using 44GB memory: 3108 validation images, 1024x1024x3 float32(because keras is broke at float16). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414070,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/02/2018 03:37:23",
      "content": "<p>Limit reached, I don't think my drive can go any faster</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417916,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "11/09/2018 01:19:29",
      "content": "<p>I did it in a parallelized fashion:</p>\n\n<pre><code>def save_to_dir(i):\n    if(i%1000==0): print(\"loaded sample {}\".format(i))\n    imid=df['Id'][i]\n    r = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_red.tif\"))\n    g = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_green.tif\"))\n    b = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_blue.tif\"))\n    y = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_yellow.tif\"))\n\n    image = np.dstack((r,g,b,y))\n    np.save(ROOT+\"train_full_size_np/\"+imid,image)\nnum_cores = 8\nParallel(n_jobs=num_cores, prefer=\"threads\")(delayed(save_to_dir)(i) for i in range(len(df)))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 417990,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/09/2018 04:47:52",
          "content": "<p>I tried to do something using the multiprocessing library but it didn't play along with Jupyter. I ended up just letting it run. For now I've just used mogrify and convert to make a few sets at different resolutions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418750,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "11/10/2018 14:56:39",
          "content": "<p>Try concurrent.futures import ThreadPoolExecutor\nIt works fine with Jupyter</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418907,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/10/2018 21:05:47",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 431936,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "12/03/2018 05:54:14",
      "content": "<p>Here is a faster way in ruby. This combines the png/jpg images into 512x512. I use this for converting the separate files consistently. Contest data is PNG grayscale, but HPA is JPG RGB with the colors you select in the URL. This script will take both formats and output singular RGBA png files with the same intensity.</p>\n\n<p><a href=\"https://pastebin.com/v10a1Ckq\">https://pastebin.com/v10a1Ckq</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "413432": "Starting to explore the tif files, they are quite large. Uncompressed the training set is over 600GB. To shrink this down I am converting them to RGB files stored as png. Here is how I am doing it. Locally it looks like this is going to take about 15 hours, so far it looks like it will decrease the size significantly. Especially so because I am using 3 channels instead of 4.\n\n<pre># use convert to merge images to rgb\nfrom tqdm import tqdm_notebook\nfrom pathlib import Path\nimport os\n\ninput_path = Path(\"../input/human_protein_atlas\")\ntrain_csv = input_path / \"train.csv\"\n\ninput_path_large = Path(\"J:/hpa_images\")\ntrain_dir = input_path_large / \"train_full_size\"\ntrain_combined_dir = input_path_fast / \"train_combined_rgb\"\n\ndata = pd.read_csv(str(train_csv))\nids = data['Id'].tolist()\n\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(train_dir / img_id)\n    red_img = (img_path + \"_red.tif\")\n    yellow_img = (img_path + \"_yellow.tif\")\n    blue_img = (img_path + \"_blue.tif\")\n    green_img = (img_path + \"_green.tif\")\n    \n    out_img = str(train_combined_dir / img_id) + \"_rgb.png\"\n    \n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s  -set colorspace RGB -combine %s\" % (red_img,green_img,blue_img,out_img) \n    #print(cmd)\n    os.system(cmd)\n</pre>",
    "413471": "Actually, the images can be directly read from archive using zipfile library, so no need to extract them first. As an example, the following function, used in another competition, reads images from archive 1 by 1, resizes them, and writes into another archive:\n\n    def resize(filename, sz):\n        with zipfile.ZipFile(os.path.join(PATH, f'{filename}.zip'), 'r') as archive, \\\n          zipfile.ZipFile(os.path.join(PATH, filename + str(sz) + '.zip'), 'w') as archive_out:\n            for name in archive.namelist():\n                img = Image.open(io.BytesIO(archive.read(name)))\n                output = io.BytesIO()\n                img.resize((sz,sz)).save(output, format='png')\n                archive_out.writestr(name, output.getvalue())",
    "413515": "Assuming that you have ImageMagic's `convert` installed, a single line will do the same trick on Linux. Just go to the appropriate directory and paste:\n\n`find . -name \"*_blue.tif\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.tif//g' | xargs -i convert \"{}\"_red.tif \"{}\"_green.tif \"{}\"_blue.tif -set colorspace RGB -combine \"{}\"_rgb.tif`\n\nObviously, the same can be done with .png files:\n\n`find . -name \"*_blue.png\" | sort | perl -pi -e 's/\\.\\///g' | perl -pi -e 's/_blue\\.png//g' | xargs -i convert \"{}\"_red.png \"{}\"_green.png \"{}\"_blue.png -set colorspace RGB -combine \"{}\"_rgb.png`",
    "413521": "I ended up switching away from linux to my gaming computer because of the graphics card. This probably will work under git bash or cygwin too.",
    "413523": "I might be able to do something like this. I'm using Keras currently, it seemed easier to combine them first externally.",
    "413938": "Overnight it finished and it looks like it has worked well.\n\n* Before: 618GB, RGBY in separate tif\n* After: 155GB, RGB png files. Y discarded.",
    "414004": "Using 44GB memory: 3108 validation images, 1024x1024x3 float32(because keras is broke at float16).",
    "414070": "Limit reached, I don't think my drive can go any faster",
    "416677": "Full size TIF data is 7z compressed...",
    "417916": "I did it in a parallelized fashion:\n\n    def save_to_dir(i):\n        if(i%1000==0): print(\"loaded sample {}\".format(i))\n        imid=df['Id'][i]\n        r = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_red.tif\"))\n        g = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_green.tif\"))\n        b = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_blue.tif\"))\n        y = np.array(Image.open(ROOT+'train_full_size/'+imid+\"_yellow.tif\"))\n\n        image = np.dstack((r,g,b,y))\n        np.save(ROOT+\"train_full_size_np/\"+imid,image)\n    num_cores = 8\n    Parallel(n_jobs=num_cores, prefer=\"threads\")(delayed(save_to_dir)(i) for i in range(len(df)))",
    "417990": "I tried to do something using the multiprocessing library but it didn't play along with Jupyter. I ended up just letting it run. For now I've just used mogrify and convert to make a few sets at different resolutions.",
    "418750": "Try concurrent.futures import ThreadPoolExecutor\nIt works fine with Jupyter",
    "418907": "Thanks!",
    "425652": "Awesome!",
    "427702": "Hi,TomomiMoriyama,how did you get such a high score, is it convenient to say something?",
    "430856": "Hi! willer, I'm not using TIF :( yet.\n\nModel:\n　Iafoss's [pretrained ResNet34 with RGBY (0.460 public LB)]\n　　arch = resnet34 changed to resnet50\n　　sz = 256 changed to 512\nDown Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image\n\nStratification:\n　　Trent's [Multilabel Stratification Python Package]\n\nData:\n　　512 RGBY png image\nExternal Data:\n　HumanProteinAtras v18[Official pre-trained models and external data thread]\n　　including Uncertain\n→got LB 0.528: My model's True Score\n　　contains similar image with test\n　　found 259 matching image by Tilii's phash [A list of identical and near-identical images]\n→with this leakage , got LB 0.588, so it's a cheat and I feel sorry for that.",
    "431236": "That's pretty much how I got my score but with my own model. I need to work better on my leakage, I only used 126 images to boost things. My model is only 1/10th the size of resnet though.",
    "431259": "Hi, @Brian\nI'm almost new to kaggle competition,\nand I read carefully all the discussion thread.\nI just followed and tried techniques on the commet(most of it is yours)...\nI'm scared to see my score ranked 1st, and thought I did something wrong or break the rules...\nIs it ok to use similarity data?",
    "431268": "This is only my second contest, from what I understand as long as you list it in the thread and the license allows you to use it, then it is ok. You found more images in the HPA than I did, I am currently downloading them with yellow from the list you posted on the other thread.",
    "431297": "Thank you @Brian\nI feel safe to hear its within the rule. Thank you.",
    "431314": "Phew, I thought I was doing something really wrong.... Almost gave up on this competition few weeks back. Thank you guys very much for the information!",
    "431490": "TomomiMoriyama: the rules are related to the use of external data, they do not address directly the leakage topic\". What is not yet clear to me is why organizers in this competition do not enforce (yet) people using the HPA external data to make it available to everyone. Until now, I though this was the rule: you can use external data as long as it is easily accessible to everybody. Easily means: there is a file with structured data, weights, whatever (no scraping, parsing involved).\nWhen I tried in a previous competition to make use of external data by scraping public webpages, when I wrote my approach in the \"external data thread\" the organizers replied that it was allowed only if I provided the final .csv to everyone through a public Kaggle dataset and so I did. This is the example I am referring to:\nhttps://www.kaggle.com/stecasasso/russian-city-population-from-wikipedia\n\nHope it helps",
    "431531": "Thank you @Chase the Trane .\n I'll upload how I scraped public webpage and processed data on  \"external data thread\" .\nFinal dataset-size is 60GB, so it's difficult reproduce on kaggle kernels.",
    "431560": "Maybe you can contact directly the organizers about that. They have the final word.\nGood luck and thanks for your sharing!",
    "431936": "Here is a faster way in ruby. This combines the png/jpg images into 512x512. I use this for converting the separate files consistently. Contest data is PNG grayscale, but HPA is JPG RGB with the colors you select in the URL. This script will take both formats and output singular RGBA png files with the same intensity.\n\nhttps://pastebin.com/v10a1Ckq",
    "432441": "Thanks for sharing your external data and explaining what you have done. I think it's not cheating if everybody knows about it. The similarity is only 2.5%, so my point of view  is that it's not that significant. I have only a third of the data you shared (the one that Brain shared) and I thought that everybody had it. Oh my, I believed that those 0.58 were training with 2048x2048 images, lol. Let's see how the competition evolves from now on ;).",
    "432445": "I've gone through again with the yellow information included and better image processing. My new list is up to 216.  We will see how it goes when I can do some more submissions.",
    "439047": "tomomimoriyama,\n\nWhen you downsample, are you trying to remove single label image ? Or to keep them ? What is the idea behind this ? Is it better for your pretrained model to get images with only 1 label ? Or as many as possible ? Does someone has any explanation on this ? \n\nThanks,\n\nDown Sample:\n　　randomly down sample up to 40% \n　　only plentiful only-one-class-labeled image",
    "439234": "Hi, @areveillon\n\nI just wanted to **cope with class-imbalance**.\nand didn't know proper way at that time.\n\nThere are more **sophisticated way**!:\n- class-weight:https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\n- oversampling:https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374\n\nI faithfuly followed @Brian 's ideas on the discussion \nand tried the idea one by one, integrating it into Iafoss's kernel. \nHe once mentioned about **dawnsampling**,\nand my implementation ended up like this:\nI reduced only one-class-labeld sample, \nbecuse it is more easy to implement.\nThere is no theoretical foundation. \n\n\\#LB 0.502 ;) without-downsampling use all-train-HPAv18\n\\#saji = [1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1]\n\\#LB 0.528 :)\n\\#saji = [0.6,0.9,0.7,0.8,0.7,0.8,0.9,0.7,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.7,1,0.6,1,0.6,1,1]\n\\#LB 0.498 :(\n\\#saji = [0.2,0.9,0.5,0.6,0.5,0.6,0.9,0.5,1,1,1,0.9,1,1,0.9,1,1,1,1,1,1,0.5,1,0.5,1,0.5,1,1]\n\\#LB 0.047 :(\n\\#saji = [0.1,0.7,0.4,0.5,0.4,0.5,0.7,0.4,1,1,1,0.7,1,1,0.8,1,1,1,1,1,1,0.4,1,0.3,1,0.2,1,1]\n\ndef getTrainDataset_RandomSample(data):\n    \n    paths = []\n    labels = []\n    \n    for name, lbl in zip(data['Id'], data['Target'].str.split(' ')):\n        y = np.zeros(28)\n        for key in lbl:\n            y[int(key)] = 1\n        if len(lbl) == 1:\n            s = np.random.random_sample()\n            if s &gt; saji[int(lbl[0])]:\n                continue\n        paths.append(name)\n        labels.append(y)\n\n    return np.array(paths), np.array(labels)\n\nYour idea is great!\nI'll try and feed-back the result \npretrain only one-class (there is no class 15 sample though)\nfine-tune with multi-labeled sample.\nI think no one has mentioned about **Curriculum Learning**.\n\n12/17 Update the **result**:\nTrained on Only 1 label image: LB 0.494\nthen fine-tune with all data: LB 0.025 ... :(",
    "439249": "Thanks a lot ! I was just wondering if it has an impact to keep only 1 label image versus \"many labels\" images... Thanks for taking the time to share your experiments !"
  },
  "source": "meta"
}