{
  "id": 67920,
  "title": "Understanding",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/67920",
  "author_name": "",
  "post_date": "2018-10-07T09:41:21.879457600Z",
  "votes": 8,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Can anyone explain to me what does the following sentence mean : \"The green filter should hence be used to predict the label, and the other filters are used as references\". What do they mean by \"references\" ? Am I understanding correctly that these other filters should not be used for prediction ?  Why are they even given then ? </p>",
  "messages": [
    {
      "id": "399988",
      "postDate": "10/07/2018 09:41:21",
      "content": "<p>Can anyone explain to me what does the following sentence mean : \"The green filter should hence be used to predict the label, and the other filters are used as references\". What do they mean by \"references\" ? Am I understanding correctly that these other filters should not be used for prediction ?  Why are they even given then ? </p>",
      "rawMarkdown": "Can anyone explain to me what does the following sentence mean : \"The green filter should hence be used to predict the label, and the other filters are used as references\". What do they mean by \"references\" ? Am I understanding correctly that these other filters should not be used for prediction ?  Why are they even given then ?",
      "votes": null
    },
    {
      "id": "400207",
      "postDate": "10/07/2018 21:01:39",
      "content": "<p>I just merged the 4 images into a RGBA image, and the predicitons are better. Just my 2 cents.</p>",
      "rawMarkdown": "I just merged the 4 images into a RGBA image, and the predicitons are better. Just my 2 cents.",
      "votes": null
    },
    {
      "id": "400234",
      "postDate": "10/08/2018 00:19:57",
      "content": "<p>Did the same using imagemagick:</p>\n\n<pre>ids = submission_df['Id'].tolist()\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(test_dir / img_id)\n    red_img = (img_path + \"_red.png\")\n    yellow_img = (img_path + \"_yellow.png\")\n    blue_img = (img_path + \"_blue.png\")\n    green_img = (img_path + \"_green.png\")\n\n    out_img = str(test_combined_dir / img_id) + \"_rgba.png\"\n\n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s %s  -set colorspace RGBA -combine %s\" % (red_img,green_img,blue_img,yellow_img,out_img) \n    os.system(cmd)\n</pre>",
      "rawMarkdown": "Did the same using imagemagick:\n\n<pre>ids = submission_df['Id'].tolist()\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(test_dir / img_id)\n    red_img = (img_path + \"_red.png\")\n    yellow_img = (img_path + \"_yellow.png\")\n    blue_img = (img_path + \"_blue.png\")\n    green_img = (img_path + \"_green.png\")\n    \n    out_img = str(test_combined_dir / img_id) + \"_rgba.png\"\n    \n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s %s  -set colorspace RGBA -combine %s\" % (red_img,green_img,blue_img,yellow_img,out_img) \n    os.system(cmd)\n</pre>",
      "votes": null
    },
    {
      "id": "400403",
      "postDate": "10/08/2018 09:07:16",
      "content": "<p>The reference filters will most likely help with the prediction. The green filter shows the labels that are provided in the training set, ie the green filter is the protein of interest. The other reference filters always show the same organelles, blue - nucleus, red - microtubules, yellow - endoplasmic reticulum. To be able to say what organelle the green filter is showing it usually helps to compare the green filter with the reference filters.</p>",
      "rawMarkdown": "The reference filters will most likely help with the prediction. The green filter shows the labels that are provided in the training set, ie the green filter is the protein of interest. The other reference filters always show the same organelles, blue - nucleus, red - microtubules, yellow - endoplasmic reticulum. To be able to say what organelle the green filter is showing it usually helps to compare the green filter with the reference filters.",
      "votes": null
    },
    {
      "id": "400467",
      "postDate": "10/08/2018 11:21:59",
      "content": "<p>I don't get why with a good 0.7 score in F1 (in the val set) when I submit it get crashed to 0.04</p>",
      "rawMarkdown": "I don't get why with a good 0.7 score in F1 (in the val set) when I submit it get crashed to 0.04",
      "votes": null
    },
    {
      "id": "400936",
      "postDate": "10/09/2018 06:56:18",
      "content": "<p>You are probably overfitting. Also, do you account for the class imbalance?</p>",
      "rawMarkdown": "You are probably overfitting. Also, do you account for the class imbalance?",
      "votes": null
    },
    {
      "id": "400975",
      "postDate": "10/09/2018 08:14:58",
      "content": "<p>Yeah, Macro F1 Score is really sensitive to the class balance and \nprobability threshold. In general, there are two 28-dimensional vectors for tuning... and a small discrepancy from the optimal point can lead to fatal degradation of quality on the test.</p>",
      "rawMarkdown": "Yeah, Macro F1 Score is really sensitive to the class balance and \nprobability threshold. In general, there are two 28-dimensional vectors for tuning... and a small discrepancy from the optimal point can lead to fatal degradation of quality on the test.",
      "votes": null
    },
    {
      "id": "401980",
      "postDate": "10/11/2018 01:49:38",
      "content": "<p>Did the dataset become smaller in terms of total filesize? If so I would appreciate the smaller PNG and larger TIFF dataset to be compacted before downloading.</p>\n\n<p>Are the unused color channels identically exactly 0? or is there some color bleeding or mysterious noise?</p>",
      "rawMarkdown": "Did the dataset become smaller in terms of total filesize? If so I would appreciate the smaller PNG and larger TIFF dataset to be compacted before downloading.\n\nAre the unused color channels identically exactly 0? or is there some color bleeding or mysterious noise?",
      "votes": null
    },
    {
      "id": "402248",
      "postDate": "10/11/2018 12:15:11",
      "content": "<p>The \"references\" delineate compartments within the cells. Without this information there is no way you can predict where the target protein (identified by the green filter) is present. For example, when the green and blue channels display overlaps, then you know that particular protein is likely present in the nucleus.</p>",
      "rawMarkdown": "The \"references\" delineate compartments within the cells. Without this information there is no way you can predict where the target protein (identified by the green filter) is present. For example, when the green and blue channels display overlaps, then you know that particular protein is likely present in the nucleus.",
      "votes": null
    },
    {
      "id": "402488",
      "postDate": "10/11/2018 19:39:40",
      "content": "<p>With RGBA there is no unused channels. I put the yellow in the alpha channel.  I didn't do a detailed inspection but they look ok visually.</p>\n\n<p>Size wise they increased minimally. Uncompressed train folder is 13GB for the separate images, 14GB for the combined. I didn't try to do any png optimizations.</p>",
      "rawMarkdown": "With RGBA there is no unused channels. I put the yellow in the alpha channel.  I didn't do a detailed inspection but they look ok visually.\n\nSize wise they increased minimally. Uncompressed train folder is 13GB for the separate images, 14GB for the combined. I didn't try to do any png optimizations.",
      "votes": null
    },
    {
      "id": "402636",
      "postDate": "10/12/2018 03:07:39",
      "content": "<p>Did you resolve this issue? I just have exactly the same(</p>",
      "rawMarkdown": "Did you resolve this issue? I just have exactly the same(",
      "votes": null
    },
    {
      "id": "402817",
      "postDate": "10/12/2018 11:50:13",
      "content": "<p>One question, I observe that the accuracy on the public test data increases if we train on the merged images and submit after merging the test data as well. But during testing on the hidden test data, will the hosts check accuracy on the green images or merged?</p>",
      "rawMarkdown": "One question, I observe that the accuracy on the public test data increases if we train on the merged images and submit after merging the test data as well. But during testing on the hidden test data, will the hosts check accuracy on the green images or merged?",
      "votes": null
    },
    {
      "id": "404129",
      "postDate": "10/15/2018 09:39:19",
      "content": "<p>Indeed! I wonder how well we can score with a simple colocalization analysis, without any \"real\" machine learning.</p>",
      "rawMarkdown": "Indeed! I wonder how well we can score with a simple colocalization analysis, without any \"real\" machine learning.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 400207,
      "author_name": "tcapelle",
      "author_url": "",
      "post_date": "10/07/2018 21:01:39",
      "content": "<p>I just merged the 4 images into a RGBA image, and the predicitons are better. Just my 2 cents.</p>",
      "votes": null,
      "replies": [
        {
          "id": 400234,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/08/2018 00:19:57",
          "content": "<p>Did the same using imagemagick:</p>\n\n<pre>ids = submission_df['Id'].tolist()\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(test_dir / img_id)\n    red_img = (img_path + \"_red.png\")\n    yellow_img = (img_path + \"_yellow.png\")\n    blue_img = (img_path + \"_blue.png\")\n    green_img = (img_path + \"_green.png\")\n\n    out_img = str(test_combined_dir / img_id) + \"_rgba.png\"\n\n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s %s  -set colorspace RGBA -combine %s\" % (red_img,green_img,blue_img,yellow_img,out_img) \n    os.system(cmd)\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 401980,
          "author_name": "ludwigmaes",
          "author_url": "",
          "post_date": "10/11/2018 01:49:38",
          "content": "<p>Did the dataset become smaller in terms of total filesize? If so I would appreciate the smaller PNG and larger TIFF dataset to be compacted before downloading.</p>\n\n<p>Are the unused color channels identically exactly 0? or is there some color bleeding or mysterious noise?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 402488,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/11/2018 19:39:40",
          "content": "<p>With RGBA there is no unused channels. I put the yellow in the alpha channel.  I didn't do a detailed inspection but they look ok visually.</p>\n\n<p>Size wise they increased minimally. Uncompressed train folder is 13GB for the separate images, 14GB for the combined. I didn't try to do any png optimizations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 400403,
      "author_name": "martinhjelmare",
      "author_url": "",
      "post_date": "10/08/2018 09:07:16",
      "content": "<p>The reference filters will most likely help with the prediction. The green filter shows the labels that are provided in the training set, ie the green filter is the protein of interest. The other reference filters always show the same organelles, blue - nucleus, red - microtubules, yellow - endoplasmic reticulum. To be able to say what organelle the green filter is showing it usually helps to compare the green filter with the reference filters.</p>",
      "votes": null,
      "replies": [
        {
          "id": 402817,
          "author_name": "tonmoyj",
          "author_url": "",
          "post_date": "10/12/2018 11:50:13",
          "content": "<p>One question, I observe that the accuracy on the public test data increases if we train on the merged images and submit after merging the test data as well. But during testing on the hidden test data, will the hosts check accuracy on the green images or merged?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 400467,
      "author_name": "tcapelle",
      "author_url": "",
      "post_date": "10/08/2018 11:21:59",
      "content": "<p>I don't get why with a good 0.7 score in F1 (in the val set) when I submit it get crashed to 0.04</p>",
      "votes": null,
      "replies": [
        {
          "id": 400936,
          "author_name": "kostaspapastamos",
          "author_url": "",
          "post_date": "10/09/2018 06:56:18",
          "content": "<p>You are probably overfitting. Also, do you account for the class imbalance?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 400975,
          "author_name": "sggpls",
          "author_url": "",
          "post_date": "10/09/2018 08:14:58",
          "content": "<p>Yeah, Macro F1 Score is really sensitive to the class balance and \nprobability threshold. In general, there are two 28-dimensional vectors for tuning... and a small discrepancy from the optimal point can lead to fatal degradation of quality on the test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 402636,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/12/2018 03:07:39",
          "content": "<p>Did you resolve this issue? I just have exactly the same(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 402248,
      "author_name": "monogenea",
      "author_url": "",
      "post_date": "10/11/2018 12:15:11",
      "content": "<p>The \"references\" delineate compartments within the cells. Without this information there is no way you can predict where the target protein (identified by the green filter) is present. For example, when the green and blue channels display overlaps, then you know that particular protein is likely present in the nucleus.</p>",
      "votes": null,
      "replies": [
        {
          "id": 404129,
          "author_name": "jschnab",
          "author_url": "",
          "post_date": "10/15/2018 09:39:19",
          "content": "<p>Indeed! I wonder how well we can score with a simple colocalization analysis, without any \"real\" machine learning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "399988": "Can anyone explain to me what does the following sentence mean : \"The green filter should hence be used to predict the label, and the other filters are used as references\". What do they mean by \"references\" ? Am I understanding correctly that these other filters should not be used for prediction ?  Why are they even given then ?",
    "400207": "I just merged the 4 images into a RGBA image, and the predicitons are better. Just my 2 cents.",
    "400234": "Did the same using imagemagick:\n\n<pre>ids = submission_df['Id'].tolist()\n\nfor img_id in tqdm_notebook(ids):\n    img_path = str(test_dir / img_id)\n    red_img = (img_path + \"_red.png\")\n    yellow_img = (img_path + \"_yellow.png\")\n    blue_img = (img_path + \"_blue.png\")\n    green_img = (img_path + \"_green.png\")\n    \n    out_img = str(test_combined_dir / img_id) + \"_rgba.png\"\n    \n    cmd= \"\\\"H:\\\\image_magick\\\\convert.exe\\\" %s %s %s %s  -set colorspace RGBA -combine %s\" % (red_img,green_img,blue_img,yellow_img,out_img) \n    os.system(cmd)\n</pre>",
    "400403": "The reference filters will most likely help with the prediction. The green filter shows the labels that are provided in the training set, ie the green filter is the protein of interest. The other reference filters always show the same organelles, blue - nucleus, red - microtubules, yellow - endoplasmic reticulum. To be able to say what organelle the green filter is showing it usually helps to compare the green filter with the reference filters.",
    "400467": "I don't get why with a good 0.7 score in F1 (in the val set) when I submit it get crashed to 0.04",
    "400936": "You are probably overfitting. Also, do you account for the class imbalance?",
    "400975": "Yeah, Macro F1 Score is really sensitive to the class balance and \nprobability threshold. In general, there are two 28-dimensional vectors for tuning... and a small discrepancy from the optimal point can lead to fatal degradation of quality on the test.",
    "401980": "Did the dataset become smaller in terms of total filesize? If so I would appreciate the smaller PNG and larger TIFF dataset to be compacted before downloading.\n\nAre the unused color channels identically exactly 0? or is there some color bleeding or mysterious noise?",
    "402248": "The \"references\" delineate compartments within the cells. Without this information there is no way you can predict where the target protein (identified by the green filter) is present. For example, when the green and blue channels display overlaps, then you know that particular protein is likely present in the nucleus.",
    "402488": "With RGBA there is no unused channels. I put the yellow in the alpha channel.  I didn't do a detailed inspection but they look ok visually.\n\nSize wise they increased minimally. Uncompressed train folder is 13GB for the separate images, 14GB for the combined. I didn't try to do any png optimizations.",
    "402636": "Did you resolve this issue? I just have exactly the same(",
    "402817": "One question, I observe that the accuracy on the public test data increases if we train on the merged images and submit after merging the test data as well. But during testing on the hidden test data, will the hosts check accuracy on the green images or merged?",
    "404129": "Indeed! I wonder how well we can score with a simple colocalization analysis, without any \"real\" machine learning."
  },
  "source": "meta"
}