{
  "id": 70770,
  "title": "Useful Statistical Data for Threshold ... etc",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70770",
  "author_name": "",
  "post_date": "2018-11-07T06:06:14.605670700Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I put these data calculated through the training set so that you can use them. (If you got a different result, please notify me!) Feel free to add some more.</p>\n\n<p>Mean and Variance for normalization:</p>\n\n<pre><code>\"\"\"\nTrain Data:\n    Mean = [0.0804419,  0.05262986, 0.05474701, 0.08270896]\n    Var  = [0.0025557  0.0023054  0.0012995  0.00293925]\n    Var1 = [0.00255578 0.00230547 0.00129955 0.00293934]\n\"\"\"\n\n\"\"\"\nTest Data:\n    Mean = [0.05908022, 0.04532852, 0.04065233, 0.05923426]\n    Var  = [0.00235361 0.0020402  0.00137821 0.00246495]\n    Var1 = [0.00235381 0.00204037 0.00137833 0.00246516]\n\"\"\"\n</code></pre>\n\n<p>Mean of train+test data: [0.07459783,  0.05063238,  0.05089102,  0.07628681]</p>\n\n<p>(Also see: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462</a>)</p>\n\n<p>The label distribution for weight loss:</p>\n\n<pre><code>\"\"\"\n0     12885\n25     8228\n21     3777\n2      3621\n23     2965\n7      2822\n5      2513\n4      1858\n3      1561\n19     1482\n1      1254\n11     1093\n14     1066\n6      1008\n18      902\n22      802\n12      688\n13      537\n16      530\n26      328\n24      322\n17      210\n20      172\n8        53\n9        45\n10       28\n15       21\n27       11\n\"\"\"\n</code></pre>\n\n<p>LB Probing (from @Iafoss <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678</a>)</p>\n\n<pre><code>\"\"\"LB Probing\n0 -&amp;gt; 0.019 -&amp;gt; 0.36239782\n1 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n2 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n3 -&amp;gt; 0.004 -&amp;gt; 0.059322034\n4 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n5 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n6 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n7 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n8 -&amp;gt; 0 -&amp;gt; 0\n9 -&amp;gt; 0 -&amp;gt; 0\n10 -&amp;gt; 0 -&amp;gt; 0\n11 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n12 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n13 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n14 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n15 -&amp;gt; 0 -&amp;gt; 0\n16 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n17 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n18 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n19 -&amp;gt; 0.004 -&amp;gt; 0.059322034\n20 -&amp;gt; 0 -&amp;gt; 0\n21 -&amp;gt; 0.008 -&amp;gt; 0.126126126\n22 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n23 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n24 -&amp;gt; 0 -&amp;gt; 0\n25 -&amp;gt; 0.013 -&amp;gt; 0.222493888\n26 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n27 -&amp;gt; 0 -&amp;gt; 0\n\"\"\"\n</code></pre>\n\n<p>File sizes from @Alexander Liao and <a href=\"/anokas\">@anokas</a> (<a href=\"https://www.kaggle.com/alexanderliao/exploratory-data-analysis/notebook\">https://www.kaggle.com/alexanderliao/exploratory-data-analysis/notebook</a>)</p>\n\n<pre><code>train                         14.0321076GB (124288 files)\ntest                          4.6823579GB (46808 files)\ntrain.csv                     0.0012744GB\nsample_submission.csv         0.0004564GB\n\n    {\n    0: 'Nucleoplasm',\n    1: 'Nuclear membrane',\n    2: 'Nucleoli',\n    3: 'Nucleoli fibrillar center',\n    4: 'Nuclear speckles',\n    5: 'Nuclear bodies',\n    6: 'Endoplasmic reticulum',\n    7: 'Golgi apparatus',\n    8: 'Peroxisomes',\n    9: 'Endosomes',\n    10: 'Lysosomes',\n    11: 'Intermediate filaments',\n    12: 'Actin filaments',\n    13: 'Focal adhesion sites',\n    14: 'Microtubules',\n    15: 'Microtubule ends',\n    16: 'Cytokinetic bridge',\n    17: 'Mitotic spindle',\n    18: 'Microtubule organizing center',\n    19: 'Centrosome',\n    20: 'Lipid droplets',\n    21: 'Plasma membrane',\n    22: 'Cell junctions',\n    23: 'Mitochondria',\n    24: 'Aggresome',\n    25: 'Cytosol',\n    26: 'Cytoplasmic bodies',\n    27: 'Rods &amp;amp; rings'}\n</code></pre>",
  "messages": [
    {
      "id": "416701",
      "postDate": "11/07/2018 06:06:14",
      "content": "<p>I put these data calculated through the training set so that you can use them. (If you got a different result, please notify me!) Feel free to add some more.</p>\n\n<p>Mean and Variance for normalization:</p>\n\n<pre><code>\"\"\"\nTrain Data:\n    Mean = [0.0804419,  0.05262986, 0.05474701, 0.08270896]\n    Var  = [0.0025557  0.0023054  0.0012995  0.00293925]\n    Var1 = [0.00255578 0.00230547 0.00129955 0.00293934]\n\"\"\"\n\n\"\"\"\nTest Data:\n    Mean = [0.05908022, 0.04532852, 0.04065233, 0.05923426]\n    Var  = [0.00235361 0.0020402  0.00137821 0.00246495]\n    Var1 = [0.00235381 0.00204037 0.00137833 0.00246516]\n\"\"\"\n</code></pre>\n\n<p>Mean of train+test data: [0.07459783,  0.05063238,  0.05089102,  0.07628681]</p>\n\n<p>(Also see: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462</a>)</p>\n\n<p>The label distribution for weight loss:</p>\n\n<pre><code>\"\"\"\n0     12885\n25     8228\n21     3777\n2      3621\n23     2965\n7      2822\n5      2513\n4      1858\n3      1561\n19     1482\n1      1254\n11     1093\n14     1066\n6      1008\n18      902\n22      802\n12      688\n13      537\n16      530\n26      328\n24      322\n17      210\n20      172\n8        53\n9        45\n10       28\n15       21\n27       11\n\"\"\"\n</code></pre>\n\n<p>LB Probing (from @Iafoss <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678</a>)</p>\n\n<pre><code>\"\"\"LB Probing\n0 -&amp;gt; 0.019 -&amp;gt; 0.36239782\n1 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n2 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n3 -&amp;gt; 0.004 -&amp;gt; 0.059322034\n4 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n5 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n6 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n7 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n8 -&amp;gt; 0 -&amp;gt; 0\n9 -&amp;gt; 0 -&amp;gt; 0\n10 -&amp;gt; 0 -&amp;gt; 0\n11 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n12 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n13 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n14 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n15 -&amp;gt; 0 -&amp;gt; 0\n16 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n17 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n18 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n19 -&amp;gt; 0.004 -&amp;gt; 0.059322034\n20 -&amp;gt; 0 -&amp;gt; 0\n21 -&amp;gt; 0.008 -&amp;gt; 0.126126126\n22 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n23 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n24 -&amp;gt; 0 -&amp;gt; 0\n25 -&amp;gt; 0.013 -&amp;gt; 0.222493888\n26 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n27 -&amp;gt; 0 -&amp;gt; 0\n\"\"\"\n</code></pre>\n\n<p>File sizes from @Alexander Liao and <a href=\"/anokas\">@anokas</a> (<a href=\"https://www.kaggle.com/alexanderliao/exploratory-data-analysis/notebook\">https://www.kaggle.com/alexanderliao/exploratory-data-analysis/notebook</a>)</p>\n\n<pre><code>train                         14.0321076GB (124288 files)\ntest                          4.6823579GB (46808 files)\ntrain.csv                     0.0012744GB\nsample_submission.csv         0.0004564GB\n\n    {\n    0: 'Nucleoplasm',\n    1: 'Nuclear membrane',\n    2: 'Nucleoli',\n    3: 'Nucleoli fibrillar center',\n    4: 'Nuclear speckles',\n    5: 'Nuclear bodies',\n    6: 'Endoplasmic reticulum',\n    7: 'Golgi apparatus',\n    8: 'Peroxisomes',\n    9: 'Endosomes',\n    10: 'Lysosomes',\n    11: 'Intermediate filaments',\n    12: 'Actin filaments',\n    13: 'Focal adhesion sites',\n    14: 'Microtubules',\n    15: 'Microtubule ends',\n    16: 'Cytokinetic bridge',\n    17: 'Mitotic spindle',\n    18: 'Microtubule organizing center',\n    19: 'Centrosome',\n    20: 'Lipid droplets',\n    21: 'Plasma membrane',\n    22: 'Cell junctions',\n    23: 'Mitochondria',\n    24: 'Aggresome',\n    25: 'Cytosol',\n    26: 'Cytoplasmic bodies',\n    27: 'Rods &amp;amp; rings'}\n</code></pre>",
      "rawMarkdown": "I put these data calculated through the training set so that you can use them. (If you got a different result, please notify me!) Feel free to add some more.\n\nMean and Variance for normalization:\n\n    \"\"\"\n    Train Data:\n        Mean = [0.0804419,  0.05262986, 0.05474701, 0.08270896]\n        Var  = [0.0025557  0.0023054  0.0012995  0.00293925]\n        Var1 = [0.00255578 0.00230547 0.00129955 0.00293934]\n    \"\"\"\n\n    \"\"\"\n    Test Data:\n        Mean = [0.05908022, 0.04532852, 0.04065233, 0.05923426]\n        Var  = [0.00235361 0.0020402  0.00137821 0.00246495]\n        Var1 = [0.00235381 0.00204037 0.00137833 0.00246516]\n    \"\"\"\n\nMean of train+test data: [0.07459783,  0.05063238,  0.05089102,  0.07628681]\n\n\n(Also see: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462)\n\nThe label distribution for weight loss:\n\n    \"\"\"\n    0     12885\n    25     8228\n    21     3777\n    2      3621\n    23     2965\n    7      2822\n    5      2513\n    4      1858\n    3      1561\n    19     1482\n    1      1254\n    11     1093\n    14     1066\n    6      1008\n    18      902\n    22      802\n    12      688\n    13      537\n    16      530\n    26      328\n    24      322\n    17      210\n    20      172\n    8        53\n    9        45\n    10       28\n    15       21\n    27       11\n    \"\"\"\n\nLB Probing (from @Iafoss https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678)\n\n    \"\"\"LB Probing\n    0 -&gt; 0.019 -&gt; 0.36239782\n    1 -&gt; 0.003 -&gt; 0.043841336\n    2 -&gt; 0.005 -&gt; 0.075268817\n    3 -&gt; 0.004 -&gt; 0.059322034\n    4 -&gt; 0.005 -&gt; 0.075268817\n    5 -&gt; 0.005 -&gt; 0.075268817\n    6 -&gt; 0.003 -&gt; 0.043841336\n    7 -&gt; 0.005 -&gt; 0.075268817\n    8 -&gt; 0 -&gt; 0\n    9 -&gt; 0 -&gt; 0\n    10 -&gt; 0 -&gt; 0\n    11 -&gt; 0.003 -&gt; 0.043841336\n    12 -&gt; 0.003 -&gt; 0.043841336\n    13 -&gt; 0.001 -&gt; 0.014198783\n    14 -&gt; 0.003 -&gt; 0.043841336\n    15 -&gt; 0 -&gt; 0\n    16 -&gt; 0.002 -&gt; 0.028806584\n    17 -&gt; 0.001 -&gt; 0.014198783\n    18 -&gt; 0.002 -&gt; 0.028806584\n    19 -&gt; 0.004 -&gt; 0.059322034\n    20 -&gt; 0 -&gt; 0\n    21 -&gt; 0.008 -&gt; 0.126126126\n    22 -&gt; 0.002 -&gt; 0.028806584\n    23 -&gt; 0.005 -&gt; 0.075268817\n    24 -&gt; 0 -&gt; 0\n    25 -&gt; 0.013 -&gt; 0.222493888\n    26 -&gt; 0.002 -&gt; 0.028806584\n    27 -&gt; 0 -&gt; 0\n    \"\"\"\n\nFile sizes from @Alexander Liao and @anokas (https://www.kaggle.com/alexanderliao/exploratory-data-analysis/notebook)\n\n    train                         14.0321076GB (124288 files)\n    test                          4.6823579GB (46808 files)\n    train.csv                     0.0012744GB\n    sample_submission.csv         0.0004564GB\n\n        {\n        0: 'Nucleoplasm',\n        1: 'Nuclear membrane',\n        2: 'Nucleoli',\n        3: 'Nucleoli fibrillar center',\n        4: 'Nuclear speckles',\n        5: 'Nuclear bodies',\n        6: 'Endoplasmic reticulum',\n        7: 'Golgi apparatus',\n        8: 'Peroxisomes',\n        9: 'Endosomes',\n        10: 'Lysosomes',\n        11: 'Intermediate filaments',\n        12: 'Actin filaments',\n        13: 'Focal adhesion sites',\n        14: 'Microtubules',\n        15: 'Microtubule ends',\n        16: 'Cytokinetic bridge',\n        17: 'Mitotic spindle',\n        18: 'Microtubule organizing center',\n        19: 'Centrosome',\n        20: 'Lipid droplets',\n        21: 'Plasma membrane',\n        22: 'Cell junctions',\n        23: 'Mitochondria',\n        24: 'Aggresome',\n        25: 'Cytosol',\n        26: 'Cytoplasmic bodies',\n        27: 'Rods &amp; rings'}",
      "votes": null
    },
    {
      "id": "416723",
      "postDate": "11/07/2018 07:08:52",
      "content": "<p>Wouldn't it be better to use combined stats of train and test datasets? Also, would be great if you could share this as a kernel.</p>",
      "rawMarkdown": "Wouldn't it be better to use combined stats of train and test datasets? Also, would be great if you could share this as a kernel.",
      "votes": null
    },
    {
      "id": "417225",
      "postDate": "11/08/2018 00:50:32",
      "content": "<p>It would make sense to use combined stats of train and test. But the data deviation above shows that it is unlikely for the train and test set to come from the same distribution.</p>",
      "rawMarkdown": "It would make sense to use combined stats of train and test. But the data deviation above shows that it is unlikely for the train and test set to come from the same distribution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 416723,
      "author_name": "sidujjain",
      "author_url": "",
      "post_date": "11/07/2018 07:08:52",
      "content": "<p>Wouldn't it be better to use combined stats of train and test datasets? Also, would be great if you could share this as a kernel.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417225,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/08/2018 00:50:32",
          "content": "<p>It would make sense to use combined stats of train and test. But the data deviation above shows that it is unlikely for the train and test set to come from the same distribution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "416701": "I put these data calculated through the training set so that you can use them. (If you got a different result, please notify me!) Feel free to add some more.\n\nMean and Variance for normalization:\n\n    \"\"\"\n    Train Data:\n        Mean = [0.0804419,  0.05262986, 0.05474701, 0.08270896]\n        Var  = [0.0025557  0.0023054  0.0012995  0.00293925]\n        Var1 = [0.00255578 0.00230547 0.00129955 0.00293934]\n    \"\"\"\n\n    \"\"\"\n    Test Data:\n        Mean = [0.05908022, 0.04532852, 0.04065233, 0.05923426]\n        Var  = [0.00235361 0.0020402  0.00137821 0.00246495]\n        Var1 = [0.00235381 0.00204037 0.00137833 0.00246516]\n    \"\"\"\n\nMean of train+test data: [0.07459783,  0.05063238,  0.05089102,  0.07628681]\n\n\n(Also see: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462)\n\nThe label distribution for weight loss:\n\n    \"\"\"\n    0     12885\n    25     8228\n    21     3777\n    2      3621\n    23     2965\n    7      2822\n    5      2513\n    4      1858\n    3      1561\n    19     1482\n    1      1254\n    11     1093\n    14     1066\n    6      1008\n    18      902\n    22      802\n    12      688\n    13      537\n    16      530\n    26      328\n    24      322\n    17      210\n    20      172\n    8        53\n    9        45\n    10       28\n    15       21\n    27       11\n    \"\"\"\n\nLB Probing (from @Iafoss https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678)\n\n    \"\"\"LB Probing\n    0 -&gt; 0.019 -&gt; 0.36239782\n    1 -&gt; 0.003 -&gt; 0.043841336\n    2 -&gt; 0.005 -&gt; 0.075268817\n    3 -&gt; 0.004 -&gt; 0.059322034\n    4 -&gt; 0.005 -&gt; 0.075268817\n    5 -&gt; 0.005 -&gt; 0.075268817\n    6 -&gt; 0.003 -&gt; 0.043841336\n    7 -&gt; 0.005 -&gt; 0.075268817\n    8 -&gt; 0 -&gt; 0\n    9 -&gt; 0 -&gt; 0\n    10 -&gt; 0 -&gt; 0\n    11 -&gt; 0.003 -&gt; 0.043841336\n    12 -&gt; 0.003 -&gt; 0.043841336\n    13 -&gt; 0.001 -&gt; 0.014198783\n    14 -&gt; 0.003 -&gt; 0.043841336\n    15 -&gt; 0 -&gt; 0\n    16 -&gt; 0.002 -&gt; 0.028806584\n    17 -&gt; 0.001 -&gt; 0.014198783\n    18 -&gt; 0.002 -&gt; 0.028806584\n    19 -&gt; 0.004 -&gt; 0.059322034\n    20 -&gt; 0 -&gt; 0\n    21 -&gt; 0.008 -&gt; 0.126126126\n    22 -&gt; 0.002 -&gt; 0.028806584\n    23 -&gt; 0.005 -&gt; 0.075268817\n    24 -&gt; 0 -&gt; 0\n    25 -&gt; 0.013 -&gt; 0.222493888\n    26 -&gt; 0.002 -&gt; 0.028806584\n    27 -&gt; 0 -&gt; 0\n    \"\"\"\n\nFile sizes from @Alexander Liao and @anokas (https://www.kaggle.com/alexanderliao/exploratory-data-analysis/notebook)\n\n    train                         14.0321076GB (124288 files)\n    test                          4.6823579GB (46808 files)\n    train.csv                     0.0012744GB\n    sample_submission.csv         0.0004564GB\n\n        {\n        0: 'Nucleoplasm',\n        1: 'Nuclear membrane',\n        2: 'Nucleoli',\n        3: 'Nucleoli fibrillar center',\n        4: 'Nuclear speckles',\n        5: 'Nuclear bodies',\n        6: 'Endoplasmic reticulum',\n        7: 'Golgi apparatus',\n        8: 'Peroxisomes',\n        9: 'Endosomes',\n        10: 'Lysosomes',\n        11: 'Intermediate filaments',\n        12: 'Actin filaments',\n        13: 'Focal adhesion sites',\n        14: 'Microtubules',\n        15: 'Microtubule ends',\n        16: 'Cytokinetic bridge',\n        17: 'Mitotic spindle',\n        18: 'Microtubule organizing center',\n        19: 'Centrosome',\n        20: 'Lipid droplets',\n        21: 'Plasma membrane',\n        22: 'Cell junctions',\n        23: 'Mitochondria',\n        24: 'Aggresome',\n        25: 'Cytosol',\n        26: 'Cytoplasmic bodies',\n        27: 'Rods &amp; rings'}",
    "416723": "Wouldn't it be better to use combined stats of train and test datasets? Also, would be great if you could share this as a kernel.",
    "417225": "It would make sense to use combined stats of train and test. But the data deviation above shows that it is unlikely for the train and test set to come from the same distribution."
  },
  "source": "meta"
}