{
  "id": 230751,
  "title": "Public data as Kaggle dataset released",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/230751",
  "author_name": "Phil Culliton",
  "post_date": "2021-04-05T14:02:40.099000",
  "votes": 22,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Sorry for the delay about this - there were some issues to solve on my end before we could make it fully available. The public data is now available as a separate Kaggle dataset <a href=\"https://www.kaggle.com/philculliton/hpa-challenge-2021-extra-train-images\" target=\"_blank\">here</a>. Please let me know if you run into issues using it.</p>",
  "messages": [
    {
      "id": 1263552,
      "postDate": "2021-04-05T14:02:40.100Z",
      "content": "<p>Hi all,</p>\n<p>Sorry for the delay about this - there were some issues to solve on my end before we could make it fully available. The public data is now available as a separate Kaggle dataset <a href=\"https://www.kaggle.com/philculliton/hpa-challenge-2021-extra-train-images\" target=\"_blank\">here</a>. Please let me know if you run into issues using it.</p>",
      "rawMarkdown": "Hi all,\n\nSorry for the delay about this - there were some issues to solve on my end before we could make it fully available. The public data is now available as a separate Kaggle dataset [here](https://www.kaggle.com/philculliton/hpa-challenge-2021-extra-train-images). Please let me know if you run into issues using it.",
      "votes": 21
    },
    {
      "id": 1263973,
      "postDate": "2021-04-05T19:31:34.693Z",
      "content": "<p>Why are there only 92.5k files (around 23k 4-channel images). <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">This notebook</a> showed that there should be around 67k extra images in the public dataset. Did I miss something ?</p>",
      "rawMarkdown": "Why are there only 92.5k files (around 23k 4-channel images). [This notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) showed that there should be around 67k extra images in the public dataset. Did I miss something ?",
      "votes": 4,
      "replies": [
        {
          "id": 1282972,
          "postDate": "2021-04-24T12:54:02.673Z",
          "content": "<p>Hi! It is true that this dataset is smaller than the whole public dataset. This is because Kaggle limit on the size of the hosted dataset, so we had to make some selection. Here I picked all the rarer classes, and then prioritize diverse cell lines and multi-labels. </p>",
          "rawMarkdown": "Hi! It is true that this dataset is smaller than the whole public dataset. This is because Kaggle limit on the size of the hosted dataset, so we had to make some selection. Here I picked all the rarer classes, and then prioritize diverse cell lines and multi-labels. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 1281716,
      "postDate": "2021-04-23T08:07:31.717Z",
      "content": "<p><a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> <a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nAs others have asked, this dataset is small.<br>\nWhy is it small? How is the data being sampled?</p>",
      "rawMarkdown": "@lnhtrang @emmalumpan @philculliton \nAs others have asked, this dataset is small.\nWhy is it small? How is the data being sampled?",
      "replies": [
        {
          "id": 1282973,
          "postDate": "2021-04-24T12:55:15.977Z",
          "content": "<p>Hi! It is true that this dataset is smaller than the whole public dataset. This is because Kaggle limit on the size of the hosted dataset, so we had to make some selection. Here I picked all the rarer classes, and then prioritize diverse cell lines and multi-labels. </p>",
          "rawMarkdown": "Hi! It is true that this dataset is smaller than the whole public dataset. This is because Kaggle limit on the size of the hosted dataset, so we had to make some selection. Here I picked all the rarer classes, and then prioritize diverse cell lines and multi-labels. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 1264979,
      "postDate": "2021-04-06T14:30:54.413Z",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Thanks a lot for providing this dataset, super helpful! Can you share the logic which images were sampled from the bigger public HPA dataset? </p>",
      "rawMarkdown": "@philculliton Thanks a lot for providing this dataset, super helpful! Can you share the logic which images were sampled from the bigger public HPA dataset? "
    },
    {
      "id": 1263709,
      "postDate": "2021-04-05T16:09:38.590Z",
      "content": "<p>The labels can be found in the public dataset  <a href=\"https://www.kaggle.com/lnhtrang/publichpa-withcellline\" target=\"_blank\">https://www.kaggle.com/lnhtrang/publichpa-withcellline</a></p>",
      "rawMarkdown": "The labels can be found in the public dataset  https://www.kaggle.com/lnhtrang/publichpa-withcellline\n",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1263973,
      "author_name": "Duc-Kinh Le Tran",
      "author_url": "",
      "post_date": "2021-04-05T19:31:34.693000",
      "content": "<p>Why are there only 92.5k files (around 23k 4-channel images). <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">This notebook</a> showed that there should be around 67k extra images in the public dataset. Did I miss something ?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1282972,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-24T12:54:02.673000",
          "content": "<p>Hi! It is true that this dataset is smaller than the whole public dataset. This is because Kaggle limit on the size of the hosted dataset, so we had to make some selection. Here I picked all the rarer classes, and then prioritize diverse cell lines and multi-labels. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1281716,
      "author_name": "Inoichan",
      "author_url": "",
      "post_date": "2021-04-23T08:07:31.717000",
      "content": "<p><a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> <a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nAs others have asked, this dataset is small.<br>\nWhy is it small? How is the data being sampled?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1282973,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-24T12:55:15.977000",
          "content": "<p>Hi! It is true that this dataset is smaller than the whole public dataset. This is because Kaggle limit on the size of the hosted dataset, so we had to make some selection. Here I picked all the rarer classes, and then prioritize diverse cell lines and multi-labels. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1264979,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-04-06T14:30:54.413000",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Thanks a lot for providing this dataset, super helpful! Can you share the logic which images were sampled from the bigger public HPA dataset? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1263709,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-05T16:09:38.590000",
      "content": "<p>The labels can be found in the public dataset  <a href=\"https://www.kaggle.com/lnhtrang/publichpa-withcellline\" target=\"_blank\">https://www.kaggle.com/lnhtrang/publichpa-withcellline</a></p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1263552": "Hi all,\n\nSorry for the delay about this - there were some issues to solve on my end before we could make it fully available. The public data is now available as a separate Kaggle dataset [here](https://www.kaggle.com/philculliton/hpa-challenge-2021-extra-train-images). Please let me know if you run into issues using it.",
    "1263973": "Why are there only 92.5k files (around 23k 4-channel images). [This notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) showed that there should be around 67k extra images in the public dataset. Did I miss something ?",
    "1281716": "@lnhtrang @emmalumpan @philculliton \nAs others have asked, this dataset is small.\nWhy is it small? How is the data being sampled?",
    "1264979": "@philculliton Thanks a lot for providing this dataset, super helpful! Can you share the logic which images were sampled from the bigger public HPA dataset? ",
    "1263709": "The labels can be found in the public dataset  https://www.kaggle.com/lnhtrang/publichpa-withcellline\n"
  }
}