{
  "id": 332786,
  "title": "Ways to work with small number of training samples",
  "url": "/competitions/hubmap-organ-segmentation/discussion/332786",
  "author_name": "",
  "post_date": "2022-06-23T11:26:39.363341900Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The competition has only recently begun, and my first glance at the data led me to believe that there are few training samples. There are only 351 training samples, and it is expected that we will see roughly 500+ samples in testing.</p>\n<p>Because the training data is minimal, I wanted to know whether there are any specific strategies for working with less data other than transfer learning.</p>\n<p>Also, is there some external data related to this which we can use?</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1830343",
      "postDate": "06/23/2022 11:26:39",
      "content": "<p>The competition has only recently begun, and my first glance at the data led me to believe that there are few training samples. There are only 351 training samples, and it is expected that we will see roughly 500+ samples in testing.</p>\n<p>Because the training data is minimal, I wanted to know whether there are any specific strategies for working with less data other than transfer learning.</p>\n<p>Also, is there some external data related to this which we can use?</p>\n<p>Thanks</p>",
      "rawMarkdown": "The competition has only recently begun, and my first glance at the data led me to believe that there are few training samples. There are only 351 training samples, and it is expected that we will see roughly 500+ samples in testing.\n\nBecause the training data is minimal, I wanted to know whether there are any specific strategies for working with less data other than transfer learning.\n\nAlso, is there some external data related to this which we can use?\n\nThanks",
      "votes": null
    },
    {
      "id": "1830806",
      "postDate": "06/23/2022 17:16:00",
      "content": "<p>other than using augmentations there are some sources for external data<br>\n1.) previous year comp and this year both have \"glomeruli in kidney\" in common<br>\n2.)hubmap portal might also contain some useful datasets <a href=\"https://portal.hubmapconsortium.org/\" target=\"_blank\">https://portal.hubmapconsortium.org/</a></p>",
      "rawMarkdown": "other than using augmentations there are some sources for external data\n1.) previous year comp and this year both have \"glomeruli in kidney\" in common\n2.)hubmap portal might also contain some useful datasets https://portal.hubmapconsortium.org/",
      "votes": null
    },
    {
      "id": "1833184",
      "postDate": "06/25/2022 17:53:12",
      "content": "<p>Note that since this is a segmentation (multiple predictions per image) task and not a classification (1 per image) task, the number of train images is not informative.</p>\n<p>If we are asked to classify each image then we only have 351 training data points. However since we are doing segmentation, Kaggle could have given us 1 super large image with all the images concatenated together. Hence, the amount of segmentation training data cannot be deduced from the number of train images.</p>\n<p>None-the-less using external data, data augmentation, transfer learning etc etc are all good ideas.</p>",
      "rawMarkdown": "Note that since this is a segmentation (multiple predictions per image) task and not a classification (1 per image) task, the number of train images is not informative.\n  \nIf we are asked to classify each image then we only have 351 training data points. However since we are doing segmentation, Kaggle could have given us 1 super large image with all the images concatenated together. Hence, the amount of segmentation training data cannot be deduced from the number of train images.\n\nNone-the-less using external data, data augmentation, transfer learning etc etc are all good ideas.",
      "votes": null
    },
    {
      "id": "1833281",
      "postDate": "06/25/2022 20:47:47",
      "content": "<p>you may want to consider:<br>\nhow to prove number of training samples is \"small\" in the first place?<br>\nquality vs quantity ?<br>\ntrain vs test domain ?</p>",
      "rawMarkdown": "you may want to consider:\nhow to prove number of training samples is \"small\" in the first place?\nquality vs quantity ?\ntrain vs test domain ?",
      "votes": null
    },
    {
      "id": "1833679",
      "postDate": "06/26/2022 08:09:52",
      "content": "<p>you can use  external data for pertaining.<br>\nsince you do not have segmentation ground truth, you can use image classification instead<br>\ne.g. predict organ or FTU at image level</p>",
      "rawMarkdown": "you can use  external data for pertaining.\nsince you do not have segmentation ground truth, you can use image classification instead\ne.g. predict organ or FTU at image level",
      "votes": null
    },
    {
      "id": "1835186",
      "postDate": "06/27/2022 14:54:00",
      "content": "<p>My approach would be to create a tiled dataset with some minor overlap. You can get even more images when you realize that you can now rotate/augment the individual tiles as all 2d rotations will create valid images. </p>\n<p>I think this will give us plenty of images considering how many FTUs are present in the large images.</p>",
      "rawMarkdown": "My approach would be to create a tiled dataset with some minor overlap. You can get even more images when you realize that you can now rotate/augment the individual tiles as all 2d rotations will create valid images. \n\nI think this will give us plenty of images considering how many FTUs are present in the large images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1830806,
      "author_name": "mrigendraagrawal",
      "author_url": "",
      "post_date": "06/23/2022 17:16:00",
      "content": "<p>other than using augmentations there are some sources for external data<br>\n1.) previous year comp and this year both have \"glomeruli in kidney\" in common<br>\n2.)hubmap portal might also contain some useful datasets <a href=\"https://portal.hubmapconsortium.org/\" target=\"_blank\">https://portal.hubmapconsortium.org/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1833184,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/25/2022 17:53:12",
      "content": "<p>Note that since this is a segmentation (multiple predictions per image) task and not a classification (1 per image) task, the number of train images is not informative.</p>\n<p>If we are asked to classify each image then we only have 351 training data points. However since we are doing segmentation, Kaggle could have given us 1 super large image with all the images concatenated together. Hence, the amount of segmentation training data cannot be deduced from the number of train images.</p>\n<p>None-the-less using external data, data augmentation, transfer learning etc etc are all good ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1833281,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/25/2022 20:47:47",
      "content": "<p>you may want to consider:<br>\nhow to prove number of training samples is \"small\" in the first place?<br>\nquality vs quantity ?<br>\ntrain vs test domain ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1833679,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/26/2022 08:09:52",
      "content": "<p>you can use  external data for pertaining.<br>\nsince you do not have segmentation ground truth, you can use image classification instead<br>\ne.g. predict organ or FTU at image level</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1835186,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "06/27/2022 14:54:00",
      "content": "<p>My approach would be to create a tiled dataset with some minor overlap. You can get even more images when you realize that you can now rotate/augment the individual tiles as all 2d rotations will create valid images. </p>\n<p>I think this will give us plenty of images considering how many FTUs are present in the large images.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1830343": "The competition has only recently begun, and my first glance at the data led me to believe that there are few training samples. There are only 351 training samples, and it is expected that we will see roughly 500+ samples in testing.\n\nBecause the training data is minimal, I wanted to know whether there are any specific strategies for working with less data other than transfer learning.\n\nAlso, is there some external data related to this which we can use?\n\nThanks",
    "1830806": "other than using augmentations there are some sources for external data\n1.) previous year comp and this year both have \"glomeruli in kidney\" in common\n2.)hubmap portal might also contain some useful datasets https://portal.hubmapconsortium.org/",
    "1833184": "Note that since this is a segmentation (multiple predictions per image) task and not a classification (1 per image) task, the number of train images is not informative.\n  \nIf we are asked to classify each image then we only have 351 training data points. However since we are doing segmentation, Kaggle could have given us 1 super large image with all the images concatenated together. Hence, the amount of segmentation training data cannot be deduced from the number of train images.\n\nNone-the-less using external data, data augmentation, transfer learning etc etc are all good ideas.",
    "1833281": "you may want to consider:\nhow to prove number of training samples is \"small\" in the first place?\nquality vs quantity ?\ntrain vs test domain ?",
    "1833679": "you can use  external data for pertaining.\nsince you do not have segmentation ground truth, you can use image classification instead\ne.g. predict organ or FTU at image level",
    "1835186": "My approach would be to create a tiled dataset with some minor overlap. You can get even more images when you realize that you can now rotate/augment the individual tiles as all 2d rotations will create valid images. \n\nI think this will give us plenty of images considering how many FTUs are present in the large images."
  },
  "source": "meta"
}