{
  "id": 289552,
  "title": "What manual manipulation of train data is allowed?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/289552",
  "author_name": "Slawek Biel",
  "post_date": "2021-11-20T20:56:47.604000",
  "votes": 17,
  "comment_count": 19,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/279488\" target=\"_blank\">This thread</a> makes a compelling case that the annotations we were given were incorrectly encoded (and I was also able to <a href=\"https://www.kaggle.com/slawekbiel/broken-mask-example\" target=\"_blank\">reproduce the failure case</a>. This wasn't addressed, and now halfway through the competition it doesn't seem like it will be. So what I'm now wondering is are we allowed to try to fix the most glaring cases ourselves?</p>\n<p>The rules state</p>\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n<p>I'm not sure what \"validation set\" means in this context? In the training set provided are we allowed to remove annotations based on manual inspection? Can we mark some of them and treat them differently in training? Finally can we try to redraw the boundaries by hand?</p>",
  "messages": [
    {
      "id": 1590039,
      "postDate": "2021-11-20T20:56:47.603Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/279488\" target=\"_blank\">This thread</a> makes a compelling case that the annotations we were given were incorrectly encoded (and I was also able to <a href=\"https://www.kaggle.com/slawekbiel/broken-mask-example\" target=\"_blank\">reproduce the failure case</a>. This wasn't addressed, and now halfway through the competition it doesn't seem like it will be. So what I'm now wondering is are we allowed to try to fix the most glaring cases ourselves?</p>\n<p>The rules state</p>\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n<p>I'm not sure what \"validation set\" means in this context? In the training set provided are we allowed to remove annotations based on manual inspection? Can we mark some of them and treat them differently in training? Finally can we try to redraw the boundaries by hand?</p>",
      "rawMarkdown": "[This thread](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/279488) makes a compelling case that the annotations we were given were incorrectly encoded (and I was also able to [reproduce the failure case](https://www.kaggle.com/slawekbiel/broken-mask-example). This wasn't addressed, and now halfway through the competition it doesn't seem like it will be. So what I'm now wondering is are we allowed to try to fix the most glaring cases ourselves?\n\nThe rules state\n> Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nI'm not sure what \"validation set\" means in this context? In the training set provided are we allowed to remove annotations based on manual inspection? Can we mark some of them and treat them differently in training? Finally can we try to redraw the boundaries by hand?",
      "votes": 17
    },
    {
      "id": 1590043,
      "postDate": "2021-11-20T21:03:42.110Z",
      "content": "<p>From my experience, you can do whatever you want with the training data.</p>",
      "rawMarkdown": "From my experience, you can do whatever you want with the training data.",
      "votes": 14
    },
    {
      "id": 1666188,
      "postDate": "2022-01-27T14:29:31.310Z",
      "content": "<p>I like what <a href=\"https://www.kaggle.com/Theo\" target=\"_blank\">@Theo</a> Viel <a href=\"https://www.kaggle.com/Chris\" target=\"_blank\">@Chris</a> Deotte <a href=\"https://www.kaggle.com/RB\" target=\"_blank\">@RB</a> seem to agree on, that participants are allowed to alter the training data if they want to.</p>\n<p>If altering the data is allowed and you have projected how much training data you will need, you can decide to get a method that will enable you to enhance the data in a way that will give you good results/performance.  For example, maybe</p>\n<ul>\n<li><p>Leverage semi-supervision to predict labels for data, assign them proxy labels and if the criteria match with original data, you can add to the training data set (tri-training). It can reduce the volume you need to manually label while giving you more of annotated data.</p></li>\n<li><p>Use transfer learning – using a model that has been trained on a similar dataset &amp; fine-tune it to achieve the required results. Useful when the source and target have some similarities but are not identical.</p></li>\n<li><p>Increase your data points through data augmentation. In fact, <a href=\"https://deepchecks.com/glossary/data-augmentation/\" target=\"_blank\">data augmentation</a> methods improve overall performance and several augmentation methods positively affect the model. For images, it could be something like rotation, cropping, flipping, scaling, varying brightness, translation, or color casting.</p></li>\n</ul>\n<p>Then, there’s this open-source <a href=\"https://github.com/deepchecks/deepchecks\" target=\"_blank\">package</a> that can help you examine different aspects of your data and models including duplicates, mismatch, performance_overfit, data_sample_leakage, single_feature_contribution, mixed nulls, unused features, rare_format_detection, and many more. </p>",
      "rawMarkdown": "I like what @Theo Viel @Chris Deotte @RB seem to agree on, that participants are allowed to alter the training data if they want to.\n\nIf altering the data is allowed and you have projected how much training data you will need, you can decide to get a method that will enable you to enhance the data in a way that will give you good results/performance.  For example, maybe\n\n-     Leverage semi-supervision to predict labels for data, assign them proxy labels and if the criteria match with original data, you can add to the training data set (tri-training). It can reduce the volume you need to manually label while giving you more of annotated data.\n\n-    Use transfer learning – using a model that has been trained on a similar dataset & fine-tune it to achieve the required results. Useful when the source and target have some similarities but are not identical.\n\n-    Increase your data points through data augmentation. In fact, [data augmentation](https://deepchecks.com/glossary/data-augmentation/) methods improve overall performance and several augmentation methods positively affect the model. For images, it could be something like rotation, cropping, flipping, scaling, varying brightness, translation, or color casting.\n\nThen, there’s this open-source [package](https://github.com/deepchecks/deepchecks) that can help you examine different aspects of your data and models including duplicates, mismatch, performance_overfit, data_sample_leakage, single_feature_contribution, mixed nulls, unused features, rare_format_detection, and many more. \n",
      "votes": 2
    },
    {
      "id": 1599049,
      "postDate": "2021-11-29T04:56:10.150Z",
      "content": "<p>You are allowed to alter the training data (and training labels) any way you want.</p>",
      "rawMarkdown": "You are allowed to alter the training data (and training labels) any way you want.",
      "votes": 2,
      "replies": [
        {
          "id": 1604466,
          "postDate": "2021-12-03T12:09:41.590Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi! Is hand-labeling the training semi-supervised data allowed? Just to make sure I won't break the rules if I do this. Thank you!</p>",
          "rawMarkdown": "@cdeotte Hi! Is hand-labeling the training semi-supervised data allowed? Just to make sure I won't break the rules if I do this. Thank you!"
        }
      ]
    },
    {
      "id": 1599345,
      "postDate": "2021-11-29T11:05:37.713Z",
      "content": "<p>Has anyone tried unsupervised methods for data pre-processing for images from Live Cell Dataset? If so, did it help?</p>",
      "rawMarkdown": "Has anyone tried unsupervised methods for data pre-processing for images from Live Cell Dataset? If so, did it help?"
    },
    {
      "id": 1590520,
      "postDate": "2021-11-21T11:15:46.173Z",
      "content": "<p>the thing could be used  live cell data, as mentioned in data part by host.</p>\n<p>Did anyone use it in training?</p>",
      "rawMarkdown": "the thing could be used  live cell data, as mentioned in data part by host.\n\nDid anyone use it in training?",
      "replies": [
        {
          "id": 1590590,
          "postDate": "2021-11-21T12:57:57.003Z",
          "content": "<p>I do, just added as is into new categories, except astro.</p>",
          "rawMarkdown": "I do, just added as is into new categories, except astro."
        },
        {
          "id": 1590595,
          "postDate": "2021-11-21T13:10:43.110Z",
          "content": "<p>thank for your sharing.</p>\n<p>what is the LB difference between added and none?</p>",
          "rawMarkdown": "thank for your sharing.\n\nwhat is the LB difference between added and none?"
        },
        {
          "id": 1590652,
          "postDate": "2021-11-21T14:35:52.880Z",
          "content": "<p>About 0.05. Since then I am using that data (union of test/val/train) for training only, can't say how it's affecting now. Further boosting 305-&gt;325 is due to postprocessing only.</p>",
          "rawMarkdown": "About 0.05. Since then I am using that data (union of test/val/train) for training only, can't say how it's affecting now. Further boosting 305->325 is due to postprocessing only.",
          "votes": 3
        },
        {
          "id": 1590672,
          "postDate": "2021-11-21T14:53:31.260Z",
          "content": "<p>thx for sharing.</p>",
          "rawMarkdown": "thx for sharing."
        },
        {
          "id": 1591903,
          "postDate": "2021-11-22T17:30:25.480Z",
          "content": "<p><a href=\"https://www.kaggle.com/rednikotin\" target=\"_blank\">@rednikotin</a> , may I ask whether you mean that you included the labeled LIVECELL data for SHSY5Y or also some other, unlabeled data?</p>",
          "rawMarkdown": "@rednikotin , may I ask whether you mean that you included the labeled LIVECELL data for SHSY5Y or also some other, unlabeled data?"
        },
        {
          "id": 1592035,
          "postDate": "2021-11-22T20:10:21.393Z",
          "content": "<p><a href=\"https://www.kaggle.com/omallo\" target=\"_blank\">@omallo</a> I meant that I included SHSY5Y from LIVECELL into the same category and other categories put into individual categories (3+7 total). For unlabeled data I used pseudo labeling, but it doesn't help much.</p>",
          "rawMarkdown": "@omallo I meant that I included SHSY5Y from LIVECELL into the same category and other categories put into individual categories (3+7 total). For unlabeled data I used pseudo labeling, but it doesn't help much.",
          "votes": 4
        },
        {
          "id": 1594593,
          "postDate": "2021-11-25T02:10:28.380Z",
          "content": "<blockquote>\n  <p>I do, just added as is into new categories, except astro.</p>\n</blockquote>\n<p>I used the LiveCell data too. But what did you mean by \"except astro\" ? <a href=\"https://www.kaggle.com/rednikotin\" target=\"_blank\">@rednikotin</a> </p>",
          "rawMarkdown": "> I do, just added as is into new categories, except astro.\n\nI used the LiveCell data too. But what did you mean by \"except astro\" ? @rednikotin "
        },
        {
          "id": 1594851,
          "postDate": "2021-11-25T07:02:03.667Z",
          "content": "<p>Sorry, I was not clear, I added astro into the same category as original astro and other cell type into individual categories. </p>",
          "rawMarkdown": "Sorry, I was not clear, I added astro into the same category as original astro and other cell type into individual categories. ",
          "votes": 1
        },
        {
          "id": 1594899,
          "postDate": "2021-11-25T08:09:17.907Z",
          "content": "<p>I did not find astro category in LiveCell dataset</p>",
          "rawMarkdown": "I did not find astro category in LiveCell dataset"
        },
        {
          "id": 1594905,
          "postDate": "2021-11-25T08:24:07.823Z",
          "content": "<p>Sorry <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a>, I did that SHSY5Y.</p>",
          "rawMarkdown": "Sorry @namgalielei, I did that SHSY5Y.",
          "votes": 1
        },
        {
          "id": 1597018,
          "postDate": "2021-11-27T05:45:26.743Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1601509,
          "postDate": "2021-12-01T11:12:23.990Z",
          "content": "<p>Sorry, May I ask a problem about post processing method? I am looking at some post processing methods recently. But it doesn't work as your boost. Are your methods are about how to solve the no-overlapping requirement? </p>",
          "rawMarkdown": "Sorry, May I ask a problem about post processing method? I am looking at some post processing methods recently. But it doesn't work as your boost. Are your methods are about how to solve the no-overlapping requirement? "
        }
      ]
    },
    {
      "id": 1590054,
      "postDate": "2021-11-20T21:22:28.400Z",
      "content": "<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n<p>This is one of the reason they only give us minimal test data. Like Theo said, do whatever with training dataset. </p>",
      "rawMarkdown": "> Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nThis is one of the reason they only give us minimal test data. Like Theo said, do whatever with training dataset. "
    }
  ],
  "comments": [
    {
      "id": 1590043,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-11-20T21:03:42.110000",
      "content": "<p>From my experience, you can do whatever you want with the training data.</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 1666188,
      "author_name": "Kristopher Briggs",
      "author_url": "",
      "post_date": "2022-01-27T14:29:31.310000",
      "content": "<p>I like what <a href=\"https://www.kaggle.com/Theo\" target=\"_blank\">@Theo</a> Viel <a href=\"https://www.kaggle.com/Chris\" target=\"_blank\">@Chris</a> Deotte <a href=\"https://www.kaggle.com/RB\" target=\"_blank\">@RB</a> seem to agree on, that participants are allowed to alter the training data if they want to.</p>\n<p>If altering the data is allowed and you have projected how much training data you will need, you can decide to get a method that will enable you to enhance the data in a way that will give you good results/performance.  For example, maybe</p>\n<ul>\n<li><p>Leverage semi-supervision to predict labels for data, assign them proxy labels and if the criteria match with original data, you can add to the training data set (tri-training). It can reduce the volume you need to manually label while giving you more of annotated data.</p></li>\n<li><p>Use transfer learning – using a model that has been trained on a similar dataset &amp; fine-tune it to achieve the required results. Useful when the source and target have some similarities but are not identical.</p></li>\n<li><p>Increase your data points through data augmentation. In fact, <a href=\"https://deepchecks.com/glossary/data-augmentation/\" target=\"_blank\">data augmentation</a> methods improve overall performance and several augmentation methods positively affect the model. For images, it could be something like rotation, cropping, flipping, scaling, varying brightness, translation, or color casting.</p></li>\n</ul>\n<p>Then, there’s this open-source <a href=\"https://github.com/deepchecks/deepchecks\" target=\"_blank\">package</a> that can help you examine different aspects of your data and models including duplicates, mismatch, performance_overfit, data_sample_leakage, single_feature_contribution, mixed nulls, unused features, rare_format_detection, and many more. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1599049,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-11-29T04:56:10.150000",
      "content": "<p>You are allowed to alter the training data (and training labels) any way you want.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1604466,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2021-12-03T12:09:41.590000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi! Is hand-labeling the training semi-supervised data allowed? Just to make sure I won't break the rules if I do this. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1599345,
      "author_name": "timurishmuratov",
      "author_url": "",
      "post_date": "2021-11-29T11:05:37.713000",
      "content": "<p>Has anyone tried unsupervised methods for data pre-processing for images from Live Cell Dataset? If so, did it help?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1590520,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2021-11-21T11:15:46.173000",
      "content": "<p>the thing could be used  live cell data, as mentioned in data part by host.</p>\n<p>Did anyone use it in training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1590590,
          "author_name": "Valentin Nikotin",
          "author_url": "",
          "post_date": "2021-11-21T12:57:57.003000",
          "content": "<p>I do, just added as is into new categories, except astro.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1590595,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-11-21T13:10:43.110000",
          "content": "<p>thank for your sharing.</p>\n<p>what is the LB difference between added and none?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1590652,
          "author_name": "Valentin Nikotin",
          "author_url": "",
          "post_date": "2021-11-21T14:35:52.880000",
          "content": "<p>About 0.05. Since then I am using that data (union of test/val/train) for training only, can't say how it's affecting now. Further boosting 305-&gt;325 is due to postprocessing only.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1590672,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-11-21T14:53:31.260000",
          "content": "<p>thx for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1591903,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2021-11-22T17:30:25.480000",
          "content": "<p><a href=\"https://www.kaggle.com/rednikotin\" target=\"_blank\">@rednikotin</a> , may I ask whether you mean that you included the labeled LIVECELL data for SHSY5Y or also some other, unlabeled data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1592035,
          "author_name": "Valentin Nikotin",
          "author_url": "",
          "post_date": "2021-11-22T20:10:21.393000",
          "content": "<p><a href=\"https://www.kaggle.com/omallo\" target=\"_blank\">@omallo</a> I meant that I included SHSY5Y from LIVECELL into the same category and other categories put into individual categories (3+7 total). For unlabeled data I used pseudo labeling, but it doesn't help much.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1594593,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-11-25T02:10:28.380000",
          "content": "<blockquote>\n  <p>I do, just added as is into new categories, except astro.</p>\n</blockquote>\n<p>I used the LiveCell data too. But what did you mean by \"except astro\" ? <a href=\"https://www.kaggle.com/rednikotin\" target=\"_blank\">@rednikotin</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1594851,
          "author_name": "Valentin Nikotin",
          "author_url": "",
          "post_date": "2021-11-25T07:02:03.667000",
          "content": "<p>Sorry, I was not clear, I added astro into the same category as original astro and other cell type into individual categories. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1594899,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-11-25T08:09:17.907000",
          "content": "<p>I did not find astro category in LiveCell dataset</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1594905,
          "author_name": "Valentin Nikotin",
          "author_url": "",
          "post_date": "2021-11-25T08:24:07.823000",
          "content": "<p>Sorry <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a>, I did that SHSY5Y.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1597018,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-11-27T05:45:26.743000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1601509,
          "author_name": "shinewine",
          "author_url": "",
          "post_date": "2021-12-01T11:12:23.990000",
          "content": "<p>Sorry, May I ask a problem about post processing method? I am looking at some post processing methods recently. But it doesn't work as your boost. Are your methods are about how to solve the no-overlapping requirement? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1590054,
      "author_name": "RB",
      "author_url": "",
      "post_date": "2021-11-20T21:22:28.400000",
      "content": "<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n<p>This is one of the reason they only give us minimal test data. Like Theo said, do whatever with training dataset. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1590039": "[This thread](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/279488) makes a compelling case that the annotations we were given were incorrectly encoded (and I was also able to [reproduce the failure case](https://www.kaggle.com/slawekbiel/broken-mask-example). This wasn't addressed, and now halfway through the competition it doesn't seem like it will be. So what I'm now wondering is are we allowed to try to fix the most glaring cases ourselves?\n\nThe rules state\n> Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nI'm not sure what \"validation set\" means in this context? In the training set provided are we allowed to remove annotations based on manual inspection? Can we mark some of them and treat them differently in training? Finally can we try to redraw the boundaries by hand?",
    "1590043": "From my experience, you can do whatever you want with the training data.",
    "1666188": "I like what @Theo Viel @Chris Deotte @RB seem to agree on, that participants are allowed to alter the training data if they want to.\n\nIf altering the data is allowed and you have projected how much training data you will need, you can decide to get a method that will enable you to enhance the data in a way that will give you good results/performance.  For example, maybe\n\n-     Leverage semi-supervision to predict labels for data, assign them proxy labels and if the criteria match with original data, you can add to the training data set (tri-training). It can reduce the volume you need to manually label while giving you more of annotated data.\n\n-    Use transfer learning – using a model that has been trained on a similar dataset & fine-tune it to achieve the required results. Useful when the source and target have some similarities but are not identical.\n\n-    Increase your data points through data augmentation. In fact, [data augmentation](https://deepchecks.com/glossary/data-augmentation/) methods improve overall performance and several augmentation methods positively affect the model. For images, it could be something like rotation, cropping, flipping, scaling, varying brightness, translation, or color casting.\n\nThen, there’s this open-source [package](https://github.com/deepchecks/deepchecks) that can help you examine different aspects of your data and models including duplicates, mismatch, performance_overfit, data_sample_leakage, single_feature_contribution, mixed nulls, unused features, rare_format_detection, and many more. \n",
    "1599049": "You are allowed to alter the training data (and training labels) any way you want.",
    "1599345": "Has anyone tried unsupervised methods for data pre-processing for images from Live Cell Dataset? If so, did it help?",
    "1590520": "the thing could be used  live cell data, as mentioned in data part by host.\n\nDid anyone use it in training?",
    "1590054": "> Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nThis is one of the reason they only give us minimal test data. Like Theo said, do whatever with training dataset. "
  }
}