{
  "id": 30750,
  "title": "Manually creating bounding boxes on training data",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30750",
  "author_name": "",
  "post_date": "2017-03-28T00:26:30.726990100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>@Wendy:</p>\n\n<p>Could one of the admins of this competition please clarify / provide a statement on manual annotation of the training data.</p>\n\n<p>To identify cervix areas in the training images and to train a detector network on that data I want to annotate (some of) the training images with bounding boxes telling the detector which areas to look at to identify the cervix. </p>\n\n<p>A similar question has already been asked in this thread: <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30471\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30471</a></p>\n\n<p>The rules only state that \"...<em>Submissions may not use or incorporate information from hand labeling or human prediction of the <strong>validation dataset</strong> or <strong>test data records</strong></em>...\"</p>\n\n<p>Do I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?</p>",
  "messages": [
    {
      "id": "170917",
      "postDate": "03/28/2017 00:26:30",
      "content": "<p>@Wendy:</p>\n\n<p>Could one of the admins of this competition please clarify / provide a statement on manual annotation of the training data.</p>\n\n<p>To identify cervix areas in the training images and to train a detector network on that data I want to annotate (some of) the training images with bounding boxes telling the detector which areas to look at to identify the cervix. </p>\n\n<p>A similar question has already been asked in this thread: <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30471\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30471</a></p>\n\n<p>The rules only state that \"...<em>Submissions may not use or incorporate information from hand labeling or human prediction of the <strong>validation dataset</strong> or <strong>test data records</strong></em>...\"</p>\n\n<p>Do I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?</p>",
      "rawMarkdown": "Wendy:\n\nCould one of the admins of this competition please clarify / provide a statement on manual annotation of the training data.\n\nTo identify cervix areas in the training images and to train a detector network on that data I want to annotate (some of) the training images with bounding boxes telling the detector which areas to look at to identify the cervix. \n\nA similar question has already been asked in this thread: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30471\n\nThe rules only state that \"...*Submissions may not use or incorporate information from hand labeling or human prediction of the **validation dataset** or **test data records***...\"\n\nDo I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?",
      "votes": null
    },
    {
      "id": "170918",
      "postDate": "03/28/2017 00:35:31",
      "content": "<blockquote>\n  <p><strong>FPP_UK wrote</strong></p>\n  \n  <p>Do I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?</p>\n</blockquote>\n\n<p>You are correct, you are allowed to do whatever you want with the training dataset, and you don't need to share it publicly either (though nobody here is going to complain if you want to :-).</p>\n\n<p>After training, though, you must only use automated methods to produce the test set labels for submission. You can use pseudo-labeling of the test set to further train your classifier, but you can't manually label it.</p>",
      "rawMarkdown": "> **FPP_UK wrote**\n> \n> Do I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?\n> \n\nYou are correct, you are allowed to do whatever you want with the training dataset, and you don't need to share it publicly either (though nobody here is going to complain if you want to :-).\n\nAfter training, though, you must only use automated methods to produce the test set labels for submission. You can use pseudo-labeling of the test set to further train your classifier, but you can't manually label it.",
      "votes": null
    },
    {
      "id": "171824",
      "postDate": "03/31/2017 15:49:45",
      "content": "<p>Does anyone know if manual annotations would count as \"external data\" sources? The rules state that external data (along with pre-trained models) are totally fine as long as they are shared to the forum. Would manual annotations fall into the same restriction?</p>",
      "rawMarkdown": "Does anyone know if manual annotations would count as \"external data\" sources? The rules state that external data (along with pre-trained models) are totally fine as long as they are shared to the forum. Would manual annotations fall into the same restriction?",
      "votes": null
    },
    {
      "id": "173904",
      "postDate": "04/09/2017 08:01:21",
      "content": "<p>Manual Annotations don't count as external data. Annotating the training set and additional set you have been given is totally fine and within bounds. You don't even need to share the annotations.</p>",
      "rawMarkdown": "Manual Annotations don't count as external data. Annotating the training set and additional set you have been given is totally fine and within bounds. You don't even need to share the annotations.",
      "votes": null
    },
    {
      "id": "175043",
      "postDate": "04/13/2017 15:18:05",
      "content": "<p>Hi, </p>\n\n<p>Just to let you know that I did label the Type_1 dataset. The script to label and annotations are available here:\n<a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565</a></p>",
      "rawMarkdown": "Hi, \n\nJust to let you know that I did label the Type_1 dataset. The script to label and annotations are available here:\nhttps://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 170918,
      "author_name": "mumech",
      "author_url": "",
      "post_date": "03/28/2017 00:35:31",
      "content": "<blockquote>\n  <p><strong>FPP_UK wrote</strong></p>\n  \n  <p>Do I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?</p>\n</blockquote>\n\n<p>You are correct, you are allowed to do whatever you want with the training dataset, and you don't need to share it publicly either (though nobody here is going to complain if you want to :-).</p>\n\n<p>After training, though, you must only use automated methods to produce the test set labels for submission. You can use pseudo-labeling of the test set to further train your classifier, but you can't manually label it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 175043,
          "author_name": "deveaup",
          "author_url": "",
          "post_date": "04/13/2017 15:18:05",
          "content": "<p>Hi, </p>\n\n<p>Just to let you know that I did label the Type_1 dataset. The script to label and annotations are available here:\n<a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 171824,
      "author_name": "gkericks",
      "author_url": "",
      "post_date": "03/31/2017 15:49:45",
      "content": "<p>Does anyone know if manual annotations would count as \"external data\" sources? The rules state that external data (along with pre-trained models) are totally fine as long as they are shared to the forum. Would manual annotations fall into the same restriction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 173904,
          "author_name": "yadavsarthak",
          "author_url": "",
          "post_date": "04/09/2017 08:01:21",
          "content": "<p>Manual Annotations don't count as external data. Annotating the training set and additional set you have been given is totally fine and within bounds. You don't even need to share the annotations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "170917": "Wendy:\n\nCould one of the admins of this competition please clarify / provide a statement on manual annotation of the training data.\n\nTo identify cervix areas in the training images and to train a detector network on that data I want to annotate (some of) the training images with bounding boxes telling the detector which areas to look at to identify the cervix. \n\nA similar question has already been asked in this thread: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30471\n\nThe rules only state that \"...*Submissions may not use or incorporate information from hand labeling or human prediction of the **validation dataset** or **test data records***...\"\n\nDo I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?",
    "170918": "> **FPP_UK wrote**\n> \n> Do I understand this correctly that we can pre-process the training data in any way we see fit (manually and/or automatically), but all following steps have to be automated (i.e. without manual annotations / steps)?\n> \n\nYou are correct, you are allowed to do whatever you want with the training dataset, and you don't need to share it publicly either (though nobody here is going to complain if you want to :-).\n\nAfter training, though, you must only use automated methods to produce the test set labels for submission. You can use pseudo-labeling of the test set to further train your classifier, but you can't manually label it.",
    "171824": "Does anyone know if manual annotations would count as \"external data\" sources? The rules state that external data (along with pre-trained models) are totally fine as long as they are shared to the forum. Would manual annotations fall into the same restriction?",
    "173904": "Manual Annotations don't count as external data. Annotating the training set and additional set you have been given is totally fine and within bounds. You don't even need to share the annotations.",
    "175043": "Hi, \n\nJust to let you know that I did label the Type_1 dataset. The script to label and annotations are available here:\nhttps://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565"
  },
  "source": "meta"
}