{
  "id": 30471,
  "title": "Short tutorial on how to (humanly) recognize cervix types",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30471",
  "author_name": "",
  "post_date": "2017-03-21T21:02:40.265046Z",
  "votes": 46,
  "comment_count": 15,
  "views": 2,
  "content": "<p>Hi! @Jay Patravali asked how to manually determine which is which ( see this thread: <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30401\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30401</a> )</p>\n\n<p>I added a short answer, but now, I will answer with pictures! In full color! </p>\n\n<p>Disclaimer: 1) I have some medical training but it was long ago. 2) The images are <strong>highly graphic</strong>, you might lose your sex drive for a while . . .</p>\n\n<p>Let's begin. The most visible feature that helps us distinguish is the endocervical skin. It is red (one layer, very thin, you can see the blood and muscles) and it is in high contrast with the pinkish aspect of the rest of the cervix. </p>\n\n<p>Example of a nice Type 1 <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6156/classical_type_1.jpg\" alt=\"classical_type_1\" title=\"\"></p>\n\n<p>Type 3 has all this tissue inside the uterus so only pink skin is visible from outside <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6157/classical_type_3.jpg\" alt=\"classical_type_3.jpg\" title=\"\"></p>\n\n<p>Type 2 is a mix between them. So, part of the endocervical red skin is visible <strong>BUT</strong> also some of the pink skin is inside. This is a key difference between I and II. And imho the hardest thing to detect in this competition. <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6158/easy_type_2.jpg\" alt=\"easy_type_2.jpg\" title=\"\"></p>\n\n<p>Now, guess why this is type 2 and not 1? <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6160/type2_like_type1_01.jpg\" alt=\"type2_like_type1_01.jpg\" title=\"\"> All unilayer skin is visible. Maybe the os is dilated and you can see inside the cervix tunnel? (normally you can't)</p>\n\n<p>And this one <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6159/type2_like_type3_01.jpg\" alt=\"type2_like_type3_01.jpg\" title=\"\"> looks like type 3. Imho there is a subtle shift in color from one region to another. So even if the contrast is not good, there is \"red\" skin visible outside. So type 2. And for the classical Type 1 the red tissue looks like it gets out of the os rims. But is just a guess. </p>\n\n<p>And now let's ramp up the fun! Blurry images, blood, pathological conditions, IUDs, you name it! Check: <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6161/nasty_examples.jpg\" alt=\"nasty_examples.jpg\" title=\"\"></p>\n\n<p>I think this will be a 2 stage approach. First, detect the os/cervix and then classify. To my knowledge, a similar competition [2 stage (detect,classify) + a lot of grief] is <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition\">noaa-right-whale-recognition</a>. </p>\n\n<p>So, anybody in the mood to paint the Os on ~1400 images? Or maybe it is easier to mark clear cases on 512 images and score some leaderboard points ;)</p>\n\n<p>Hope it helps!\nCristi</p>\n\n<p>p.s. If you know how to \"insert\" the pictures in the text, without having the urls, let me know. This new Kaggle UI beats me. </p>\n\n<p>edit: hacked! 2 step process ;)</p>",
  "messages": [
    {
      "id": "169634",
      "postDate": "03/21/2017 21:02:40",
      "content": "<p>Hi! @Jay Patravali asked how to manually determine which is which ( see this thread: <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30401\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30401</a> )</p>\n\n<p>I added a short answer, but now, I will answer with pictures! In full color! </p>\n\n<p>Disclaimer: 1) I have some medical training but it was long ago. 2) The images are <strong>highly graphic</strong>, you might lose your sex drive for a while . . .</p>\n\n<p>Let's begin. The most visible feature that helps us distinguish is the endocervical skin. It is red (one layer, very thin, you can see the blood and muscles) and it is in high contrast with the pinkish aspect of the rest of the cervix. </p>\n\n<p>Example of a nice Type 1 <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6156/classical_type_1.jpg\" alt=\"classical_type_1\" title=\"\"></p>\n\n<p>Type 3 has all this tissue inside the uterus so only pink skin is visible from outside <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6157/classical_type_3.jpg\" alt=\"classical_type_3.jpg\" title=\"\"></p>\n\n<p>Type 2 is a mix between them. So, part of the endocervical red skin is visible <strong>BUT</strong> also some of the pink skin is inside. This is a key difference between I and II. And imho the hardest thing to detect in this competition. <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6158/easy_type_2.jpg\" alt=\"easy_type_2.jpg\" title=\"\"></p>\n\n<p>Now, guess why this is type 2 and not 1? <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6160/type2_like_type1_01.jpg\" alt=\"type2_like_type1_01.jpg\" title=\"\"> All unilayer skin is visible. Maybe the os is dilated and you can see inside the cervix tunnel? (normally you can't)</p>\n\n<p>And this one <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6159/type2_like_type3_01.jpg\" alt=\"type2_like_type3_01.jpg\" title=\"\"> looks like type 3. Imho there is a subtle shift in color from one region to another. So even if the contrast is not good, there is \"red\" skin visible outside. So type 2. And for the classical Type 1 the red tissue looks like it gets out of the os rims. But is just a guess. </p>\n\n<p>And now let's ramp up the fun! Blurry images, blood, pathological conditions, IUDs, you name it! Check: <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/169634/6161/nasty_examples.jpg\" alt=\"nasty_examples.jpg\" title=\"\"></p>\n\n<p>I think this will be a 2 stage approach. First, detect the os/cervix and then classify. To my knowledge, a similar competition [2 stage (detect,classify) + a lot of grief] is <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition\">noaa-right-whale-recognition</a>. </p>\n\n<p>So, anybody in the mood to paint the Os on ~1400 images? Or maybe it is easier to mark clear cases on 512 images and score some leaderboard points ;)</p>\n\n<p>Hope it helps!\nCristi</p>\n\n<p>p.s. If you know how to \"insert\" the pictures in the text, without having the urls, let me know. This new Kaggle UI beats me. </p>\n\n<p>edit: hacked! 2 step process ;)</p>",
      "rawMarkdown": "Hi! @Jay Patravali asked how to manually determine which is which ( see this thread: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30401 )\n\nI added a short answer, but now, I will answer with pictures! In full color! \n\nDisclaimer: 1) I have some medical training but it was long ago. 2) The images are **highly graphic**, you might lose your sex drive for a while . . .\n\nLet's begin. The most visible feature that helps us distinguish is the endocervical skin. It is red (one layer, very thin, you can see the blood and muscles) and it is in high contrast with the pinkish aspect of the rest of the cervix. \n\nExample of a nice Type 1 ![classical_type_1](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6156/classical_type_1.jpg)\n\nType 3 has all this tissue inside the uterus so only pink skin is visible from outside ![classical_type_3.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6157/classical_type_3.jpg)\n\nType 2 is a mix between them. So, part of the endocervical red skin is visible **BUT** also some of the pink skin is inside. This is a key difference between I and II. And imho the hardest thing to detect in this competition. ![easy_type_2.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6158/easy_type_2.jpg)\n\nNow, guess why this is type 2 and not 1? ![type2_like_type1_01.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6160/type2_like_type1_01.jpg) All unilayer skin is visible. Maybe the os is dilated and you can see inside the cervix tunnel? (normally you can't)\n\n And this one ![type2_like_type3_01.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6159/type2_like_type3_01.jpg) looks like type 3. Imho there is a subtle shift in color from one region to another. So even if the contrast is not good, there is \"red\" skin visible outside. So type 2. And for the classical Type 1 the red tissue looks like it gets out of the os rims. But is just a guess. \n\nAnd now let's ramp up the fun! Blurry images, blood, pathological conditions, IUDs, you name it! Check: ![nasty_examples.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6161/nasty_examples.jpg)\n\nI think this will be a 2 stage approach. First, detect the os/cervix and then classify. To my knowledge, a similar competition [2 stage (detect,classify) + a lot of grief] is [noaa-right-whale-recognition](https://www.kaggle.com/c/noaa-right-whale-recognition). \n\nSo, anybody in the mood to paint the Os on ~1400 images? Or maybe it is easier to mark clear cases on 512 images and score some leaderboard points ;)\n\nHope it helps!\nCristi\n\np.s. If you know how to \"insert\" the pictures in the text, without having the urls, let me know. This new Kaggle UI beats me. \n\nedit: hacked! 2 step process ;)",
      "votes": null
    },
    {
      "id": "169639",
      "postDate": "03/21/2017 21:18:51",
      "content": "<p>I labeled 400 train set images with bounding box and then trained detector networks on them. Detectors worked well, but cv score was not immediatly better for classification and have not got around to generating submission to see if helps with lb score.  </p>\n\n<p>When doing bounding box, or pixel labeling it was not clear to me how small to make the bounding box, so my initial work may not have been best way.  Thanks for detailed post will help with better bounding box labeling. </p>\n\n<p>edit: I guess an important question is how far from the os does the relevent data extend. It may be easier to detect cervix then os alone as there is quite a bit of variation between images.</p>",
      "rawMarkdown": "I labeled 400 train set images with bounding box and then trained detector networks on them. Detectors worked well, but cv score was not immediatly better for classification and have not got around to generating submission to see if helps with lb score.  \n\nWhen doing bounding box, or pixel labeling it was not clear to me how small to make the bounding box, so my initial work may not have been best way.  Thanks for detailed post will help with better bounding box labeling. \n\nedit: I guess an important question is how far from the os does the relevent data extend. It may be easier to detect cervix then os alone as there is quite a bit of variation between images.",
      "votes": null
    },
    {
      "id": "169640",
      "postDate": "03/21/2017 21:19:30",
      "content": "<p>Great, thanks! Do you know whether the data is correctly labelled in most cases? Additional data seems to add noise to model training - Part me was thinking that the additional images may have many more wrong labels ?</p>",
      "rawMarkdown": "Great, thanks! Do you know whether the data is correctly labelled in most cases? Additional data seems to add noise to model training - Part me was thinking that the additional images may have many more wrong labels ?",
      "votes": null
    },
    {
      "id": "169651",
      "postDate": "03/21/2017 23:05:32",
      "content": "<p>Thanks a bunch.  I was going to bother my nurse sister about this, but now can avoid this...</p>\n\n<p>I would be curious what the human rate of agreement is on these categories is -- the distinction between 1 and 2 is not always clear to me, and I don't know whether this is really more picking a variable on a scale than a classification task.</p>",
      "rawMarkdown": "Thanks a bunch.  I was going to bother my nurse sister about this, but now can avoid this...\n\nI would be curious what the human rate of agreement is on these categories is -- the distinction between 1 and 2 is not always clear to me, and I don't know whether this is really more picking a variable on a scale than a classification task.",
      "votes": null
    },
    {
      "id": "169738",
      "postDate": "03/22/2017 11:25:30",
      "content": "<p>thanks for the clear explanation</p>",
      "rawMarkdown": "thanks for the clear explanation",
      "votes": null
    },
    {
      "id": "169739",
      "postDate": "03/22/2017 11:29:23",
      "content": "<p>hi, is this generally \"ok\" / not against the rules? I mean I always assumed that all steps which lead to the result must be done \"automatically\" - i.e. no manual labelling of test images (which is obvious) but I also thought no manual classification / segmentation of the train data as preparation is allowed, was I wrong?</p>",
      "rawMarkdown": "hi, is this generally \"ok\" / not against the rules? I mean I always assumed that all steps which lead to the result must be done \"automatically\" - i.e. no manual labelling of test images (which is obvious) but I also thought no manual classification / segmentation of the train data as preparation is allowed, was I wrong?",
      "votes": null
    },
    {
      "id": "169747",
      "postDate": "03/22/2017 12:18:16",
      "content": "<p>You can do everything with the train data :) \nWhen you train your algorithms on manually annotated data, they still automatically make predictions on test date.</p>",
      "rawMarkdown": "You can do everything with the train data :) \nWhen you train your algorithms on manually annotated data, they still automatically make predictions on test date.",
      "votes": null
    },
    {
      "id": "169784",
      "postDate": "03/22/2017 15:49:50",
      "content": "<p>Nice @Devin! I am planning to do that myself when I will have some quality hours to put in.</p>\n\n<p>I have no idea how large the bbox should be. Maybe the entire cervix? Just to make sure that you scale everything to the same proportion. Otherwise is tricky to take the same percent every time.</p>\n\n<p>Indeed, cervix is more stable than the os. And is roughly in the middle (given or take some projections) </p>\n\n<p>@Wolhow123 true, you can do whatever you what with train data. This is also true with test data as long as is <strong>automated</strong>. Of course, nobody is checking this for non top 3 models. (aka you might get in top 5 manually labeling everything)</p>",
      "rawMarkdown": "Nice @Devin! I am planning to do that myself when I will have some quality hours to put in.\n\nI have no idea how large the bbox should be. Maybe the entire cervix? Just to make sure that you scale everything to the same proportion. Otherwise is tricky to take the same percent every time.\n\nIndeed, cervix is more stable than the os. And is roughly in the middle (given or take some projections) \n\n@Wolhow123 true, you can do whatever you what with train data. This is also true with test data as long as is **automated**. Of course, nobody is checking this for non top 3 models. (aka you might get in top 5 manually labeling everything)",
      "votes": null
    },
    {
      "id": "169825",
      "postDate": "03/22/2017 20:26:06",
      "content": "<p>Imho a classification task. And you can bother your sister with corner cases and probably for some more insights!</p>",
      "rawMarkdown": "Imho a classification task. And you can bother your sister with corner cases and probably for some more insights!",
      "votes": null
    },
    {
      "id": "169826",
      "postDate": "03/22/2017 20:28:45",
      "content": "<p>@Devin what are you using for annotations? I used <a href=\"https://github.com/cvhciKIT/sloth\">Sloth</a> a lot but is a bit sluggish and not as productive as I would like it to be.</p>",
      "rawMarkdown": "Devin what are you using for annotations? I used [Sloth](https://github.com/cvhciKIT/sloth) a lot but is a bit sluggish and not as productive as I would like it to be.",
      "votes": null
    },
    {
      "id": "169831",
      "postDate": "03/22/2017 20:55:12",
      "content": "<p>I am currently using imagej(fiji), with a macro to save the coordinates of bounding box with single keypress. I can load the next image with one command(<code>shift-o</code>), and then save with one keypress. An easy improvement would be to autoload the next image after saving the previous. This is pretty fast method though for bounding box.  You have to install the macro at the start of the session to get it to work. </p>\n\n<pre><code>macro \"Macro 1 [g]\" {\nname = getInfo(\"image.filename\")\nsaveAs(\"XY Coordinates\", \"/path/where/you/save/annotations/\"+name+\".txt\");\n} \n</code></pre>",
      "rawMarkdown": "I am currently using imagej(fiji), with a macro to save the coordinates of bounding box with single keypress. I can load the next image with one command(`shift-o`), and then save with one keypress. An easy improvement would be to autoload the next image after saving the previous. This is pretty fast method though for bounding box.  You have to install the macro at the start of the session to get it to work. \n\n    macro \"Macro 1 [g]\" {\n\tname = getInfo(\"image.filename\")\n\tsaveAs(\"XY Coordinates\", \"/path/where/you/save/annotations/\"+name+\".txt\");\n    }",
      "votes": null
    },
    {
      "id": "169919",
      "postDate": "03/23/2017 08:18:24",
      "content": "<p>Hehe! My good old friend forgot in the drawer! I had some nice scripts once for imageJ.</p>",
      "rawMarkdown": "Hehe! My good old friend forgot in the drawer! I had some nice scripts once for imageJ.",
      "votes": null
    },
    {
      "id": "169943",
      "postDate": "03/23/2017 11:36:24",
      "content": "<p>I wrote a script that saves selections as ROI's To run it:</p>\n\n<ul>\n<li>save it to a file in /ImageJ/macros</li>\n<li>install the script</li>\n<li>Open an image</li>\n<li>Click a or s to move to prev/next image. This is necessary to trigger the Show ROI command.</li>\n<li><p>Edit, clear, etc. as you see fit. When changing the images with a or s, the ROI will be automatically saved.</p>\n\n<p><a href=\"http://pastebin.com/s7dum97c\">http://pastebin.com/s7dum97c</a>  because code formatting fails a bit right now.</p></li>\n</ul>\n\n<p>Hope it helps!</p>\n\n<p>p.s. python code to read the ROIs:  <a href=\"https://pypi.python.org/pypi/read-roi\">https://pypi.python.org/pypi/read-roi</a></p>\n\n<p>p.p.s. I think I will select the whole cervix not only the os.</p>",
      "rawMarkdown": "I wrote a script that saves selections as ROI's To run it:\n\n - save it to a file in /ImageJ/macros\n - install the script\n - Open an image\n - Click a or s to move to prev/next image. This is necessary to trigger the Show ROI command.\n - Edit, clear, etc. as you see fit. When changing the images with a or s, the ROI will be automatically saved.\n\n [http://pastebin.com/s7dum97c](http://pastebin.com/s7dum97c)  because code formatting fails a bit right now.\n\nHope it helps!\n\np.s. python code to read the ROIs:  https://pypi.python.org/pypi/read-roi\n\np.p.s. I think I will select the whole cervix not only the os.",
      "votes": null
    },
    {
      "id": "170043",
      "postDate": "03/23/2017 19:56:19",
      "content": "<p>Thanks for the code and link to roi reader in python! For a previous project I spent several frustrated hours trying to figure out how to get roi info into python, and ended just using save xy coordinates to get around it.</p>",
      "rawMarkdown": "Thanks for the code and link to roi reader in python! For a previous project I spent several frustrated hours trying to figure out how to get roi info into python, and ended just using save xy coordinates to get around it.",
      "votes": null
    },
    {
      "id": "171835",
      "postDate": "03/31/2017 16:22:08",
      "content": "<p>Hi guys. First off, very informative discussion and special thanks to @visoft for the informative tutorial! </p>\n\n<p>For those hand-labeling the data-set, are you planning to share the results? I will undertake this task myself as well, but am wondering if it is worth doing others will be distributing a marked-up version (like what happened with the whale competition). I apologize if that question is out of line, I am more curious than anything.</p>",
      "rawMarkdown": "Hi guys. First off, very informative discussion and special thanks to @visoft for the informative tutorial! \n\nFor those hand-labeling the data-set, are you planning to share the results? I will undertake this task myself as well, but am wondering if it is worth doing others will be distributing a marked-up version (like what happened with the whale competition). I apologize if that question is out of line, I am more curious than anything.",
      "votes": null
    },
    {
      "id": "2585254",
      "postDate": "01/03/2024 13:43:23",
      "content": "<p>thank you that was helpful </p>",
      "rawMarkdown": "thank you that was helpful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2585254,
      "author_name": "meryemelfatimi",
      "author_url": "",
      "post_date": "01/03/2024 13:43:23",
      "content": "<p>thank you that was helpful </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 169639,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "03/21/2017 21:18:51",
      "content": "<p>I labeled 400 train set images with bounding box and then trained detector networks on them. Detectors worked well, but cv score was not immediatly better for classification and have not got around to generating submission to see if helps with lb score.  </p>\n\n<p>When doing bounding box, or pixel labeling it was not clear to me how small to make the bounding box, so my initial work may not have been best way.  Thanks for detailed post will help with better bounding box labeling. </p>\n\n<p>edit: I guess an important question is how far from the os does the relevent data extend. It may be easier to detect cervix then os alone as there is quite a bit of variation between images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 169739,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "03/22/2017 11:29:23",
          "content": "<p>hi, is this generally \"ok\" / not against the rules? I mean I always assumed that all steps which lead to the result must be done \"automatically\" - i.e. no manual labelling of test images (which is obvious) but I also thought no manual classification / segmentation of the train data as preparation is allowed, was I wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169747,
          "author_name": "wolhow123",
          "author_url": "",
          "post_date": "03/22/2017 12:18:16",
          "content": "<p>You can do everything with the train data :) \nWhen you train your algorithms on manually annotated data, they still automatically make predictions on test date.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169784,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/22/2017 15:49:50",
          "content": "<p>Nice @Devin! I am planning to do that myself when I will have some quality hours to put in.</p>\n\n<p>I have no idea how large the bbox should be. Maybe the entire cervix? Just to make sure that you scale everything to the same proportion. Otherwise is tricky to take the same percent every time.</p>\n\n<p>Indeed, cervix is more stable than the os. And is roughly in the middle (given or take some projections) </p>\n\n<p>@Wolhow123 true, you can do whatever you what with train data. This is also true with test data as long as is <strong>automated</strong>. Of course, nobody is checking this for non top 3 models. (aka you might get in top 5 manually labeling everything)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169826,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/22/2017 20:28:45",
          "content": "<p>@Devin what are you using for annotations? I used <a href=\"https://github.com/cvhciKIT/sloth\">Sloth</a> a lot but is a bit sluggish and not as productive as I would like it to be.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169831,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "03/22/2017 20:55:12",
          "content": "<p>I am currently using imagej(fiji), with a macro to save the coordinates of bounding box with single keypress. I can load the next image with one command(<code>shift-o</code>), and then save with one keypress. An easy improvement would be to autoload the next image after saving the previous. This is pretty fast method though for bounding box.  You have to install the macro at the start of the session to get it to work. </p>\n\n<pre><code>macro \"Macro 1 [g]\" {\nname = getInfo(\"image.filename\")\nsaveAs(\"XY Coordinates\", \"/path/where/you/save/annotations/\"+name+\".txt\");\n} \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169919,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/23/2017 08:18:24",
          "content": "<p>Hehe! My good old friend forgot in the drawer! I had some nice scripts once for imageJ.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169943,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/23/2017 11:36:24",
          "content": "<p>I wrote a script that saves selections as ROI's To run it:</p>\n\n<ul>\n<li>save it to a file in /ImageJ/macros</li>\n<li>install the script</li>\n<li>Open an image</li>\n<li>Click a or s to move to prev/next image. This is necessary to trigger the Show ROI command.</li>\n<li><p>Edit, clear, etc. as you see fit. When changing the images with a or s, the ROI will be automatically saved.</p>\n\n<p><a href=\"http://pastebin.com/s7dum97c\">http://pastebin.com/s7dum97c</a>  because code formatting fails a bit right now.</p></li>\n</ul>\n\n<p>Hope it helps!</p>\n\n<p>p.s. python code to read the ROIs:  <a href=\"https://pypi.python.org/pypi/read-roi\">https://pypi.python.org/pypi/read-roi</a></p>\n\n<p>p.p.s. I think I will select the whole cervix not only the os.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 170043,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "03/23/2017 19:56:19",
          "content": "<p>Thanks for the code and link to roi reader in python! For a previous project I spent several frustrated hours trying to figure out how to get roi info into python, and ended just using save xy coordinates to get around it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 171835,
          "author_name": "gkericks",
          "author_url": "",
          "post_date": "03/31/2017 16:22:08",
          "content": "<p>Hi guys. First off, very informative discussion and special thanks to @visoft for the informative tutorial! </p>\n\n<p>For those hand-labeling the data-set, are you planning to share the results? I will undertake this task myself as well, but am wondering if it is worth doing others will be distributing a marked-up version (like what happened with the whale competition). I apologize if that question is out of line, I am more curious than anything.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 169640,
      "author_name": "craigglastonbury",
      "author_url": "",
      "post_date": "03/21/2017 21:19:30",
      "content": "<p>Great, thanks! Do you know whether the data is correctly labelled in most cases? Additional data seems to add noise to model training - Part me was thinking that the additional images may have many more wrong labels ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 169651,
      "author_name": "jamest",
      "author_url": "",
      "post_date": "03/21/2017 23:05:32",
      "content": "<p>Thanks a bunch.  I was going to bother my nurse sister about this, but now can avoid this...</p>\n\n<p>I would be curious what the human rate of agreement is on these categories is -- the distinction between 1 and 2 is not always clear to me, and I don't know whether this is really more picking a variable on a scale than a classification task.</p>",
      "votes": null,
      "replies": [
        {
          "id": 169825,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/22/2017 20:26:06",
          "content": "<p>Imho a classification task. And you can bother your sister with corner cases and probably for some more insights!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 169738,
      "author_name": "steelrose",
      "author_url": "",
      "post_date": "03/22/2017 11:25:30",
      "content": "<p>thanks for the clear explanation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "169634": "Hi! @Jay Patravali asked how to manually determine which is which ( see this thread: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30401 )\n\nI added a short answer, but now, I will answer with pictures! In full color! \n\nDisclaimer: 1) I have some medical training but it was long ago. 2) The images are **highly graphic**, you might lose your sex drive for a while . . .\n\nLet's begin. The most visible feature that helps us distinguish is the endocervical skin. It is red (one layer, very thin, you can see the blood and muscles) and it is in high contrast with the pinkish aspect of the rest of the cervix. \n\nExample of a nice Type 1 ![classical_type_1](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6156/classical_type_1.jpg)\n\nType 3 has all this tissue inside the uterus so only pink skin is visible from outside ![classical_type_3.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6157/classical_type_3.jpg)\n\nType 2 is a mix between them. So, part of the endocervical red skin is visible **BUT** also some of the pink skin is inside. This is a key difference between I and II. And imho the hardest thing to detect in this competition. ![easy_type_2.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6158/easy_type_2.jpg)\n\nNow, guess why this is type 2 and not 1? ![type2_like_type1_01.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6160/type2_like_type1_01.jpg) All unilayer skin is visible. Maybe the os is dilated and you can see inside the cervix tunnel? (normally you can't)\n\n And this one ![type2_like_type3_01.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6159/type2_like_type3_01.jpg) looks like type 3. Imho there is a subtle shift in color from one region to another. So even if the contrast is not good, there is \"red\" skin visible outside. So type 2. And for the classical Type 1 the red tissue looks like it gets out of the os rims. But is just a guess. \n\nAnd now let's ramp up the fun! Blurry images, blood, pathological conditions, IUDs, you name it! Check: ![nasty_examples.jpg](https://kaggle2.blob.core.windows.net/forum-message-attachments/169634/6161/nasty_examples.jpg)\n\nI think this will be a 2 stage approach. First, detect the os/cervix and then classify. To my knowledge, a similar competition [2 stage (detect,classify) + a lot of grief] is [noaa-right-whale-recognition](https://www.kaggle.com/c/noaa-right-whale-recognition). \n\nSo, anybody in the mood to paint the Os on ~1400 images? Or maybe it is easier to mark clear cases on 512 images and score some leaderboard points ;)\n\nHope it helps!\nCristi\n\np.s. If you know how to \"insert\" the pictures in the text, without having the urls, let me know. This new Kaggle UI beats me. \n\nedit: hacked! 2 step process ;)",
    "169639": "I labeled 400 train set images with bounding box and then trained detector networks on them. Detectors worked well, but cv score was not immediatly better for classification and have not got around to generating submission to see if helps with lb score.  \n\nWhen doing bounding box, or pixel labeling it was not clear to me how small to make the bounding box, so my initial work may not have been best way.  Thanks for detailed post will help with better bounding box labeling. \n\nedit: I guess an important question is how far from the os does the relevent data extend. It may be easier to detect cervix then os alone as there is quite a bit of variation between images.",
    "169640": "Great, thanks! Do you know whether the data is correctly labelled in most cases? Additional data seems to add noise to model training - Part me was thinking that the additional images may have many more wrong labels ?",
    "169651": "Thanks a bunch.  I was going to bother my nurse sister about this, but now can avoid this...\n\nI would be curious what the human rate of agreement is on these categories is -- the distinction between 1 and 2 is not always clear to me, and I don't know whether this is really more picking a variable on a scale than a classification task.",
    "169738": "thanks for the clear explanation",
    "169739": "hi, is this generally \"ok\" / not against the rules? I mean I always assumed that all steps which lead to the result must be done \"automatically\" - i.e. no manual labelling of test images (which is obvious) but I also thought no manual classification / segmentation of the train data as preparation is allowed, was I wrong?",
    "169747": "You can do everything with the train data :) \nWhen you train your algorithms on manually annotated data, they still automatically make predictions on test date.",
    "169784": "Nice @Devin! I am planning to do that myself when I will have some quality hours to put in.\n\nI have no idea how large the bbox should be. Maybe the entire cervix? Just to make sure that you scale everything to the same proportion. Otherwise is tricky to take the same percent every time.\n\nIndeed, cervix is more stable than the os. And is roughly in the middle (given or take some projections) \n\n@Wolhow123 true, you can do whatever you what with train data. This is also true with test data as long as is **automated**. Of course, nobody is checking this for non top 3 models. (aka you might get in top 5 manually labeling everything)",
    "169825": "Imho a classification task. And you can bother your sister with corner cases and probably for some more insights!",
    "169826": "Devin what are you using for annotations? I used [Sloth](https://github.com/cvhciKIT/sloth) a lot but is a bit sluggish and not as productive as I would like it to be.",
    "169831": "I am currently using imagej(fiji), with a macro to save the coordinates of bounding box with single keypress. I can load the next image with one command(`shift-o`), and then save with one keypress. An easy improvement would be to autoload the next image after saving the previous. This is pretty fast method though for bounding box.  You have to install the macro at the start of the session to get it to work. \n\n    macro \"Macro 1 [g]\" {\n\tname = getInfo(\"image.filename\")\n\tsaveAs(\"XY Coordinates\", \"/path/where/you/save/annotations/\"+name+\".txt\");\n    }",
    "169919": "Hehe! My good old friend forgot in the drawer! I had some nice scripts once for imageJ.",
    "169943": "I wrote a script that saves selections as ROI's To run it:\n\n - save it to a file in /ImageJ/macros\n - install the script\n - Open an image\n - Click a or s to move to prev/next image. This is necessary to trigger the Show ROI command.\n - Edit, clear, etc. as you see fit. When changing the images with a or s, the ROI will be automatically saved.\n\n [http://pastebin.com/s7dum97c](http://pastebin.com/s7dum97c)  because code formatting fails a bit right now.\n\nHope it helps!\n\np.s. python code to read the ROIs:  https://pypi.python.org/pypi/read-roi\n\np.p.s. I think I will select the whole cervix not only the os.",
    "170043": "Thanks for the code and link to roi reader in python! For a previous project I spent several frustrated hours trying to figure out how to get roi info into python, and ended just using save xy coordinates to get around it.",
    "171835": "Hi guys. First off, very informative discussion and special thanks to @visoft for the informative tutorial! \n\nFor those hand-labeling the data-set, are you planning to share the results? I will undertake this task myself as well, but am wondering if it is worth doing others will be distributing a marked-up version (like what happened with the whale competition). I apologize if that question is out of line, I am more curious than anything.",
    "2585254": "thank you that was helpful"
  },
  "source": "meta"
}