{
  "id": 35269,
  "title": "27th place solution: Polar Coordinates",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/entanglement-27th-place-solution-polar-coordinates",
  "author_name": "",
  "post_date": "2017-07-18T07:56:59.470Z",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Actually placed at 26th after LB update. \nIdea was to exploit the approximate radial symmetry in the problem to a number of advantages that I would describe below, of course, with some definite LB gain! (Any other physicists in the room!?) Polar coordinates seem to be a \"natural\" frame for the features that were important to the cervix type detection (transformation zone being a ring).</p>\n\n<p>I think there is a lot more scope for improvement here as I very hastily implemented this in the last few days of the competition. Posting this in case it turns out helpful for MobileODT in some way.</p>\n\n<h2>First, a few Basic transformations applied on the Images:</h2>\n\n<p>$ Since images could be green filtered images, and it seemed color didn't add any information, I had replicated the green channel for R, G, B (so that I can make use of the pretrained models), instead of using grayscale values for R, G, B as greyscale values of a green filtered test images would be from a different distribution then, as average = (g + 0 + 0)/3 in that case, making the mean 3 times smaller than expected distribution. </p>\n\n<p>$ These G, G, G channel images were then resized to a maximum dimension of 598x598 with aspect ratio preserved.</p>\n\n<p>$ Then a center crop of 448x448 (75% of 598) was extracted without any resizing.\nThese were used as the base images before any more processing.</p>\n\n<p>Then trying out a bunch of pretrained finetuned models (Vgg16BN, InceptionV3, ResNet, Xception) on these base \"cartesian\" images gave a best single model of around 0.69-0.71 PBLB at best, which gave a fairly good score after weighted ensembling. For tuning these weights of the ensemble, I had used a separate hold-out set of data. At this point adding any other model to the ensemble didn't add much to the final score.</p>\n\n<h2>Polar Coordinates</h2>\n\n<p>A polar transform is a map (not one to one, actually) that takes (x,y) from Cartesian coordinates to (r, theta) in polar coordinates as follows:</p>\n\n<p>$$ (x,y) \\to (r,\\theta)=\\left(\\sqrt{x^2 + y^2}, ~\\text{tan}^{-1}\\frac{y}{x} \\right) $$.</p>\n\n<p>Notice that this is not a one to one map (it is a one-to-many map). For example, $$(r, \\theta), ~\\text{and} ~(r,\\theta + 2n\\pi)$$ trivially maps to the same point in cartesian coordinates. Also, the center/origin in x-y coordinates maps to an infinite number of points  $$(r=0,~\\theta)$$ for any theta (basically the theta axis).</p>\n\n<p>At first I tried to implement this in a python function (fairly straightforward), but then resorted to ImageJ's plugin <a href=\"https://imagej.nih.gov/ij/plugins/polar-transformer.html\">PolarTransformer.class</a> due to its fast pixel crunching.\nAn example of this transform on a spherically symmetric image:</p>\n\n<p><img src=\"http://i.imgur.com/WCCyVab.png?raw=True\" alt=\"\" title=\"\"></p>\n\n<p>One may analyze the effect of this on cervix images by considering the below illustrative picture where the blue/green represents some background noise (the arrows are drawn to indicate the direction of the convolutional kernels over the image):\n<img src=\"http://i.imgur.com/4VEXVO2.jpg\" alt=\"\" title=\"\">\nThe black region towards the right most end of the polar image is due to the fact that the original image is rectangular shaped and not a circular one, and hence there are always some missing pixels over the circular ring with radius r &gt; r_0; where r_0 = size/2 (size being the width or height of the square image in cartesian coordinates).</p>\n\n<p>The noise can now be easily cropped and then checked for quality by reconstructing Cartesian from Polar using the inverse map which is a many-to-one map:\n$$ (r, \\theta) \\to (x, y) = (r \\cos\\theta, r\\sin\\theta) $$\nas follows:\n<img src=\"http://i.imgur.com/6Yvho2R.jpg\" alt=\"\" title=\"\"></p>\n\n<p>The advantages with this transform are two-fold:</p>\n\n<p>1 - Considering the green and blue parts above to be the black borders of the MobileODT tube or some non-cervix background, that can be now, easily eliminated by cropping out upto a limited value along the radius axis of the polar image (width) leading to less overfit models.\nExample: \nConsider train_data/Type_2/1038.jpg:\n<img src=\"http://i.imgur.com/brk4HXG.jpg\" alt=\"\" title=\"\">\n(360 here corresponds to max theta in degrees)\nIn my case I used these 216x360 cropped images for bulding models over polar data. To inspect the quality of information in the cropped polar images, one may check by reconstructing images by inverse transform:\n<img src=\"http://i.imgur.com/oU8uxRz.jpg\" alt=\"\" title=\"\">\nPretty good reconstruction! The black background of the original image can now be seen painted in the reconstructed image, which is due to those pixel maps that have multiple cartesian coordinates for the same polar coordinate (many to one). \nOne can easily find a good cutoff for such noise cropping manually as region around a cervix over a certain radius is anyhow not much important for cervix type classification (ofcourse assuming cervix centered images).</p>\n\n<p>2 - Applying CNNs on polar images makes the square kernels \"compatible\" with the transformation zone ring. A kernel running over Cartesian images is represented by the green arrow above, while a kernel running over the Polar ones is represented by the yellow arrow.</p>\n\n<p>But when pretrained models were built on these polar images alone, single best model was much worse than cartesian -  standing at around 0.78. But I was atleast sure this was not as much overfit.</p>\n\n<p>To get the best features out of the cartesian and polar (216x360), finally, I resorted to using the flattened last convolution layers from two separately finetuned resnet models on cartesian and polar, and feeding them to xgboost. \nAdding this model to the ensemble gave me about a 0.07 boost on the private LB.</p>\n\n<p>The xgboost used was not cross-validated across various parameters as the competition was nearing end. There could be a scope for improvement here. \nAdditionally, I didn't use any detect framework to get the cervix-centered crops, instead, I assumed images were more or less centered anyway and applied the polar transform. This is another possible area of improvement. Also, one may experiment with the reconstructed cartesian images from cropped polar because the pitch black background generally kills the kernels and overfits.</p>\n\n<p>Let me know if you find these features helpful. Thanks!</p>",
  "messages": [
    {
      "id": "195921",
      "postDate": "06/25/2017 19:04:24",
      "content": "<p>Actually placed at 26th after LB update. \nIdea was to exploit the approximate radial symmetry in the problem to a number of advantages that I would describe below, of course, with some definite LB gain! (Any other physicists in the room!?) Polar coordinates seem to be a \"natural\" frame for the features that were important to the cervix type detection (transformation zone being a ring).</p>\n\n<p>I think there is a lot more scope for improvement here as I very hastily implemented this in the last few days of the competition. Posting this in case it turns out helpful for MobileODT in some way.</p>\n\n<h2>First, a few Basic transformations applied on the Images:</h2>\n\n<p>$ Since images could be green filtered images, and it seemed color didn't add any information, I had replicated the green channel for R, G, B (so that I can make use of the pretrained models), instead of using grayscale values for R, G, B as greyscale values of a green filtered test images would be from a different distribution then, as average = (g + 0 + 0)/3 in that case, making the mean 3 times smaller than expected distribution. </p>\n\n<p>$ These G, G, G channel images were then resized to a maximum dimension of 598x598 with aspect ratio preserved.</p>\n\n<p>$ Then a center crop of 448x448 (75% of 598) was extracted without any resizing.\nThese were used as the base images before any more processing.</p>\n\n<p>Then trying out a bunch of pretrained finetuned models (Vgg16BN, InceptionV3, ResNet, Xception) on these base \"cartesian\" images gave a best single model of around 0.69-0.71 PBLB at best, which gave a fairly good score after weighted ensembling. For tuning these weights of the ensemble, I had used a separate hold-out set of data. At this point adding any other model to the ensemble didn't add much to the final score.</p>\n\n<h2>Polar Coordinates</h2>\n\n<p>A polar transform is a map (not one to one, actually) that takes (x,y) from Cartesian coordinates to (r, theta) in polar coordinates as follows:</p>\n\n<p>$$ (x,y) \\to (r,\\theta)=\\left(\\sqrt{x^2 + y^2}, ~\\text{tan}^{-1}\\frac{y}{x} \\right) $$.</p>\n\n<p>Notice that this is not a one to one map (it is a one-to-many map). For example, $$(r, \\theta), ~\\text{and} ~(r,\\theta + 2n\\pi)$$ trivially maps to the same point in cartesian coordinates. Also, the center/origin in x-y coordinates maps to an infinite number of points  $$(r=0,~\\theta)$$ for any theta (basically the theta axis).</p>\n\n<p>At first I tried to implement this in a python function (fairly straightforward), but then resorted to ImageJ's plugin <a href=\"https://imagej.nih.gov/ij/plugins/polar-transformer.html\">PolarTransformer.class</a> due to its fast pixel crunching.\nAn example of this transform on a spherically symmetric image:</p>\n\n<p><img src=\"http://i.imgur.com/WCCyVab.png?raw=True\" alt=\"\" title=\"\"></p>\n\n<p>One may analyze the effect of this on cervix images by considering the below illustrative picture where the blue/green represents some background noise (the arrows are drawn to indicate the direction of the convolutional kernels over the image):\n<img src=\"http://i.imgur.com/4VEXVO2.jpg\" alt=\"\" title=\"\">\nThe black region towards the right most end of the polar image is due to the fact that the original image is rectangular shaped and not a circular one, and hence there are always some missing pixels over the circular ring with radius r &gt; r_0; where r_0 = size/2 (size being the width or height of the square image in cartesian coordinates).</p>\n\n<p>The noise can now be easily cropped and then checked for quality by reconstructing Cartesian from Polar using the inverse map which is a many-to-one map:\n$$ (r, \\theta) \\to (x, y) = (r \\cos\\theta, r\\sin\\theta) $$\nas follows:\n<img src=\"http://i.imgur.com/6Yvho2R.jpg\" alt=\"\" title=\"\"></p>\n\n<p>The advantages with this transform are two-fold:</p>\n\n<p>1 - Considering the green and blue parts above to be the black borders of the MobileODT tube or some non-cervix background, that can be now, easily eliminated by cropping out upto a limited value along the radius axis of the polar image (width) leading to less overfit models.\nExample: \nConsider train_data/Type_2/1038.jpg:\n<img src=\"http://i.imgur.com/brk4HXG.jpg\" alt=\"\" title=\"\">\n(360 here corresponds to max theta in degrees)\nIn my case I used these 216x360 cropped images for bulding models over polar data. To inspect the quality of information in the cropped polar images, one may check by reconstructing images by inverse transform:\n<img src=\"http://i.imgur.com/oU8uxRz.jpg\" alt=\"\" title=\"\">\nPretty good reconstruction! The black background of the original image can now be seen painted in the reconstructed image, which is due to those pixel maps that have multiple cartesian coordinates for the same polar coordinate (many to one). \nOne can easily find a good cutoff for such noise cropping manually as region around a cervix over a certain radius is anyhow not much important for cervix type classification (ofcourse assuming cervix centered images).</p>\n\n<p>2 - Applying CNNs on polar images makes the square kernels \"compatible\" with the transformation zone ring. A kernel running over Cartesian images is represented by the green arrow above, while a kernel running over the Polar ones is represented by the yellow arrow.</p>\n\n<p>But when pretrained models were built on these polar images alone, single best model was much worse than cartesian -  standing at around 0.78. But I was atleast sure this was not as much overfit.</p>\n\n<p>To get the best features out of the cartesian and polar (216x360), finally, I resorted to using the flattened last convolution layers from two separately finetuned resnet models on cartesian and polar, and feeding them to xgboost. \nAdding this model to the ensemble gave me about a 0.07 boost on the private LB.</p>\n\n<p>The xgboost used was not cross-validated across various parameters as the competition was nearing end. There could be a scope for improvement here. \nAdditionally, I didn't use any detect framework to get the cervix-centered crops, instead, I assumed images were more or less centered anyway and applied the polar transform. This is another possible area of improvement. Also, one may experiment with the reconstructed cartesian images from cropped polar because the pitch black background generally kills the kernels and overfits.</p>\n\n<p>Let me know if you find these features helpful. Thanks!</p>",
      "rawMarkdown": "Actually placed at 26th after LB update. \nIdea was to exploit the approximate radial symmetry in the problem to a number of advantages that I would describe below, of course, with some definite LB gain! (Any other physicists in the room!?) Polar coordinates seem to be a \"natural\" frame for the features that were important to the cervix type detection (transformation zone being a ring).\n\nI think there is a lot more scope for improvement here as I very hastily implemented this in the last few days of the competition. Posting this in case it turns out helpful for MobileODT in some way.\n\n## First, a few Basic transformations applied on the Images: ##\n$ Since images could be green filtered images, and it seemed color didn't add any information, I had replicated the green channel for R, G, B (so that I can make use of the pretrained models), instead of using grayscale values for R, G, B as greyscale values of a green filtered test images would be from a different distribution then, as average = (g + 0 + 0)/3 in that case, making the mean 3 times smaller than expected distribution. \n\n$ These G, G, G channel images were then resized to a maximum dimension of 598x598 with aspect ratio preserved.\n\n$ Then a center crop of 448x448 (75% of 598) was extracted without any resizing.\nThese were used as the base images before any more processing.\n\nThen trying out a bunch of pretrained finetuned models (Vgg16BN, InceptionV3, ResNet, Xception) on these base \"cartesian\" images gave a best single model of around 0.69-0.71 PBLB at best, which gave a fairly good score after weighted ensembling. For tuning these weights of the ensemble, I had used a separate hold-out set of data. At this point adding any other model to the ensemble didn't add much to the final score.\n\n## Polar Coordinates ##\nA polar transform is a map (not one to one, actually) that takes (x,y) from Cartesian coordinates to (r, theta) in polar coordinates as follows:\n\n$$ (x,y) \\to (r,\\theta)=\\left(\\sqrt{x^2 + y^2}, ~\\text{tan}^{-1}\\frac{y}{x} \\right) $$.\n\nNotice that this is not a one to one map (it is a one-to-many map). For example, $$(r, \\theta), ~\\text{and} ~(r,\\theta + 2n\\pi)$$ trivially maps to the same point in cartesian coordinates. Also, the center/origin in x-y coordinates maps to an infinite number of points  $$(r=0,~\\theta)$$ for any theta (basically the theta axis).\n\nAt first I tried to implement this in a python function (fairly straightforward), but then resorted to ImageJ's plugin [PolarTransformer.class][1] due to its fast pixel crunching.\nAn example of this transform on a spherically symmetric image:\n\n![](http://i.imgur.com/WCCyVab.png?raw=True)\n\nOne may analyze the effect of this on cervix images by considering the below illustrative picture where the blue/green represents some background noise (the arrows are drawn to indicate the direction of the convolutional kernels over the image):\n![](http://i.imgur.com/4VEXVO2.jpg)\nThe black region towards the right most end of the polar image is due to the fact that the original image is rectangular shaped and not a circular one, and hence there are always some missing pixels over the circular ring with radius r &gt; r_0; where r_0 = size/2 (size being the width or height of the square image in cartesian coordinates).\n\nThe noise can now be easily cropped and then checked for quality by reconstructing Cartesian from Polar using the inverse map which is a many-to-one map:\n$$ (r, \\theta) \\to (x, y) = (r \\cos\\theta, r\\sin\\theta) $$\nas follows:\n![](http://i.imgur.com/6Yvho2R.jpg)\n\nThe advantages with this transform are two-fold:\n\n1 - Considering the green and blue parts above to be the black borders of the MobileODT tube or some non-cervix background, that can be now, easily eliminated by cropping out upto a limited value along the radius axis of the polar image (width) leading to less overfit models.\nExample: \nConsider train_data/Type_2/1038.jpg:\n![](http://i.imgur.com/brk4HXG.jpg)\n(360 here corresponds to max theta in degrees)\nIn my case I used these 216x360 cropped images for bulding models over polar data. To inspect the quality of information in the cropped polar images, one may check by reconstructing images by inverse transform:\n![](http://i.imgur.com/oU8uxRz.jpg)\nPretty good reconstruction! The black background of the original image can now be seen painted in the reconstructed image, which is due to those pixel maps that have multiple cartesian coordinates for the same polar coordinate (many to one). \nOne can easily find a good cutoff for such noise cropping manually as region around a cervix over a certain radius is anyhow not much important for cervix type classification (ofcourse assuming cervix centered images).\n\n2 - Applying CNNs on polar images makes the square kernels \"compatible\" with the transformation zone ring. A kernel running over Cartesian images is represented by the green arrow above, while a kernel running over the Polar ones is represented by the yellow arrow.\n\nBut when pretrained models were built on these polar images alone, single best model was much worse than cartesian -  standing at around 0.78. But I was atleast sure this was not as much overfit.\n\nTo get the best features out of the cartesian and polar (216x360), finally, I resorted to using the flattened last convolution layers from two separately finetuned resnet models on cartesian and polar, and feeding them to xgboost. \nAdding this model to the ensemble gave me about a 0.07 boost on the private LB.\n\nThe xgboost used was not cross-validated across various parameters as the competition was nearing end. There could be a scope for improvement here. \nAdditionally, I didn't use any detect framework to get the cervix-centered crops, instead, I assumed images were more or less centered anyway and applied the polar transform. This is another possible area of improvement. Also, one may experiment with the reconstructed cartesian images from cropped polar because the pitch black background generally kills the kernels and overfits.\n\nLet me know if you find these features helpful. Thanks!\n\n  [1]: https://imagej.nih.gov/ij/plugins/polar-transformer.html",
      "votes": null
    },
    {
      "id": "195925",
      "postDate": "06/25/2017 19:12:36",
      "content": "<p>Clever !</p>",
      "rawMarkdown": "Clever !",
      "votes": null
    },
    {
      "id": "196707",
      "postDate": "06/28/2017 00:59:28",
      "content": "<p>Thanks entanglement for the nice writeup and interesting approach.   Wondering if throwing out the red channel is wise given the three types appear to often be associated with the amount of visible blood.   Also, I think the percentage of green filtered images is fairly small.   Nonetheless, polar coordinates seem applicable in a variety of contexts and could prove to be useful here.   One of the upvotes is mine.</p>",
      "rawMarkdown": "Thanks entanglement for the nice writeup and interesting approach.   Wondering if throwing out the red channel is wise given the three types appear to often be associated with the amount of visible blood.   Also, I think the percentage of green filtered images is fairly small.   Nonetheless, polar coordinates seem applicable in a variety of contexts and could prove to be useful here.   One of the upvotes is mine.",
      "votes": null
    },
    {
      "id": "196935",
      "postDate": "06/28/2017 12:37:06",
      "content": "<p>How did you use InceptionV3 for Type 1,2,3 classification ?Is it not pre trained for 1000 classes as mentioned <a href=\"https://www.tensorflow.org/tutorials/image_recognition\">here</a>.\nI am still trying to understand the application of these.</p>\n\n<p>This is a very clever way! Congrats!</p>",
      "rawMarkdown": "How did you use InceptionV3 for Type 1,2,3 classification ?Is it not pre trained for 1000 classes as mentioned [here][1].\nI am still trying to understand the application of these.\n\nThis is a very clever way! Congrats!\n  [1]: https://www.tensorflow.org/tutorials/image_recognition",
      "votes": null
    },
    {
      "id": "196948",
      "postDate": "06/28/2017 13:15:37",
      "content": "<p>Just look up what's called as \"transfer learning\", where you pluck off the final 1000 way layer (and maybe a few other layers above), and train the network with your new 3-way softmax.</p>",
      "rawMarkdown": "Just look up what's called as \"transfer learning\", where you pluck off the final 1000 way layer (and maybe a few other layers above), and train the network with your new 3-way softmax.",
      "votes": null
    },
    {
      "id": "198359",
      "postDate": "07/02/2017 03:08:04",
      "content": "<p>But isn't the amount of visible blood information also present as black regions in the G channel?</p>",
      "rawMarkdown": "But isn't the amount of visible blood information also present as black regions in the G channel?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 195925,
      "author_name": "carlosaguayo",
      "author_url": "",
      "post_date": "06/25/2017 19:12:36",
      "content": "<p>Clever !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 196707,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "06/28/2017 00:59:28",
      "content": "<p>Thanks entanglement for the nice writeup and interesting approach.   Wondering if throwing out the red channel is wise given the three types appear to often be associated with the amount of visible blood.   Also, I think the percentage of green filtered images is fairly small.   Nonetheless, polar coordinates seem applicable in a variety of contexts and could prove to be useful here.   One of the upvotes is mine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 198359,
          "author_name": "bvineeth007",
          "author_url": "",
          "post_date": "07/02/2017 03:08:04",
          "content": "<p>But isn't the amount of visible blood information also present as black regions in the G channel?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196935,
      "author_name": "gandalf1292",
      "author_url": "",
      "post_date": "06/28/2017 12:37:06",
      "content": "<p>How did you use InceptionV3 for Type 1,2,3 classification ?Is it not pre trained for 1000 classes as mentioned <a href=\"https://www.tensorflow.org/tutorials/image_recognition\">here</a>.\nI am still trying to understand the application of these.</p>\n\n<p>This is a very clever way! Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 196948,
          "author_name": "bvineeth007",
          "author_url": "",
          "post_date": "06/28/2017 13:15:37",
          "content": "<p>Just look up what's called as \"transfer learning\", where you pluck off the final 1000 way layer (and maybe a few other layers above), and train the network with your new 3-way softmax.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "195921": "Actually placed at 26th after LB update. \nIdea was to exploit the approximate radial symmetry in the problem to a number of advantages that I would describe below, of course, with some definite LB gain! (Any other physicists in the room!?) Polar coordinates seem to be a \"natural\" frame for the features that were important to the cervix type detection (transformation zone being a ring).\n\nI think there is a lot more scope for improvement here as I very hastily implemented this in the last few days of the competition. Posting this in case it turns out helpful for MobileODT in some way.\n\n## First, a few Basic transformations applied on the Images: ##\n$ Since images could be green filtered images, and it seemed color didn't add any information, I had replicated the green channel for R, G, B (so that I can make use of the pretrained models), instead of using grayscale values for R, G, B as greyscale values of a green filtered test images would be from a different distribution then, as average = (g + 0 + 0)/3 in that case, making the mean 3 times smaller than expected distribution. \n\n$ These G, G, G channel images were then resized to a maximum dimension of 598x598 with aspect ratio preserved.\n\n$ Then a center crop of 448x448 (75% of 598) was extracted without any resizing.\nThese were used as the base images before any more processing.\n\nThen trying out a bunch of pretrained finetuned models (Vgg16BN, InceptionV3, ResNet, Xception) on these base \"cartesian\" images gave a best single model of around 0.69-0.71 PBLB at best, which gave a fairly good score after weighted ensembling. For tuning these weights of the ensemble, I had used a separate hold-out set of data. At this point adding any other model to the ensemble didn't add much to the final score.\n\n## Polar Coordinates ##\nA polar transform is a map (not one to one, actually) that takes (x,y) from Cartesian coordinates to (r, theta) in polar coordinates as follows:\n\n$$ (x,y) \\to (r,\\theta)=\\left(\\sqrt{x^2 + y^2}, ~\\text{tan}^{-1}\\frac{y}{x} \\right) $$.\n\nNotice that this is not a one to one map (it is a one-to-many map). For example, $$(r, \\theta), ~\\text{and} ~(r,\\theta + 2n\\pi)$$ trivially maps to the same point in cartesian coordinates. Also, the center/origin in x-y coordinates maps to an infinite number of points  $$(r=0,~\\theta)$$ for any theta (basically the theta axis).\n\nAt first I tried to implement this in a python function (fairly straightforward), but then resorted to ImageJ's plugin [PolarTransformer.class][1] due to its fast pixel crunching.\nAn example of this transform on a spherically symmetric image:\n\n![](http://i.imgur.com/WCCyVab.png?raw=True)\n\nOne may analyze the effect of this on cervix images by considering the below illustrative picture where the blue/green represents some background noise (the arrows are drawn to indicate the direction of the convolutional kernels over the image):\n![](http://i.imgur.com/4VEXVO2.jpg)\nThe black region towards the right most end of the polar image is due to the fact that the original image is rectangular shaped and not a circular one, and hence there are always some missing pixels over the circular ring with radius r &gt; r_0; where r_0 = size/2 (size being the width or height of the square image in cartesian coordinates).\n\nThe noise can now be easily cropped and then checked for quality by reconstructing Cartesian from Polar using the inverse map which is a many-to-one map:\n$$ (r, \\theta) \\to (x, y) = (r \\cos\\theta, r\\sin\\theta) $$\nas follows:\n![](http://i.imgur.com/6Yvho2R.jpg)\n\nThe advantages with this transform are two-fold:\n\n1 - Considering the green and blue parts above to be the black borders of the MobileODT tube or some non-cervix background, that can be now, easily eliminated by cropping out upto a limited value along the radius axis of the polar image (width) leading to less overfit models.\nExample: \nConsider train_data/Type_2/1038.jpg:\n![](http://i.imgur.com/brk4HXG.jpg)\n(360 here corresponds to max theta in degrees)\nIn my case I used these 216x360 cropped images for bulding models over polar data. To inspect the quality of information in the cropped polar images, one may check by reconstructing images by inverse transform:\n![](http://i.imgur.com/oU8uxRz.jpg)\nPretty good reconstruction! The black background of the original image can now be seen painted in the reconstructed image, which is due to those pixel maps that have multiple cartesian coordinates for the same polar coordinate (many to one). \nOne can easily find a good cutoff for such noise cropping manually as region around a cervix over a certain radius is anyhow not much important for cervix type classification (ofcourse assuming cervix centered images).\n\n2 - Applying CNNs on polar images makes the square kernels \"compatible\" with the transformation zone ring. A kernel running over Cartesian images is represented by the green arrow above, while a kernel running over the Polar ones is represented by the yellow arrow.\n\nBut when pretrained models were built on these polar images alone, single best model was much worse than cartesian -  standing at around 0.78. But I was atleast sure this was not as much overfit.\n\nTo get the best features out of the cartesian and polar (216x360), finally, I resorted to using the flattened last convolution layers from two separately finetuned resnet models on cartesian and polar, and feeding them to xgboost. \nAdding this model to the ensemble gave me about a 0.07 boost on the private LB.\n\nThe xgboost used was not cross-validated across various parameters as the competition was nearing end. There could be a scope for improvement here. \nAdditionally, I didn't use any detect framework to get the cervix-centered crops, instead, I assumed images were more or less centered anyway and applied the polar transform. This is another possible area of improvement. Also, one may experiment with the reconstructed cartesian images from cropped polar because the pitch black background generally kills the kernels and overfits.\n\nLet me know if you find these features helpful. Thanks!\n\n  [1]: https://imagej.nih.gov/ij/plugins/polar-transformer.html",
    "195925": "Clever !",
    "196707": "Thanks entanglement for the nice writeup and interesting approach.   Wondering if throwing out the red channel is wise given the three types appear to often be associated with the amount of visible blood.   Also, I think the percentage of green filtered images is fairly small.   Nonetheless, polar coordinates seem applicable in a variety of contexts and could prove to be useful here.   One of the upvotes is mine.",
    "196935": "How did you use InceptionV3 for Type 1,2,3 classification ?Is it not pre trained for 1000 classes as mentioned [here][1].\nI am still trying to understand the application of these.\n\nThis is a very clever way! Congrats!\n  [1]: https://www.tensorflow.org/tutorials/image_recognition",
    "196948": "Just look up what's called as \"transfer learning\", where you pluck off the final 1000 way layer (and maybe a few other layers above), and train the network with your new 3-way softmax.",
    "198359": "But isn't the amount of visible blood information also present as black regions in the G channel?"
  },
  "source": "meta"
}