{
  "id": 68431,
  "title": "A humble introduction to the art of medicine & machine learning",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/68431",
  "author_name": "",
  "post_date": "2018-10-12T15:56:48.537999400Z",
  "votes": 16,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Good morning, evening and night!</p>\n\n<p><strong>A tribute to the organizers</strong>\nJust an amazing competition. I've seen people from all sorts of backgrounds coming together to learn and compete. Hats off. </p>\n\n<p><strong>What is this thread about?</strong>\nNot everyone here is an expert in medicine and/or machine learning/statistics etc. so I figured for people trying to learn or improve their models I would do my best to offer some guidance and tips. For me personally, Kaggle is all about people coming together and learning new, intriguing things. I realise I write this a bit late into the competition but nevertheless, here we are. </p>\n\n<p><strong>Why should anyone even care about what you have to say?</strong>\nThis is most certainly not a definite guide for success just some general and specific things that I think newcomers (and more advanced learners) could find useful. I think there are more questions than answers in this thread, a good way to learn. Also, see the Kaggle getting started thread \"How to win Kaggle Competitions\" for Zeeshan-ul-hassan Usmans general input on winning competitions. </p>\n\n<p><strong>Who are you?</strong>\nI am <strong>NOT</strong> an expert in all fields mentioned above. <strong>I am a 24 years old MD/PhD student.</strong> I have almost finished my medical studies (MD) and I am also completing my PhD studies in medical data analysis, a rather fitting background aligned with my interests in data-driven healthcare. This is my first time posting on Kaggle, sadly, I do not have the time to engage in this competition and therefore hope to contribute with this.</p>\n\n<p><em>Let's get started.</em></p>\n\n<p><strong>1. Understand the problem you are trying to solve</strong></p>\n\n<p>You might think that this is an obvious statement, but the complexity is not to be underestimated here. <em>Please allow me to expand on this,</em> \nWhat are we trying to predict in this competition? Are we trying to predict whether someone has pneumonia or not? Or are we trying to predict one or more \"bounding boxes\" in a given image? What is the difference between these problems?</p>\n\n<p>\"Ground-truth bounding boxes\" (some prefer \"reference bounding boxes\"), manually drawn boxes on an image by radiologists highlighting opacities or other suspect changes in the chest x-rays. (see Guy Zahavis excellent kernel). </p>\n\n<p>If one really wants to learn, I think for some questions it might be better to search for the answers yourselves, I believe that asking the right question is harder than finding the answer in many cases. I will offer some guidance for anyone who wants a little push.\nWe are trying to predict bounding boxes in these images, and we are also trying to predict if someone has opacities (indicating pneumonia) or not. There is an evident difference between these problems, one of them we can call \"object detection\" and the other \"image classification\". So, as a newcomer I would approach this very simply by understanding what models others have used for object detection tasks in the past and do the same for image classification (more on this later).</p>\n\n<p>From a competition standpoint the evaluation metric is crucial to understand, I would honestly spend hours and hours trying to understand just the metric (Intersect over Union) if I wanted to improve my score. Some key things that I would look for would be, first of all what is this mysterious IoU thing? How is it affected by false prediction, in the cases my model predicts say a bounding box where there shouldn't be one? Is it maybe worse to make a false positive prediction, than no prediction at all when there should be one? etc. This is important, because you can build your model accordingly and increase your score. Does it matter whether your model correctly predicts if someone has pneumonia or not? How is this evaluated and relatively, how important is this compared to predicting the bounding boxes? That is, what should I focus my time and effort on? </p>\n\n<p>There are many more things to consider before we even get started with trying to solve the problem but I don't want to write a book that no one will end up reading so lets move on. Just remember, you can endlessly and blindly test your way forward or you can really take the time to understand and formulate the problem, which I guarantee will not only lead to a better score but also to a deeper understanding. </p>\n\n<p><strong>2. Understand all the available data</strong></p>\n\n<p>Let's say that we have formulated the problem and now have a good grasp of the evaluation metric. A logical next step is trying to grasp what kind of data we are working with, I am not talking about any statistics yet. I literally mean, what features/variables are available to us? What do they mean in the relevant medical context? A good example is the AP-PA discussions that I briefly saw some people discussing, in which AP (Anterio-Posterior) and PA (Posterio-Anterior) represent different values in the available dicom-meta data indicating whether a patient was lying down during the scan or standing up. Some smart people quickly realised the meaning of this feature from a medical perspective, where a patient who was lying down for the chest-x-ray most likely is more weakened than a patient who could stand up, logically inferring that such a patient is more likely to have opacities.</p>\n\n<p>Another important part is to understand the weaknesses of the data that you are training on comparatively to the data you will be testing on. One such event with this data is perhaps, will the same radiologists that labelled the training images also label the testing images? What is the inter-individual (between persons) variation? </p>\n\n<p>A side note is to explore whether you can expand or improve your data set i.e. better or additional labels with external sources etc. (if allowed). </p>\n\n<p><strong>3. Explore the data [EDA = Exploratory data analysis]</strong></p>\n\n<p>Cool, now we have a problem to solve and we know what data we are working with. You will notice that a large part of what I am discussing in this thread has nothing to with complex mathematics and advanced machine learning network architecture. This is because I believe that the fundamentals are equally important as all the complexities of building a model. I think a lot of people would agree with me when I say that the EDA is extremely important to the outcome of your models. The EDA consists of many parts, only limited to ones imagination, time and knowledge. I will highlight some parts that I think are important and underestimated. </p>\n\n<p>Do not skip the simple stuff like checking distribution, outliers and simple correlations. It might be boring, and one can probably argue that it won't affect the outcome. I personally believe that it is a good habit to explore these simple things to better understand your data regardless. A big part of why I believe that these things are important is because the deeper understanding of the features both context-based and statistically could be helpful for you to develop new features, and as many large names in the field of machine learning has stated i.e. Andrew Ng, feature engineering is a very important part of applied machine learning.</p>\n\n<p>\"Feature engineering\" simply means creating new features by combining features or looking at existing data from a new perspective, one such example is to look at the grey-scale histograms of the images or simple stuff like creating a feature \"bounding box area\". ENDLESS OPPORTUNITIES and smarter people than me are definitely doing this in ways that I probably wouldn't even understand in this competition. </p>\n\n<p>A more statistically \"complex\" part of EDA could be to build multi-variate models and try to predict based on a patients' age, gender, view position and whatever features you believe to be important if an image has a bounding box, or try to predict the designated class (i.e. pneumonia or not), try to predict if an image has multiple bounding boxes or just one etc. Imagine if you could predict these with such high accuracy that you wouldn't need to process some of the images in the network, and even if not you could use the information as a pre- or postprocessing step to further improve your model. </p>\n\n<p>I suggest starting simple with basic multivariate linear regression and moving slowly to more non-linear or more flexible models (if you have the time and interest).</p>\n\n<p>Feature selection, feature engineering and dimensionality reduction, similar terms yet they have different meaning. In feature selection our goal is to identify and select the most important features to reduce feature space, in feature engineering we create new features and in dimensionality reduction we combine existing features to reduce feature space. Depending on your data and goal, all three could be useful. Try principal component analysis (PCA), random forest algorithms, clustering etc. there are so many things you could do here just define your problem and approach it accordingly.</p>\n\n<p><strong>4. \"When are we gonna build our model??????\"</strong></p>\n\n<p>We have a problem, we have our data, we understand the association of each feature in our data towards what we are trying to predict and we have explored the data in many different ways. \nNaturally, you can build your model whenever you want, this is just the way I imagine my workflow. My next step would be to read, to read what other people have done with similar data and similar problem. What did they do? How well did they do? What problems did they encounter and how did they solve them? I would read papers e.g (ChexNet, U-net, YOLO), I would read the discussion forum, I would explore other peoples kernels cause there are some damn smart and helpful people on Kaggle. In some sense I would argue that this step should be done continuously, that is before, during and after the EDA.</p>\n\n<p>When you are tired of reading, reformulate your problem and read some more. \"Was it object detection that I wanted to solve or is it maybe easier to solve an image segmentation problem?\" \"This group used the U-net to segment lung images, is that better than the YOLO-architecture, how come, why and can I apply it to my problem and my data?</p>\n\n<p>As I said previously, I don't believe in blindly testing rather preparation to the extent that is possible is my method of preference. Everyone is different.</p>\n\n<p><strong>5. \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"</strong></p>\n\n<p>Trust me when I say, I know the feeling. When my interest in this field sparked, I was very confused and afraid. <strong>But with the right motivation, anyone can learn and fortunately we don't have to be experts in everything to be able to apply available tools, we can rely on the help of others and the community as a whole.</strong> </p>\n\n<p>When one is introduced to machine learning it is usually simple feed-forward multi-layer perceptron networks and then gradually we build our knowledge towards convolutional neural networks. You don't have to be an expert in linear algebra to understand convolution, naturally it is beneficial to have in-depth knowledge and it could help your results but it is not necessary for you to start exploring this field. Rather, understand the concepts and then utilize existing architectures to build your models, as many people do today. Deep learning is on many levels still a black box for everybody, even experienced people in the field!</p>\n\n<p><strong>6. Choosing the right model</strong></p>\n\n<p>Depending on how we have defined our problem in the previous steps, if we want to solve an object detection task we should look at the architectures used previously by others approaching the same problem, here are some reading recommendations: Sliding window algorithm (old school), Region-based CNN (R-CNN) and its subsequent developments, You only look once (YOLO) algorithm and its developments. I am sure you will encounter many more when reading up on these.</p>\n\n<p>If you instead decided that image segmentation is the way to go then you could naturally tweak the above mentioned architectures and I also think many are trying or have tried U-net which is just one of many options. And naturally, if you want to do image classification make sure to read up more on this.</p>\n\n<p><em>\"What model should I use?\"</em> \nOh, this is such a hard question. The honest answer is that I don't know, it takes a deep understanding and experience to know how these models will perform on a given data set. So, in the beginning I believe most people just try out different architectures and as you become more knowledgeable and comfortable you get a better feeling for how they do on different types of data.</p>\n\n<p><em>\"Does it matter which one I choose?\"\n\"I still don't know which one to choose, you suck!\"</em> \nYes, it does matter which one you choose but don't worry because this question will lead us to a very important concept of deep learning in competitions, namely <strong>Ensemble models</strong>. Combining your independent models, partnering with other people and combining models will greatly improve your score. This is a crucial step in competition, and a almost compulsory step for winners on Kaggle. <em>\"So, I'll just take the average  of all model predictions?\"</em> Well, yes if you want to try that, ensemble modelling is a complex topic and averaging, max voting is just the simple beginning. For more complex ensembling I would advise you to read up on boosting and bagging. </p>\n\n<p><strong>7. What now? I am still not winning!</strong></p>\n\n<p>Customize your architecture, rinse and repeat previous steps. Discuss with others on Kaggle, friends, family, random people on the street. Have you tried transfer learning (i.e. pretrained networks)? Did someone just post a really nice kernel suggesting something you haven't tried? Just get your brain going and I am sure you can think of a smart way to improve your model further, and don't be afraid to try new things. </p>\n\n<p><strong>8. \"This was all well and good, but I don't know how to use tensorflow or pytorch, let alone implement complex convolutional neural networks\"</strong></p>\n\n<p>For many this might be the case and sadly I don't have a better answer than \"it's never too late to learn\". You can also look for someone with complementary knowledge to yours and try to team up. Many network architectures are already implemented in pytorch and tensorflow so you don't have to reinvent the wheel, but some knowledge is necessary to be able to handle this. I believe in you and isn't it just so fun to learn. :)</p>\n\n<p>That was all I had the time to write up, the train has now stopped sadly. I am sure there are many things I have overlooked and missed. If anything, please complement and correct me in the comments. I hope someone finds this useful as a simple learning tool, please don't hesitate to drop a comment so that we can learn together! Best of luck everybody.</p>\n\n<p><em><strong>Sincerely,\nMe.</strong></em></p>",
  "messages": [
    {
      "id": "402945",
      "postDate": "10/12/2018 15:56:48",
      "content": "<p>Good morning, evening and night!</p>\n\n<p><strong>A tribute to the organizers</strong>\nJust an amazing competition. I've seen people from all sorts of backgrounds coming together to learn and compete. Hats off. </p>\n\n<p><strong>What is this thread about?</strong>\nNot everyone here is an expert in medicine and/or machine learning/statistics etc. so I figured for people trying to learn or improve their models I would do my best to offer some guidance and tips. For me personally, Kaggle is all about people coming together and learning new, intriguing things. I realise I write this a bit late into the competition but nevertheless, here we are. </p>\n\n<p><strong>Why should anyone even care about what you have to say?</strong>\nThis is most certainly not a definite guide for success just some general and specific things that I think newcomers (and more advanced learners) could find useful. I think there are more questions than answers in this thread, a good way to learn. Also, see the Kaggle getting started thread \"How to win Kaggle Competitions\" for Zeeshan-ul-hassan Usmans general input on winning competitions. </p>\n\n<p><strong>Who are you?</strong>\nI am <strong>NOT</strong> an expert in all fields mentioned above. <strong>I am a 24 years old MD/PhD student.</strong> I have almost finished my medical studies (MD) and I am also completing my PhD studies in medical data analysis, a rather fitting background aligned with my interests in data-driven healthcare. This is my first time posting on Kaggle, sadly, I do not have the time to engage in this competition and therefore hope to contribute with this.</p>\n\n<p><em>Let's get started.</em></p>\n\n<p><strong>1. Understand the problem you are trying to solve</strong></p>\n\n<p>You might think that this is an obvious statement, but the complexity is not to be underestimated here. <em>Please allow me to expand on this,</em> \nWhat are we trying to predict in this competition? Are we trying to predict whether someone has pneumonia or not? Or are we trying to predict one or more \"bounding boxes\" in a given image? What is the difference between these problems?</p>\n\n<p>\"Ground-truth bounding boxes\" (some prefer \"reference bounding boxes\"), manually drawn boxes on an image by radiologists highlighting opacities or other suspect changes in the chest x-rays. (see Guy Zahavis excellent kernel). </p>\n\n<p>If one really wants to learn, I think for some questions it might be better to search for the answers yourselves, I believe that asking the right question is harder than finding the answer in many cases. I will offer some guidance for anyone who wants a little push.\nWe are trying to predict bounding boxes in these images, and we are also trying to predict if someone has opacities (indicating pneumonia) or not. There is an evident difference between these problems, one of them we can call \"object detection\" and the other \"image classification\". So, as a newcomer I would approach this very simply by understanding what models others have used for object detection tasks in the past and do the same for image classification (more on this later).</p>\n\n<p>From a competition standpoint the evaluation metric is crucial to understand, I would honestly spend hours and hours trying to understand just the metric (Intersect over Union) if I wanted to improve my score. Some key things that I would look for would be, first of all what is this mysterious IoU thing? How is it affected by false prediction, in the cases my model predicts say a bounding box where there shouldn't be one? Is it maybe worse to make a false positive prediction, than no prediction at all when there should be one? etc. This is important, because you can build your model accordingly and increase your score. Does it matter whether your model correctly predicts if someone has pneumonia or not? How is this evaluated and relatively, how important is this compared to predicting the bounding boxes? That is, what should I focus my time and effort on? </p>\n\n<p>There are many more things to consider before we even get started with trying to solve the problem but I don't want to write a book that no one will end up reading so lets move on. Just remember, you can endlessly and blindly test your way forward or you can really take the time to understand and formulate the problem, which I guarantee will not only lead to a better score but also to a deeper understanding. </p>\n\n<p><strong>2. Understand all the available data</strong></p>\n\n<p>Let's say that we have formulated the problem and now have a good grasp of the evaluation metric. A logical next step is trying to grasp what kind of data we are working with, I am not talking about any statistics yet. I literally mean, what features/variables are available to us? What do they mean in the relevant medical context? A good example is the AP-PA discussions that I briefly saw some people discussing, in which AP (Anterio-Posterior) and PA (Posterio-Anterior) represent different values in the available dicom-meta data indicating whether a patient was lying down during the scan or standing up. Some smart people quickly realised the meaning of this feature from a medical perspective, where a patient who was lying down for the chest-x-ray most likely is more weakened than a patient who could stand up, logically inferring that such a patient is more likely to have opacities.</p>\n\n<p>Another important part is to understand the weaknesses of the data that you are training on comparatively to the data you will be testing on. One such event with this data is perhaps, will the same radiologists that labelled the training images also label the testing images? What is the inter-individual (between persons) variation? </p>\n\n<p>A side note is to explore whether you can expand or improve your data set i.e. better or additional labels with external sources etc. (if allowed). </p>\n\n<p><strong>3. Explore the data [EDA = Exploratory data analysis]</strong></p>\n\n<p>Cool, now we have a problem to solve and we know what data we are working with. You will notice that a large part of what I am discussing in this thread has nothing to with complex mathematics and advanced machine learning network architecture. This is because I believe that the fundamentals are equally important as all the complexities of building a model. I think a lot of people would agree with me when I say that the EDA is extremely important to the outcome of your models. The EDA consists of many parts, only limited to ones imagination, time and knowledge. I will highlight some parts that I think are important and underestimated. </p>\n\n<p>Do not skip the simple stuff like checking distribution, outliers and simple correlations. It might be boring, and one can probably argue that it won't affect the outcome. I personally believe that it is a good habit to explore these simple things to better understand your data regardless. A big part of why I believe that these things are important is because the deeper understanding of the features both context-based and statistically could be helpful for you to develop new features, and as many large names in the field of machine learning has stated i.e. Andrew Ng, feature engineering is a very important part of applied machine learning.</p>\n\n<p>\"Feature engineering\" simply means creating new features by combining features or looking at existing data from a new perspective, one such example is to look at the grey-scale histograms of the images or simple stuff like creating a feature \"bounding box area\". ENDLESS OPPORTUNITIES and smarter people than me are definitely doing this in ways that I probably wouldn't even understand in this competition. </p>\n\n<p>A more statistically \"complex\" part of EDA could be to build multi-variate models and try to predict based on a patients' age, gender, view position and whatever features you believe to be important if an image has a bounding box, or try to predict the designated class (i.e. pneumonia or not), try to predict if an image has multiple bounding boxes or just one etc. Imagine if you could predict these with such high accuracy that you wouldn't need to process some of the images in the network, and even if not you could use the information as a pre- or postprocessing step to further improve your model. </p>\n\n<p>I suggest starting simple with basic multivariate linear regression and moving slowly to more non-linear or more flexible models (if you have the time and interest).</p>\n\n<p>Feature selection, feature engineering and dimensionality reduction, similar terms yet they have different meaning. In feature selection our goal is to identify and select the most important features to reduce feature space, in feature engineering we create new features and in dimensionality reduction we combine existing features to reduce feature space. Depending on your data and goal, all three could be useful. Try principal component analysis (PCA), random forest algorithms, clustering etc. there are so many things you could do here just define your problem and approach it accordingly.</p>\n\n<p><strong>4. \"When are we gonna build our model??????\"</strong></p>\n\n<p>We have a problem, we have our data, we understand the association of each feature in our data towards what we are trying to predict and we have explored the data in many different ways. \nNaturally, you can build your model whenever you want, this is just the way I imagine my workflow. My next step would be to read, to read what other people have done with similar data and similar problem. What did they do? How well did they do? What problems did they encounter and how did they solve them? I would read papers e.g (ChexNet, U-net, YOLO), I would read the discussion forum, I would explore other peoples kernels cause there are some damn smart and helpful people on Kaggle. In some sense I would argue that this step should be done continuously, that is before, during and after the EDA.</p>\n\n<p>When you are tired of reading, reformulate your problem and read some more. \"Was it object detection that I wanted to solve or is it maybe easier to solve an image segmentation problem?\" \"This group used the U-net to segment lung images, is that better than the YOLO-architecture, how come, why and can I apply it to my problem and my data?</p>\n\n<p>As I said previously, I don't believe in blindly testing rather preparation to the extent that is possible is my method of preference. Everyone is different.</p>\n\n<p><strong>5. \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"</strong></p>\n\n<p>Trust me when I say, I know the feeling. When my interest in this field sparked, I was very confused and afraid. <strong>But with the right motivation, anyone can learn and fortunately we don't have to be experts in everything to be able to apply available tools, we can rely on the help of others and the community as a whole.</strong> </p>\n\n<p>When one is introduced to machine learning it is usually simple feed-forward multi-layer perceptron networks and then gradually we build our knowledge towards convolutional neural networks. You don't have to be an expert in linear algebra to understand convolution, naturally it is beneficial to have in-depth knowledge and it could help your results but it is not necessary for you to start exploring this field. Rather, understand the concepts and then utilize existing architectures to build your models, as many people do today. Deep learning is on many levels still a black box for everybody, even experienced people in the field!</p>\n\n<p><strong>6. Choosing the right model</strong></p>\n\n<p>Depending on how we have defined our problem in the previous steps, if we want to solve an object detection task we should look at the architectures used previously by others approaching the same problem, here are some reading recommendations: Sliding window algorithm (old school), Region-based CNN (R-CNN) and its subsequent developments, You only look once (YOLO) algorithm and its developments. I am sure you will encounter many more when reading up on these.</p>\n\n<p>If you instead decided that image segmentation is the way to go then you could naturally tweak the above mentioned architectures and I also think many are trying or have tried U-net which is just one of many options. And naturally, if you want to do image classification make sure to read up more on this.</p>\n\n<p><em>\"What model should I use?\"</em> \nOh, this is such a hard question. The honest answer is that I don't know, it takes a deep understanding and experience to know how these models will perform on a given data set. So, in the beginning I believe most people just try out different architectures and as you become more knowledgeable and comfortable you get a better feeling for how they do on different types of data.</p>\n\n<p><em>\"Does it matter which one I choose?\"\n\"I still don't know which one to choose, you suck!\"</em> \nYes, it does matter which one you choose but don't worry because this question will lead us to a very important concept of deep learning in competitions, namely <strong>Ensemble models</strong>. Combining your independent models, partnering with other people and combining models will greatly improve your score. This is a crucial step in competition, and a almost compulsory step for winners on Kaggle. <em>\"So, I'll just take the average  of all model predictions?\"</em> Well, yes if you want to try that, ensemble modelling is a complex topic and averaging, max voting is just the simple beginning. For more complex ensembling I would advise you to read up on boosting and bagging. </p>\n\n<p><strong>7. What now? I am still not winning!</strong></p>\n\n<p>Customize your architecture, rinse and repeat previous steps. Discuss with others on Kaggle, friends, family, random people on the street. Have you tried transfer learning (i.e. pretrained networks)? Did someone just post a really nice kernel suggesting something you haven't tried? Just get your brain going and I am sure you can think of a smart way to improve your model further, and don't be afraid to try new things. </p>\n\n<p><strong>8. \"This was all well and good, but I don't know how to use tensorflow or pytorch, let alone implement complex convolutional neural networks\"</strong></p>\n\n<p>For many this might be the case and sadly I don't have a better answer than \"it's never too late to learn\". You can also look for someone with complementary knowledge to yours and try to team up. Many network architectures are already implemented in pytorch and tensorflow so you don't have to reinvent the wheel, but some knowledge is necessary to be able to handle this. I believe in you and isn't it just so fun to learn. :)</p>\n\n<p>That was all I had the time to write up, the train has now stopped sadly. I am sure there are many things I have overlooked and missed. If anything, please complement and correct me in the comments. I hope someone finds this useful as a simple learning tool, please don't hesitate to drop a comment so that we can learn together! Best of luck everybody.</p>\n\n<p><em><strong>Sincerely,\nMe.</strong></em></p>",
      "rawMarkdown": "Good morning, evening and night!\n\n**A tribute to the organizers**\nJust an amazing competition. I've seen people from all sorts of backgrounds coming together to learn and compete. Hats off. \n\n**What is this thread about?**\nNot everyone here is an expert in medicine and/or machine learning/statistics etc. so I figured for people trying to learn or improve their models I would do my best to offer some guidance and tips. For me personally, Kaggle is all about people coming together and learning new, intriguing things. I realise I write this a bit late into the competition but nevertheless, here we are. \n\n**Why should anyone even care about what you have to say?**\nThis is most certainly not a definite guide for success just some general and specific things that I think newcomers (and more advanced learners) could find useful. I think there are more questions than answers in this thread, a good way to learn. Also, see the Kaggle getting started thread \"How to win Kaggle Competitions\" for Zeeshan-ul-hassan Usmans general input on winning competitions. \n\n**Who are you?**\nI am **NOT** an expert in all fields mentioned above. **I am a 24 years old MD/PhD student.** I have almost finished my medical studies (MD) and I am also completing my PhD studies in medical data analysis, a rather fitting background aligned with my interests in data-driven healthcare. This is my first time posting on Kaggle, sadly, I do not have the time to engage in this competition and therefore hope to contribute with this.\n\n*Let's get started.*\n\n**1. Understand the problem you are trying to solve**\n\nYou might think that this is an obvious statement, but the complexity is not to be underestimated here. *Please allow me to expand on this,* \nWhat are we trying to predict in this competition? Are we trying to predict whether someone has pneumonia or not? Or are we trying to predict one or more \"bounding boxes\" in a given image? What is the difference between these problems?\n\n\"Ground-truth bounding boxes\" (some prefer \"reference bounding boxes\"), manually drawn boxes on an image by radiologists highlighting opacities or other suspect changes in the chest x-rays. (see Guy Zahavis excellent kernel). \n\nIf one really wants to learn, I think for some questions it might be better to search for the answers yourselves, I believe that asking the right question is harder than finding the answer in many cases. I will offer some guidance for anyone who wants a little push.\nWe are trying to predict bounding boxes in these images, and we are also trying to predict if someone has opacities (indicating pneumonia) or not. There is an evident difference between these problems, one of them we can call \"object detection\" and the other \"image classification\". So, as a newcomer I would approach this very simply by understanding what models others have used for object detection tasks in the past and do the same for image classification (more on this later).\n\nFrom a competition standpoint the evaluation metric is crucial to understand, I would honestly spend hours and hours trying to understand just the metric (Intersect over Union) if I wanted to improve my score. Some key things that I would look for would be, first of all what is this mysterious IoU thing? How is it affected by false prediction, in the cases my model predicts say a bounding box where there shouldn't be one? Is it maybe worse to make a false positive prediction, than no prediction at all when there should be one? etc. This is important, because you can build your model accordingly and increase your score. Does it matter whether your model correctly predicts if someone has pneumonia or not? How is this evaluated and relatively, how important is this compared to predicting the bounding boxes? That is, what should I focus my time and effort on? \n\nThere are many more things to consider before we even get started with trying to solve the problem but I don't want to write a book that no one will end up reading so lets move on. Just remember, you can endlessly and blindly test your way forward or you can really take the time to understand and formulate the problem, which I guarantee will not only lead to a better score but also to a deeper understanding. \n\n**2. Understand all the available data**\n\nLet's say that we have formulated the problem and now have a good grasp of the evaluation metric. A logical next step is trying to grasp what kind of data we are working with, I am not talking about any statistics yet. I literally mean, what features/variables are available to us? What do they mean in the relevant medical context? A good example is the AP-PA discussions that I briefly saw some people discussing, in which AP (Anterio-Posterior) and PA (Posterio-Anterior) represent different values in the available dicom-meta data indicating whether a patient was lying down during the scan or standing up. Some smart people quickly realised the meaning of this feature from a medical perspective, where a patient who was lying down for the chest-x-ray most likely is more weakened than a patient who could stand up, logically inferring that such a patient is more likely to have opacities.\n\nAnother important part is to understand the weaknesses of the data that you are training on comparatively to the data you will be testing on. One such event with this data is perhaps, will the same radiologists that labelled the training images also label the testing images? What is the inter-individual (between persons) variation? \n\nA side note is to explore whether you can expand or improve your data set i.e. better or additional labels with external sources etc. (if allowed). \n\n**3. Explore the data [EDA = Exploratory data analysis]**\n\nCool, now we have a problem to solve and we know what data we are working with. You will notice that a large part of what I am discussing in this thread has nothing to with complex mathematics and advanced machine learning network architecture. This is because I believe that the fundamentals are equally important as all the complexities of building a model. I think a lot of people would agree with me when I say that the EDA is extremely important to the outcome of your models. The EDA consists of many parts, only limited to ones imagination, time and knowledge. I will highlight some parts that I think are important and underestimated. \n\nDo not skip the simple stuff like checking distribution, outliers and simple correlations. It might be boring, and one can probably argue that it won't affect the outcome. I personally believe that it is a good habit to explore these simple things to better understand your data regardless. A big part of why I believe that these things are important is because the deeper understanding of the features both context-based and statistically could be helpful for you to develop new features, and as many large names in the field of machine learning has stated i.e. Andrew Ng, feature engineering is a very important part of applied machine learning.\n\n\"Feature engineering\" simply means creating new features by combining features or looking at existing data from a new perspective, one such example is to look at the grey-scale histograms of the images or simple stuff like creating a feature \"bounding box area\". ENDLESS OPPORTUNITIES and smarter people than me are definitely doing this in ways that I probably wouldn't even understand in this competition. \n\nA more statistically \"complex\" part of EDA could be to build multi-variate models and try to predict based on a patients' age, gender, view position and whatever features you believe to be important if an image has a bounding box, or try to predict the designated class (i.e. pneumonia or not), try to predict if an image has multiple bounding boxes or just one etc. Imagine if you could predict these with such high accuracy that you wouldn't need to process some of the images in the network, and even if not you could use the information as a pre- or postprocessing step to further improve your model. \n\nI suggest starting simple with basic multivariate linear regression and moving slowly to more non-linear or more flexible models (if you have the time and interest).\n\nFeature selection, feature engineering and dimensionality reduction, similar terms yet they have different meaning. In feature selection our goal is to identify and select the most important features to reduce feature space, in feature engineering we create new features and in dimensionality reduction we combine existing features to reduce feature space. Depending on your data and goal, all three could be useful. Try principal component analysis (PCA), random forest algorithms, clustering etc. there are so many things you could do here just define your problem and approach it accordingly.\n\n**4. \"When are we gonna build our model??????\"**\n\nWe have a problem, we have our data, we understand the association of each feature in our data towards what we are trying to predict and we have explored the data in many different ways. \nNaturally, you can build your model whenever you want, this is just the way I imagine my workflow. My next step would be to read, to read what other people have done with similar data and similar problem. What did they do? How well did they do? What problems did they encounter and how did they solve them? I would read papers e.g (ChexNet, U-net, YOLO), I would read the discussion forum, I would explore other peoples kernels cause there are some damn smart and helpful people on Kaggle. In some sense I would argue that this step should be done continuously, that is before, during and after the EDA.\n\nWhen you are tired of reading, reformulate your problem and read some more. \"Was it object detection that I wanted to solve or is it maybe easier to solve an image segmentation problem?\" \"This group used the U-net to segment lung images, is that better than the YOLO-architecture, how come, why and can I apply it to my problem and my data?\n\nAs I said previously, I don't believe in blindly testing rather preparation to the extent that is possible is my method of preference. Everyone is different.\n \n**5. \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"**\n\nTrust me when I say, I know the feeling. When my interest in this field sparked, I was very confused and afraid. **But with the right motivation, anyone can learn and fortunately we don't have to be experts in everything to be able to apply available tools, we can rely on the help of others and the community as a whole.** \n\nWhen one is introduced to machine learning it is usually simple feed-forward multi-layer perceptron networks and then gradually we build our knowledge towards convolutional neural networks. You don't have to be an expert in linear algebra to understand convolution, naturally it is beneficial to have in-depth knowledge and it could help your results but it is not necessary for you to start exploring this field. Rather, understand the concepts and then utilize existing architectures to build your models, as many people do today. Deep learning is on many levels still a black box for everybody, even experienced people in the field!\n\n**6. Choosing the right model**\n\nDepending on how we have defined our problem in the previous steps, if we want to solve an object detection task we should look at the architectures used previously by others approaching the same problem, here are some reading recommendations: Sliding window algorithm (old school), Region-based CNN (R-CNN) and its subsequent developments, You only look once (YOLO) algorithm and its developments. I am sure you will encounter many more when reading up on these.\n\nIf you instead decided that image segmentation is the way to go then you could naturally tweak the above mentioned architectures and I also think many are trying or have tried U-net which is just one of many options. And naturally, if you want to do image classification make sure to read up more on this.\n\n*\"What model should I use?\"* \nOh, this is such a hard question. The honest answer is that I don't know, it takes a deep understanding and experience to know how these models will perform on a given data set. So, in the beginning I believe most people just try out different architectures and as you become more knowledgeable and comfortable you get a better feeling for how they do on different types of data.\n\n*\"Does it matter which one I choose?\"\n\"I still don't know which one to choose, you suck!\"* \nYes, it does matter which one you choose but don't worry because this question will lead us to a very important concept of deep learning in competitions, namely **Ensemble models**. Combining your independent models, partnering with other people and combining models will greatly improve your score. This is a crucial step in competition, and a almost compulsory step for winners on Kaggle. *\"So, I'll just take the average  of all model predictions?\"* Well, yes if you want to try that, ensemble modelling is a complex topic and averaging, max voting is just the simple beginning. For more complex ensembling I would advise you to read up on boosting and bagging. \n\n**7. What now? I am still not winning!**\n\nCustomize your architecture, rinse and repeat previous steps. Discuss with others on Kaggle, friends, family, random people on the street. Have you tried transfer learning (i.e. pretrained networks)? Did someone just post a really nice kernel suggesting something you haven't tried? Just get your brain going and I am sure you can think of a smart way to improve your model further, and don't be afraid to try new things. \n\n**8. \"This was all well and good, but I don't know how to use tensorflow or pytorch, let alone implement complex convolutional neural networks\"**\n\nFor many this might be the case and sadly I don't have a better answer than \"it's never too late to learn\". You can also look for someone with complementary knowledge to yours and try to team up. Many network architectures are already implemented in pytorch and tensorflow so you don't have to reinvent the wheel, but some knowledge is necessary to be able to handle this. I believe in you and isn't it just so fun to learn. :)\n\nThat was all I had the time to write up, the train has now stopped sadly. I am sure there are many things I have overlooked and missed. If anything, please complement and correct me in the comments. I hope someone finds this useful as a simple learning tool, please don't hesitate to drop a comment so that we can learn together! Best of luck everybody.\n\n***Sincerely,\nMe.***",
      "votes": null
    },
    {
      "id": "403637",
      "postDate": "10/14/2018 06:58:56",
      "content": "<p>Some motivation on ** \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"**</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=2HMPRXstSvQ\">Geoff Hinton: On Radiology Nov,2016</a></p>",
      "rawMarkdown": "Some motivation on ** \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"**\n\n[Geoff Hinton: On Radiology Nov,2016][1]\n\n\n  [1]: https://www.youtube.com/watch?v=2HMPRXstSvQ",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 403637,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "10/14/2018 06:58:56",
      "content": "<p>Some motivation on ** \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"**</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=2HMPRXstSvQ\">Geoff Hinton: On Radiology Nov,2016</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "402945": "Good morning, evening and night!\n\n**A tribute to the organizers**\nJust an amazing competition. I've seen people from all sorts of backgrounds coming together to learn and compete. Hats off. \n\n**What is this thread about?**\nNot everyone here is an expert in medicine and/or machine learning/statistics etc. so I figured for people trying to learn or improve their models I would do my best to offer some guidance and tips. For me personally, Kaggle is all about people coming together and learning new, intriguing things. I realise I write this a bit late into the competition but nevertheless, here we are. \n\n**Why should anyone even care about what you have to say?**\nThis is most certainly not a definite guide for success just some general and specific things that I think newcomers (and more advanced learners) could find useful. I think there are more questions than answers in this thread, a good way to learn. Also, see the Kaggle getting started thread \"How to win Kaggle Competitions\" for Zeeshan-ul-hassan Usmans general input on winning competitions. \n\n**Who are you?**\nI am **NOT** an expert in all fields mentioned above. **I am a 24 years old MD/PhD student.** I have almost finished my medical studies (MD) and I am also completing my PhD studies in medical data analysis, a rather fitting background aligned with my interests in data-driven healthcare. This is my first time posting on Kaggle, sadly, I do not have the time to engage in this competition and therefore hope to contribute with this.\n\n*Let's get started.*\n\n**1. Understand the problem you are trying to solve**\n\nYou might think that this is an obvious statement, but the complexity is not to be underestimated here. *Please allow me to expand on this,* \nWhat are we trying to predict in this competition? Are we trying to predict whether someone has pneumonia or not? Or are we trying to predict one or more \"bounding boxes\" in a given image? What is the difference between these problems?\n\n\"Ground-truth bounding boxes\" (some prefer \"reference bounding boxes\"), manually drawn boxes on an image by radiologists highlighting opacities or other suspect changes in the chest x-rays. (see Guy Zahavis excellent kernel). \n\nIf one really wants to learn, I think for some questions it might be better to search for the answers yourselves, I believe that asking the right question is harder than finding the answer in many cases. I will offer some guidance for anyone who wants a little push.\nWe are trying to predict bounding boxes in these images, and we are also trying to predict if someone has opacities (indicating pneumonia) or not. There is an evident difference between these problems, one of them we can call \"object detection\" and the other \"image classification\". So, as a newcomer I would approach this very simply by understanding what models others have used for object detection tasks in the past and do the same for image classification (more on this later).\n\nFrom a competition standpoint the evaluation metric is crucial to understand, I would honestly spend hours and hours trying to understand just the metric (Intersect over Union) if I wanted to improve my score. Some key things that I would look for would be, first of all what is this mysterious IoU thing? How is it affected by false prediction, in the cases my model predicts say a bounding box where there shouldn't be one? Is it maybe worse to make a false positive prediction, than no prediction at all when there should be one? etc. This is important, because you can build your model accordingly and increase your score. Does it matter whether your model correctly predicts if someone has pneumonia or not? How is this evaluated and relatively, how important is this compared to predicting the bounding boxes? That is, what should I focus my time and effort on? \n\nThere are many more things to consider before we even get started with trying to solve the problem but I don't want to write a book that no one will end up reading so lets move on. Just remember, you can endlessly and blindly test your way forward or you can really take the time to understand and formulate the problem, which I guarantee will not only lead to a better score but also to a deeper understanding. \n\n**2. Understand all the available data**\n\nLet's say that we have formulated the problem and now have a good grasp of the evaluation metric. A logical next step is trying to grasp what kind of data we are working with, I am not talking about any statistics yet. I literally mean, what features/variables are available to us? What do they mean in the relevant medical context? A good example is the AP-PA discussions that I briefly saw some people discussing, in which AP (Anterio-Posterior) and PA (Posterio-Anterior) represent different values in the available dicom-meta data indicating whether a patient was lying down during the scan or standing up. Some smart people quickly realised the meaning of this feature from a medical perspective, where a patient who was lying down for the chest-x-ray most likely is more weakened than a patient who could stand up, logically inferring that such a patient is more likely to have opacities.\n\nAnother important part is to understand the weaknesses of the data that you are training on comparatively to the data you will be testing on. One such event with this data is perhaps, will the same radiologists that labelled the training images also label the testing images? What is the inter-individual (between persons) variation? \n\nA side note is to explore whether you can expand or improve your data set i.e. better or additional labels with external sources etc. (if allowed). \n\n**3. Explore the data [EDA = Exploratory data analysis]**\n\nCool, now we have a problem to solve and we know what data we are working with. You will notice that a large part of what I am discussing in this thread has nothing to with complex mathematics and advanced machine learning network architecture. This is because I believe that the fundamentals are equally important as all the complexities of building a model. I think a lot of people would agree with me when I say that the EDA is extremely important to the outcome of your models. The EDA consists of many parts, only limited to ones imagination, time and knowledge. I will highlight some parts that I think are important and underestimated. \n\nDo not skip the simple stuff like checking distribution, outliers and simple correlations. It might be boring, and one can probably argue that it won't affect the outcome. I personally believe that it is a good habit to explore these simple things to better understand your data regardless. A big part of why I believe that these things are important is because the deeper understanding of the features both context-based and statistically could be helpful for you to develop new features, and as many large names in the field of machine learning has stated i.e. Andrew Ng, feature engineering is a very important part of applied machine learning.\n\n\"Feature engineering\" simply means creating new features by combining features or looking at existing data from a new perspective, one such example is to look at the grey-scale histograms of the images or simple stuff like creating a feature \"bounding box area\". ENDLESS OPPORTUNITIES and smarter people than me are definitely doing this in ways that I probably wouldn't even understand in this competition. \n\nA more statistically \"complex\" part of EDA could be to build multi-variate models and try to predict based on a patients' age, gender, view position and whatever features you believe to be important if an image has a bounding box, or try to predict the designated class (i.e. pneumonia or not), try to predict if an image has multiple bounding boxes or just one etc. Imagine if you could predict these with such high accuracy that you wouldn't need to process some of the images in the network, and even if not you could use the information as a pre- or postprocessing step to further improve your model. \n\nI suggest starting simple with basic multivariate linear regression and moving slowly to more non-linear or more flexible models (if you have the time and interest).\n\nFeature selection, feature engineering and dimensionality reduction, similar terms yet they have different meaning. In feature selection our goal is to identify and select the most important features to reduce feature space, in feature engineering we create new features and in dimensionality reduction we combine existing features to reduce feature space. Depending on your data and goal, all three could be useful. Try principal component analysis (PCA), random forest algorithms, clustering etc. there are so many things you could do here just define your problem and approach it accordingly.\n\n**4. \"When are we gonna build our model??????\"**\n\nWe have a problem, we have our data, we understand the association of each feature in our data towards what we are trying to predict and we have explored the data in many different ways. \nNaturally, you can build your model whenever you want, this is just the way I imagine my workflow. My next step would be to read, to read what other people have done with similar data and similar problem. What did they do? How well did they do? What problems did they encounter and how did they solve them? I would read papers e.g (ChexNet, U-net, YOLO), I would read the discussion forum, I would explore other peoples kernels cause there are some damn smart and helpful people on Kaggle. In some sense I would argue that this step should be done continuously, that is before, during and after the EDA.\n\nWhen you are tired of reading, reformulate your problem and read some more. \"Was it object detection that I wanted to solve or is it maybe easier to solve an image segmentation problem?\" \"This group used the U-net to segment lung images, is that better than the YOLO-architecture, how come, why and can I apply it to my problem and my data?\n\nAs I said previously, I don't believe in blindly testing rather preparation to the extent that is possible is my method of preference. Everyone is different.\n \n**5. \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"**\n\nTrust me when I say, I know the feeling. When my interest in this field sparked, I was very confused and afraid. **But with the right motivation, anyone can learn and fortunately we don't have to be experts in everything to be able to apply available tools, we can rely on the help of others and the community as a whole.** \n\nWhen one is introduced to machine learning it is usually simple feed-forward multi-layer perceptron networks and then gradually we build our knowledge towards convolutional neural networks. You don't have to be an expert in linear algebra to understand convolution, naturally it is beneficial to have in-depth knowledge and it could help your results but it is not necessary for you to start exploring this field. Rather, understand the concepts and then utilize existing architectures to build your models, as many people do today. Deep learning is on many levels still a black box for everybody, even experienced people in the field!\n\n**6. Choosing the right model**\n\nDepending on how we have defined our problem in the previous steps, if we want to solve an object detection task we should look at the architectures used previously by others approaching the same problem, here are some reading recommendations: Sliding window algorithm (old school), Region-based CNN (R-CNN) and its subsequent developments, You only look once (YOLO) algorithm and its developments. I am sure you will encounter many more when reading up on these.\n\nIf you instead decided that image segmentation is the way to go then you could naturally tweak the above mentioned architectures and I also think many are trying or have tried U-net which is just one of many options. And naturally, if you want to do image classification make sure to read up more on this.\n\n*\"What model should I use?\"* \nOh, this is such a hard question. The honest answer is that I don't know, it takes a deep understanding and experience to know how these models will perform on a given data set. So, in the beginning I believe most people just try out different architectures and as you become more knowledgeable and comfortable you get a better feeling for how they do on different types of data.\n\n*\"Does it matter which one I choose?\"\n\"I still don't know which one to choose, you suck!\"* \nYes, it does matter which one you choose but don't worry because this question will lead us to a very important concept of deep learning in competitions, namely **Ensemble models**. Combining your independent models, partnering with other people and combining models will greatly improve your score. This is a crucial step in competition, and a almost compulsory step for winners on Kaggle. *\"So, I'll just take the average  of all model predictions?\"* Well, yes if you want to try that, ensemble modelling is a complex topic and averaging, max voting is just the simple beginning. For more complex ensembling I would advise you to read up on boosting and bagging. \n\n**7. What now? I am still not winning!**\n\nCustomize your architecture, rinse and repeat previous steps. Discuss with others on Kaggle, friends, family, random people on the street. Have you tried transfer learning (i.e. pretrained networks)? Did someone just post a really nice kernel suggesting something you haven't tried? Just get your brain going and I am sure you can think of a smart way to improve your model further, and don't be afraid to try new things. \n\n**8. \"This was all well and good, but I don't know how to use tensorflow or pytorch, let alone implement complex convolutional neural networks\"**\n\nFor many this might be the case and sadly I don't have a better answer than \"it's never too late to learn\". You can also look for someone with complementary knowledge to yours and try to team up. Many network architectures are already implemented in pytorch and tensorflow so you don't have to reinvent the wheel, but some knowledge is necessary to be able to handle this. I believe in you and isn't it just so fun to learn. :)\n\nThat was all I had the time to write up, the train has now stopped sadly. I am sure there are many things I have overlooked and missed. If anything, please complement and correct me in the comments. I hope someone finds this useful as a simple learning tool, please don't hesitate to drop a comment so that we can learn together! Best of luck everybody.\n\n***Sincerely,\nMe.***",
    "403637": "Some motivation on ** \"WOW, I am just a medical doctor... Deep learning is so mathematical and hard!!\"**\n\n[Geoff Hinton: On Radiology Nov,2016][1]\n\n\n  [1]: https://www.youtube.com/watch?v=2HMPRXstSvQ"
  },
  "source": "meta"
}