{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Author**: Xiaoqi Jia, M.E. in Information and Communication Engineering, School of Computer Science and Engineering, Northeastern University.\n\n## Abstract\n\nNoisy label, resulting from omissions during the labeling process for supervised learning, will lead to overfitting and degrade the generalization performance. However, in real-world applications, it is difficult to obtain ideal datasets without any noise, especially for histopathology images, since it is time-consuming and error-prone to manually search histopathology structures of stained tissue specimens under microscopes. In this notebook, I will first discuss and visualize the impact of noise label, and then propose a method called Energy Map to minimize the gap between noisy label and clean label. This method is superior to the baseline with about 4% improvement in private leaderboard, demonstrating its effectiveness of histopathology semantic segmentation under noisy supervision.\n\nSome details and the original can be found at https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/347967\n\nThe code of submissions can be found at https://www.kaggle.com/code/gray98/learning-with-noisy-label\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## 1. Introduction\n\n\nWill noisy label impede the performance of models? Of course, yes. Noisy label will inevitably lead to model performance degradation, which is proved by many studies[1], [2]. But label noise is unavoidable in many medical image datasets. It can be caused by limited attention of human annotators, subjective nature of labeling, or errors in computerized labeling systems[3]. \n\nSo, is there noise in the annotations of our dataset in this competition?\n\nIt is hard to judge it without expertise. We might as well assume that noise exists, since we do observe some examples with inconsistent annotations, as shown in Figure 1. From my perspective, the regions within the red lines contain the unlabeled tissue which is very similar to the labeled tissue. <mark>**So my motivation is to find out if the performance of the model will be improved when we pay more attention to these unmarked tissue areas?**<mark>\n\n![截屏2022-10-01 上午7.44.52.png](https://s2.loli.net/2022/10/01/nuqR5WjN8bpJHcF.png)\n\n**It is worth noting that when conducting all experiments, I assumed that annotations of this dataset maybe noisy, and this assumption is based on my non-professional subjective perception**","metadata":{}},{"cell_type":"markdown","source":"The details of my overall process for exploring this problem are shown in Figure 2. First, I artificially simulated a dataset with a lot of label noise, and then designed a method that can reduce the impact of noisy label, and validated it on this artificially noisy dataset, as Figure 2(a). Afterwards, I applied this improvement to the actual dataset and verify whether there are omissions in this actual dataset by relabeling the original dataset, as Figure 2(b).","metadata":{}},{"cell_type":"markdown","source":"![fig2.png](https://s2.loli.net/2022/10/01/NcjnmbE8UL4khRg.png)","metadata":{}},{"cell_type":"markdown","source":"## 2. Visualization of the Influence from Noisy Label\n\n> What I cannot create, I do not understand.\n\n### Background\n\nImperfect annotations, such as missing label and noisy label, degrade the performance of models. This degradation is embodied in numerical results in most classification researches. In segmentation task, the results of segmentation prediction also manifest the negative influence of imperfect label. For classification task, many previous works provide various technologies, like DeepDream and CAM, to visualize classification models, which can understand CNN and correct training mishaps. Two examples are shown in Figure 3. However, visualization is not used in segmentation models since uncertainty in segmentation models is low. For instance, in classification task, the co-occurrence of train and railroad is high so we afraid our classification models misclassify a railroad as train. This mishap would not exist in segmentation task because segmentation supervision provides both classification and location information. However, uncertainty comes back when segmentation supervision contains omissions and noise. It intuitively occurs to us that visualizing the influence from noisy label can qualitatively tell us what our networks learn under noisy label and verify whether our improvements are effective. (This idea is from @hengck23, which can be found at https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/347967. My implementation may be a little different from the original idea.)\n\n![fig3.png](https://s2.loli.net/2022/10/01/K67VItljZLi5whr.png)\n\nThere is an easy way to understand visualization of a model. We can think of training classification networks or segmentation networks as letting the networks learn multiple-choice questions. Networks only provide a yes or no answer for all possible outcomes. For example, train a network to classify or segment cats out of cats-and-dogs datasets. After finishing the training process, I am curious about what my model thinks the cat is like and whether its cognition is the same as my subjective cognition. So I give my model a blank sheet of paper and say: please draw a cat on this paper. In the other way, to draw a cat is to let the model do essay questions or create, which can more intuitively observe the results of the model learning.\n\n### Method\n\nI adopt a very simple way to visualize of segmentation models by using generating, as shown in  Figure 4. First, train a segmentation model as a common process like Figure 4(a), and then freeze the weight of the segmentation model when the model converges. Then, use the trained model from Figure 4(a) to gradually change the input image from randomly generated images to the images that fit manual fake labels.\n\n![截屏2022-10-01 上午7.50.09.png](https://s2.loli.net/2022/10/01/DroW8njf6Ibm5CT.png)","metadata":{}},{"cell_type":"markdown","source":"### A Toy Experiment \n\n#### Simulate Noise in Clean Dataset:\n\nBefore delving into the problem of noise label, I first need to create a dataset with controllable noise. I chose all kidney histopathology images available for training in Hubmap competition data for my toy experiments, since identifying the FTUs of kidney is easier for me. I randomly selected 50% of the FTUs to mask, which simulates very strong noise.\n\n#### Train with Noisy and Clean Label:\n\nThis step is very simple. Just use noisy label and clean label separately for model training, and observe the experiment results, as shown in Figure 5.\n\n![截屏2022-10-01 上午7.48.14.png](https://s2.loli.net/2022/10/01/Uawlrdn1IuPhyCV.png)\n\n#### Visualize these Three Segmentation Models:\n\n![fig6.png](https://s2.loli.net/2022/10/01/7dGgiYyM2OzFZ6q.png)","metadata":{}},{"cell_type":"markdown","source":"**Q: Why do I generate these images by a trained segmentation model?**\n\nA: For example, when the model was trained on the noisy model, its prediction accuracy dropped from 0.9 to 0.6. Does a reduction of 0.3 mean that the model has learned one-third less information? Or did it learn nothing? When I designed a new improved method, the prediction accuracy increased from 0.6 to 0.7. Does this 0.1 improvement mean that the model has learned more effective information? To solve these questions and peek inside segmentation networks, I propose this simple generative method to visualize the learning results of segmentation networks.\n\n**Q: How to evaluate the results of generation? Do I define a metric to evaluate them?**\n\nA: The way to evaluate the generated results is very subjective, that is, whether the generated results of the model conform to your own subjective cognition. Subjective perception is very unreliable and difficult to measure, but our perception of the world is based on it.\n\nThe best way to judge whether a person knows about cats is to let him draw them on paper, and the same is true of neural networks. As Richard Feynman said, what I cannot create, I do not understand.\n\n**Q: What's the significance of this visualization method? In what situations can this method be used?**\n\nA: Visualizing segmentation models allows me to observe the training results of the model from another perspective. It helps us correct some problems that are hard to find from numerical results, or to verify the effectiveness of an improvement.","metadata":{}},{"cell_type":"markdown","source":"## 3. Learning with Noisy Label\n\nI observe two interesting experimental phenomenons when models are trained with noisy label:\n\n1. <mark>The segmentation networks tend to first fit the clean pixel-level label during an “early-learning” phase, before eventually memorizing the false annotations[4]. <mark>(https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/347967)\n\n2. <mark>In “early-learning” phase, missed detections in training set show inconspicuously higher activation than true background. (Figure 5(c))<mark>\n\nAccording to these observations, I intent to find these missed detections in training set and avoid them to mislead models. When the model converges, original noisy label will be updated by trained model for the futher training. The overview of my process is shown in Figure 7. \n\n![截屏2022-10-01 上午7.52.17.png](https://s2.loli.net/2022/10/01/8jk1UFtdsCxPciQ.png)\n\nI adopt two methods separately, OEEM (Online Easy Example Mining[5]) and Energy Map which are applied in Figure 7(c), to extend the duration of “early-learning” phase and delay the duration of “overfitting” phase. OEEM encourages the network to focus on credible supervision signals rather than noisy signals by adjusting the weight of loss computation. Energy Map, which is proposed in this notebook, is a map that dynamically relabel the missed annotations.\n\n#### Results:\n\nThe test dice of different algorithms is compared and reported in Table 1 and Table 2.\n\n![截屏2022-10-01 上午7.52.53.png](https://s2.loli.net/2022/10/01/BFANtwEgJhni7vf.png)\n\nAccording to the results, it can be found that only the segmentation of the lungs becomes better when the measures to handle noisy label are taken. For other organs, my improvements have had little or no effect. ","metadata":{}},{"cell_type":"markdown","source":"#### Other details in my experiments:\n\n**No Model Ensemble**: To simplify the noisy label problem, I only adopted a single model with 5-folds data. The model can be found at https://www.kaggle.com/datasets/hengck23/hubmap-discuss-00?select=segformer-mit-b2. \n\n**No Auxiliary Loss Function**: Based on the original code, I just removed the auxiliary loss function. The existence of the auxiliary loss function can improve the performance of the model, but it is deleted for the convenience of analyzing the noisy label problem.\n\n**A Simple Strategy of Stain Normalization**: I also adopted Vahadane stain normalization (https://www.kaggle.com/code/gray98/stain-normalization-color-transfer) in training process to tranfer HPA image to \"Hubmap\" domain by choosing the only one image in test set as target.\n\n**No external data**: All experimental data are from this competition dataset.\n\n**No multi-class**: During training and inference, all FTUs from different organs were regarded as the same class, and different organs were not trained separately.","metadata":{}},{"cell_type":"markdown","source":"## 4. Limitation and Future work\n\n### Limitation：\n\n#### 1. The Quality of Visualization：\n\nThere is still a certain gap between the distribution of the images generated by the visualization method and the distribution of the real images, since I did not impose any prior constraints in the generation process. But it has been observed that when a segmentation model works well, the texture and shape of its generated images are very similar to the target we want.\n    \n\n#### 2. Futher improvements: \n\nBecause of the limited computing power of my device, I didn't have time to apply more skills to improve the score of the leaderboard.\n\n### Future work：\n\n#### 1. Generate more realistic images：\n\nThe current visualization method proposed in this work for visualizing segmentation models training with noisy label are not yet able to generate realistic images like real histopathology images. So, I'm still figuring out how to achieve this by borrowing some operations from other generative models, such as VAE and diffusion model.\n\n\n#### 2. Verify Generalization:\n\nIn future work, I will adopt my proposed method in more datasets to demonstrate its generalizability.","metadata":{}},{"cell_type":"markdown","source":"## 5. Acknowledgements\n\nI really appreciate that @hengck23 provided many novel and practical suggestions to me. Without his guidance, I could not have enough knowledge to finish the whole work. In addition, I would like to thank miXLab and NEU for all their support. Finally, thank the staff of Hubmap  competition for all their efforts.","metadata":{}},{"cell_type":"markdown","source":"## Reference:\n\n[1] Song, Hwanjun, et al. \"Learning from noisy labels with deep neural networks: A survey.\" IEEE Transactions on Neural Networks and Learning Systems (2022).\n\n[2] Cordeiro, Filipe R., and Gustavo Carneiro. \"A Survey on Deep Learning with Noisy Labels: How to train your model when you cannot trust on the annotations?.\" 2020 33rd SIBGRAPI conference on graphics, patterns and images (SIBGRAPI). IEEE, 2020.\n\n[3] Karimi, Davood, et al. \"Deep learning with noisy labels: Exploring techniques and remedies in medical image analysis.\" Medical Image Analysis 65 (2020): 101759.\n\n[4] Liu, Sheng, et al. \"Adaptive early-learning correction for segmentation from noisy annotations.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022.\n\n[5] Li, Yi, et al. \"Online Easy Example Mining for Weakly-Supervised Gland Segmentation from Histology Images.\" International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, Cham, 2022.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}