{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Loss Functions for Medical Image Segmentation for Keras & PyTorch\n# (의학적) 이미지 분할에 사용되는 로스 함수 모음 by keras and PyTorch\n---\n","metadata":{}},{"cell_type":"markdown","source":"![](https://miro.medium.com/max/1400/1*dWBAei84Ld3HtZPx257tHQ.png)","metadata":{}},{"cell_type":"markdown","source":"This kernel provides a reference library for some popular custom loss functions that you can easily import into your code.\n\nLoss functions define how neural network models calculate the overall error from their residuals for each training batch. This in turn affects how they adjust their internal weights when performing backpropagation, so the choice of loss function has a direct influence on model performance.\n\nThe default choice of loss function for segmentation and other classification tasks is Binary Cross-Entropy (BCE). In situations where a particular metric, like the Dice Coefficient or Intersection over Union (IoU), is being used to judge model performance, competitors will sometimes experiment with loss functions that derive from these metrics - typically in the form `1 - f(x)` where `f(x)` is the metric in question. \n\nThese functions cannot simply be written in NumPy, as they must operate on tensors that also have gradient parameters which need to be calculated throughout the model during backpropagation. Accordingly, loss functions must be written using backend functions from the respective model library. This is less complicated than it sounds! For example in Keras, you would simply use the same familiar mathematical functions, albeit using the Keras backend imported as `K`, i.e `K.sum()`. Gradient calculation is handled automatically by the model libraries, although (at least in PyTorch) you can define this manually if you are confident in your mathematical ability and wish to make alterations of your own.\n\nWith multi-class classification or segmentation, we sometimes use loss functions that calculate the average loss for each class, rather than calculating loss from the prediction tensor as a whole. This kernel is meant as a template reference for the basic code, so all examples calculate loss on the entire tensor, but it should be trivial for you to modify it for multi-class averaging. \n\nI hope this kernel will be of use to you, and any corrections or suggestions are welcome.","metadata":{}},{"cell_type":"markdown","source":"손실 함수는 딥러닝 기반 의료 영상 분할 방법에서 중요한 요소 중 하나입니다. \n지난 4년 동안 다양한 분할 작업을 위해 20개 이상의 손실 함수가 제안되었습니다.\n기존 손실 함수를 4가지 의미 있는 범주로 분류하는 체계적인 분류를 제시합니다.\n\n밑에서는 이미지 분할에 자주 사용되는 손실 함수에 대한 간단한 설명과 라이브러리 구현 코드를 제공합니다.\n\n손실함수는 신경망 모델이 각 훈련 배치에서 전체 오류를 계산하는 방법을 정의합니다.\n따라서 역전파를 수행할 때 내부 가중치가 조정되는 과정에 영향을 미치므로 전체 모델 성능에도 중요한 영향을 미칩니다.\n\n이미지 분할 및 분류 작업에 대한 기본적인 손실 함수는 BCE ( 이진 교차 엔트로피 ) 입니다.\nDice 계수 혹은 IoU 손실 함수가 사용되는 상황에서도 기본적인 베이스라인으로 BCE를 사용합니다.\n\n손실 함수는 역전파가 거슬러 올라가는 동안 모델 전체에서 계산되어야 하는 텐서에서 작동해야 하므로 Numpy로 간단하게 계산될 수 없습니다.\n해당 모델 라이브러리에서 제공하는 함수를 사용해야 합니다. Keras에서는 k.sum()과 같은 함수를 사용합니다.\n그라디언트 계산은 모델 라이브러리에 의해 자동으로 계산되지만, 필요한 경우 수동으로 정의할 수 있습니다.\n\n다중 클래스 분류 및 분할에서는 전체 손실대신에 각 클래스의 평균 손실을 계산하는 손실 함수를 사용합니다.\n","metadata":{}},{"cell_type":"code","source":"import numpy\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport keras\nimport keras.backend as K","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distributation-based loss ( 분포 기반 손실 함수 )\n\n","metadata":{}},{"cell_type":"markdown","source":"### Cross entropy (CE)\nis derived from Kullback-Leibler (KL) divergence, which is a measure of dissimilarity between two distributions. For common machine learning tasks, the data distribution is given by the training set, so H(p) will be a constant.","metadata":{}},{"cell_type":"markdown","source":"크로스 엔트로피는 Kullback-Leibler(KL) divergence라는 두 분포 사이의 비유사성에 대한 측정값입니다.\n일반적인 기계 학습에서는 데이터 분포는 훈련 데이터에의해 제공되므로 H(p)는 상수입니다.","metadata":{}},{"cell_type":"markdown","source":"![](https://miro.medium.com/max/1084/1*wemp48-PUxK2q3zkREb8Zg.png)","metadata":{}},{"cell_type":"markdown","source":"**Weighted cross entropy** is an extension to CE, which assign different weight to each class. In general, the un-presented classes will be allocated larger weights.\n\n**TopK loss**  aims to force networks to focus on hard samples during training.\n\n**Focal loss** adapts the standard CE to deal with extreme foreground-background class imbalance, where the loss assigned to well-classified examples is reduced.\n\n**Distance penalized CE loss** weights cross entropy by the distance map which is derived from the ground truth mask. It aims to guide the network’s focus towards hard-to-segment boundary regions.","metadata":{}},{"cell_type":"markdown","source":"**가중 교차 엔트로피**는 CE의 확장된 버전으로 각 클래스에 다른 가중치를 할당합니다. 일반적으로 제시되지 않은 클래스에는 더 큰 가중치가 할당됩니다.\n\n**TopK loss**는 네트워크가 훈련 중에 까다로운 샘플에 집중하도록 하는 것을 목표로 합니다.\n\n**초점 손실**은 잘 분류된 예제에 할당된 손실이 감소되는 극단적인 전경-배경 **클래스 불균형**을 처리하기 위해 표준 CE를 적용합니다.\n\n**거리 패널티 CE 손실**  가중치는 실제 마스크에서 파생된 거리 맵에 의해 엔트로피를 교차합니다. 분할하기 어려운 경계 영역으로 네트워크의 초점을 안내하는 것을 목표로 합니다.","metadata":{}},{"cell_type":"markdown","source":"### BCE-Dice Loss ( BCE-Dice 결합 손실 함수 )\n---\nThis loss combines Dice loss with the standard binary cross-entropy (BCE) loss that is generally the default for segmentation models. Combining the two methods allows for some diversity in the loss, while benefitting from the stability of BCE. The equation for multi-class BCE by itself will be familiar to anyone who has studied logistic regression:\n\n![](https://wikimedia.org/api/rest_v1/media/math/render/svg/80f87a71d3a616a0939f5360cec24d702d2593a2)","metadata":{}},{"cell_type":"markdown","source":"이 손실은 이미지 분할 모델 손실 함수로서 기본값인 표준 이진 교차 엔트로피(BCE) 손실과 주사위 손실을 결합합니다. 두 가지 방법을 결합하면 손실의 다양성을 허용하면서 BCE의 안정성을 활용할 수 있습니다. 다중 클래스 BCE에 대한 방정식 자체는 로지스틱 회귀를 연구자에게 친숙한 개념이다.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nclass DiceBCELoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(DiceBCELoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice_loss = 1 - (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        Dice_BCE = BCE + dice_loss\n        \n        return Dice_BCE","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\ndef DiceBCELoss(targets, inputs, smooth=1e-6):    \n       \n    #flatten label and prediction tensors\n    inputs = K.flatten(inputs)\n    targets = K.flatten(targets)\n    \n    BCE =  binary_crossentropy(targets, inputs)\n    intersection = K.sum(K.dot(targets, inputs))    \n    dice_loss = 1 - (2*intersection + smooth) / (K.sum(targets) + K.sum(inputs) + smooth)\n    Dice_BCE = BCE + dice_loss\n    \n    return Dice_BCE","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Region-based loss ( 영역 기반 손실 함수 )\n","metadata":{}},{"cell_type":"markdown","source":"영역 기반 손실 함수는 정답과 예측된 분할 간의 불일치를 최소화하거나 중첩 영역을 최대화하는 것을 목표로 합니다.\n\n**민감도-특이성(SS) 손실**은 민감도와 특이성의 평균 제곱 차이의 가중치 합입니다. 불균형 문제를 해결하기 위해 SS는 특이성에 더 높은 가중치를 부여합니다.\n\n**주사위 손실**은 가장 일반적으로 이미지 분할에 사용되는 평가 지표인 주사위 계수를 직접 최적화합니다.\n\nDice 손실과 유사한 **IoU 손실(Jaccard 손실**이라고도 함)은 세분화 메트릭을 직접 최적화하는 데에도 사용됩니다.\n\n**Tversky 손실**은 FN(False Negative) 및 FP(False Positive)에 서로 다른 가중치를 설정합니다. 이는 FN 및 FP에 대해 동일한 가중치를 사용하는 주사위 손실과 다릅니다.\n\n**일반화된 주사위 손실**은 각 클래스의 가중치가 레이블 빈도의 제곱에 반비례하는 Dice 손실의 다중 클래스 확장입니다.\n\n**초점 트버스키 손실**은 초점 손실의 개념을 적용하여 확률이 낮은 희귀한 케이스에 초점을 맞춥니다.\n\n**페널티 손실**은 일반화된 주사위 손실에서 위음성 및 위양성에 높은 페널티를 제공하는 손실 함수입니다.","metadata":{}},{"cell_type":"markdown","source":"# Dice Loss (주사위 손실 함수 )\n---\nThe Dice coefficient, or Dice-Sørensen coefficient, is a common metric for pixel segmentation that can also be modified to act as a loss function:\n\n주사위 계수는 가장 일반적으로 이미지 분할에 사용되는 평가 지표이며, 이에 최적화 된 손실 함수.\n\n![](https://wikimedia.org/api/rest_v1/media/math/render/svg/a80a97215e1afc0b222e604af1b2099dc9363d3b)","metadata":{}},{"cell_type":"code","source":"#PyTorch\nclass DiceLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(DiceLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        \n        return 1 - dice","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\ndef DiceLoss(targets, inputs, smooth=1e-6):\n    \n    #flatten label and prediction tensors\n    inputs = K.flatten(inputs)\n    targets = K.flatten(targets)\n    \n    intersection = K.sum(K.dot(targets, inputs))\n    dice = (2*intersection + smooth) / (K.sum(targets) + K.sum(inputs) + smooth)\n    return 1 - dice","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Jaccard/Intersection over Union (IoU) Loss -> IoU 손실 함수\n# (자카드 손실 함수)\n---\nThe IoU metric, or Jaccard Index, is similar to the Dice metric and is calculated as the ratio between the overlap of the positive instances between two sets, and their mutual combined values:\n\n![](https://wikimedia.org/api/rest_v1/media/math/render/svg/eaef5aa86949f49e7dc6b9c8c3dd8b233332c9e7)\n\nLike the Dice metric, it is a common means of evaluating the performance of pixel segmentation models.\n\nIoU 평가 지표는 주사위 지표와 비슷하며 교집합 대비 합집하의 비율로 계산된다.\n주사위 지표와 함께 이미지 분할 모델에서 가장 많이 쓰이는 지표(수단)이다.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nclass IoULoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(IoULoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #intersection is equivalent to True Positive count\n        #union is the mutually inclusive area of all labels & predictions \n        intersection = (inputs * targets).sum()\n        total = (inputs + targets).sum()\n        union = total - intersection \n        \n        IoU = (intersection + smooth)/(union + smooth)\n                \n        return 1 - IoU","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\ndef IoULoss(targets, inputs, smooth=1e-6):\n    \n    #flatten label and prediction tensors\n    inputs = K.flatten(inputs)\n    targets = K.flatten(targets)\n    \n    intersection = K.sum(K.dot(targets, inputs))\n    total = K.sum(targets) + K.sum(inputs)\n    union = total - intersection\n    \n    IoU = (intersection + smooth) / (union + smooth)\n    return 1 - IoU","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Focal Loss ( 초점 손실 함수 )\n---\nFocal Loss was introduced by *Lin et al* of Facebook AI Research in 2017 as a means of combatting extremely imbalanced datasets where positive cases were relatively rare. Their paper \"Focal Loss for Dense Object Detection\" is retrievable here: https://arxiv.org/abs/1708.02002. In practice, the researchers used an alpha-modified version of the function so I have included it in this implementation.","metadata":{}},{"cell_type":"markdown","source":"**Focal Loss**는 2017년 Facebook AI Research의 *Lin et al*에 의해 양성 사례가 상대적으로 드물었던 극도로 불균형한 데이터 세트를 해결하기 위한 수단으로 도입되었습니다. \n그들의 논문 \"Focal Loss for Dense Object Detection\"은 https://arxiv.org/abs/1708.02002에서 검색할 수 있습니다. 실제로 연구원들은 함수의 알파 수정 버전을 사용했기 때문에 이 구현에 포함시켰습니다.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nALPHA = 0.8\nGAMMA = 2\n\nclass FocalLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalLoss, self).__init__()\n\n    def forward(self, inputs, targets, alpha=ALPHA, gamma=GAMMA, smooth=1):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #first compute binary cross-entropy \n        BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        BCE_EXP = torch.exp(-BCE)\n        focal_loss = alpha * (1-BCE_EXP)**gamma * BCE\n                       \n        return focal_loss","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\nALPHA = 0.8\nGAMMA = 2\n\ndef FocalLoss(targets, inputs, alpha=ALPHA, gamma=GAMMA):    \n    \n    inputs = K.flatten(inputs)\n    targets = K.flatten(targets)\n    \n    BCE = K.binary_crossentropy(targets, inputs)\n    BCE_EXP = K.exp(-BCE)\n    focal_loss = K.mean(alpha * K.pow((1-BCE_EXP), gamma) * BCE)\n    \n    return focal_loss","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tversky Loss ( 트베르스키 손실 함수 )\n---\nThis loss was introduced in \"Tversky loss function for image segmentationusing 3D fully convolutional deep networks\", retrievable here: https://arxiv.org/abs/1706.05721. It was designed to optimise segmentation on imbalanced medical datasets by utilising constants that can adjust how harshly different types of error are penalised in the loss function. From the paper:\n\n>... in the case of α=β=0.5 the Tversky index simplifies to be the same as the Dice coefficient, which is also equal to the F1 score.  With α=β=1, Equation 2 produces Tanimoto coefficient, and setting α+β=1 produces the set of Fβ scores. Larger βs weigh recall higher than precision (by placing more emphasis on false negatives).\n\nTo summarise, this loss function is weighted by the constants 'alpha' and 'beta' that penalise false positives and false negatives respectively to a higher degree in the loss function as their value is increased. The beta constant in particular has applications in situations where models can obtain misleadingly positive performance via highly conservative prediction. You may want to experiment with different values to find the optimum. With alpha==beta==0.5, this loss becomes equivalent to Dice Loss.","metadata":{}},{"cell_type":"markdown","source":"트베르스키 손실은 https://arxiv.org/abs/1706.05721 에서 검색할 수 있는 \"3D 완전 컨볼루션 심층 네트워크를 사용하는 이미지 분할을 위한 Tversky 손실 함수\"에 소개되었습니다. 손실 함수에서 서로 다른 유형의 오류가 얼마나 심하게 처벌되는지를 조정할 수 있는 상수를 활용하여 불균형 의료 데이터 세트에 대한 분할을 최적화하도록 설계되었습니다. 논문에서:\n\n>... α=β=0.5의 경우 Tversky 지수는 F1 점수와 동일한 주사위 계수와 동일하도록 단순화됩니다. α=β=1일 때 방정식 2는 Tanimoto 계수를 생성하고 α+β=1로 설정하면 Fβ 점수 세트를 생성합니다. β가 클수록 정밀도보다 재현율이 더 높습니다(위음성에 더 중점을 둠).\n\n요약하자면, 이 손실 함수는 값이 증가함에 따라 손실 함수에서 더 높은 정도로 위양성 및 위음성 각각에 페널티를 주는 상수 '알파' 및 '베타'에 의해 가중치가 부여됩니다. 특히 베타 상수는 모델이 매우 보수적인 예측을 통해 오도할 정도로 긍정적인 성능을 얻을 수 있는 상황에서 적용됩니다. 최적의 값을 찾기 위해 다양한 값으로 실험할 수 있습니다. alpha==beta==0.5에서 이 손실은 Dice Loss와 동일해집니다.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nALPHA = 0.5\nBETA = 0.5\n\nclass TverskyLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(TverskyLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1, alpha=ALPHA, beta=BETA):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #True Positives, False Positives & False Negatives\n        TP = (inputs * targets).sum()    \n        FP = ((1-targets) * inputs).sum()\n        FN = (targets * (1-inputs)).sum()\n       \n        Tversky = (TP + smooth) / (TP + alpha*FP + beta*FN + smooth)  \n        \n        return 1 - Tversky","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\nALPHA = 0.5\nBETA = 0.5\n\ndef TverskyLoss(targets, inputs, alpha=ALPHA, beta=BETA, smooth=1e-6):\n        \n        #flatten label and prediction tensors\n        inputs = K.flatten(inputs)\n        targets = K.flatten(targets)\n        \n        #True Positives, False Positives & False Negatives\n        TP = K.sum((inputs * targets))\n        FP = K.sum(((1-targets) * inputs))\n        FN = K.sum((targets * (1-inputs)))\n       \n        Tversky = (TP + smooth) / (TP + alpha*FP + beta*FN + smooth)  \n        \n        return 1 - Tversky","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Focal Tversky Loss ( 초점 트베르스키 손실 함수 )\n---\n\nA variant on the Tversky loss that also includes the gamma modifier from Focal Loss.\n\n트베르스키 손실의 변형으로서 초점 손실의 감마 수정자를 포함한다.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nALPHA = 0.5\nBETA = 0.5\nGAMMA = 1\n\nclass FocalTverskyLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalTverskyLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1, alpha=ALPHA, beta=BETA, gamma=GAMMA):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #True Positives, False Positives & False Negatives\n        TP = (inputs * targets).sum()    \n        FP = ((1-targets) * inputs).sum()\n        FN = (targets * (1-inputs)).sum()\n        \n        Tversky = (TP + smooth) / (TP + alpha*FP + beta*FN + smooth)  \n        FocalTversky = (1 - Tversky)**gamma\n                       \n        return FocalTversky","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\nALPHA = 0.5\nBETA = 0.5\nGAMMA = 1\n\ndef FocalTverskyLoss(targets, inputs, alpha=ALPHA, beta=BETA, gamma=GAMMA, smooth=1e-6):\n    \n        #flatten label and prediction tensors\n        inputs = K.flatten(inputs)\n        targets = K.flatten(targets)\n        \n        #True Positives, False Positives & False Negatives\n        TP = K.sum((inputs * targets))\n        FP = K.sum(((1-targets) * inputs))\n        FN = K.sum((targets * (1-inputs)))\n               \n        Tversky = (TP + smooth) / (TP + alpha*FP + beta*FN + smooth)  \n        FocalTversky = K.pow((1 - Tversky), gamma)\n        \n        return FocalTversky","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Lovasz Hinge Loss ( 로바즈 힌지 손실 함수 )\n---\nThis complex loss function was introduced by Berman, Triki and Blaschko in their paper \"The Lovasz-Softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks\", retrievable here: https://arxiv.org/abs/1705.08790. It is designed to optimise the Intersection over Union score for semantic segmentation, particularly for multi-class instances. Specifically, it sorts predictions by their error before calculating cumulatively how each error affects the IoU score. This gradient vector is then multiplied with the initial error vector to penalise most strongly the predictions that decreased the IoU score the most. This procedure is detailed by [jeandebleu](https://www.kaggle.com/jeandebleau) in his excellent summary [here](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/67791).\n\nThis code is taken directly from the author's github repo here: https://github.com/bermanmaxim/LovaszSoftmax and all credit is to them.\n\nIn this kernel I have implemented the flat variant that uses reshaped rank-1 tensors as inputs for PyTorch. You can modify it accordingly with the dimensions and class number of your data as needed. This code takes raw logits so ensure your model does not contain an activation layer prior to the loss calculation.\n\nI have hidden the researchers' own code below for brevity; simply load it into your kernel for the losses to function. In the case of their tensorflow implementation, I am still working to make it compatible with Keras. There are differences between the Tensorflow and Keras function libraries that complicate this.","metadata":{}},{"cell_type":"markdown","source":"로바즈 힌지 손실 함수는 Berman, Triki 및 Blaschko의 논문 \"The Lovasz-Softmax loss: A tractable surrogate for optimize of cross-over-union measure in neural network\"에서 소개되었으며 여기에서 검색할 수 있습니다: https://arxiv.org/abs/1705.08790. \n\n특히 다중 분류의 경우 의미론적 분할을 위해 Intersection over Union 점수를 최적화하도록 설계되었습니다. 특히 각 오류가 IoU 점수에 미치는 영향을 누적 계산하기 전에 오류를 기준으로 예측을 정렬합니다. 그런 다음 이 그래디언트 벡터에 초기 오류 벡터를 곱하여 IoU 점수를 가장 많이 감소시킨 예측에 가장 강력한 페널티를 부여합니다. 이 절차는 jeanderbleu의 [탁월한 요약](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/67791) 에서 자세히 설명되어 있습니다.\n\n이 코드는 https://github.com/bermanmaxim/LovaszSoftmax 작성자의 github에서 직접 가져왔으며 모든 크레딧은 이들에게 있습니다.\n\n원본 손실 함수 외에도 PyTorch에 대한 입력으로 재구성된 랭크 1 텐서를 사용하는 플랫 변형을 구현했습니다. 필요에 따라 데이터의 차원과 클래스 번호에 따라 수정할 수 있습니다. 이 코드는 원시 로짓을 사용하므로 손실 계산 전에 모델에 활성화 레이어가 포함되어 있지 않은지 확인해야 합니다.\n\n간결함을 위해 연구원의 코드를 아래에 숨겼습니다. 손실이 작동하려면 커널에 로드하기만 하면 됩니다. tensorflow 구현의 경우 여전히 Keras와 호환되도록 노력하고 있습니다. Tensorflow와 Keras 함수 라이브러리 사이에는 이를 복잡하게 만드는 차이점이 있습니다.","metadata":{}},{"cell_type":"code","source":"#PyTorch\n\ndef flatten_binary_scores(scores, labels, ignore=None):\n    \"\"\"\n    Flattens predictions in the batch (binary case)\n    Remove labels equal to 'ignore'\n    \"\"\"\n    scores = scores.view(-1)\n    labels = labels.view(-1)\n    if ignore is None:\n        return scores, labels\n    valid = (labels != ignore)\n    vscores = scores[valid]\n    vlabels = labels[valid]\n    return vscores, vlabels\n\ndef lovasz_grad(gt_sorted):\n    \"\"\"\n    Computes gradient of the Lovasz extension w.r.t sorted errors\n    See Alg. 1 in paper\n    \"\"\"\n    p = len(gt_sorted)\n    gts = gt_sorted.sum()\n    intersection = gts - gt_sorted.float().cumsum(0)\n    union = gts + (1 - gt_sorted).float().cumsum(0)\n    jaccard = 1. - intersection / union\n    if p > 1: # cover 1-pixel case\n        jaccard[1:p] = jaccard[1:p] - jaccard[0:-1]\n    return jaccard\n\ndef lovasz_hinge(logits, labels, per_image=True, ignore=None):\n    \"\"\"\n    Binary Lovasz hinge loss\n      logits: [B, H, W] Variable, logits at each pixel (between -\\infty and +\\infty)\n      labels: [B, H, W] Tensor, binary ground truth masks (0 or 1)\n      per_image: compute the loss per image instead of per batch\n      ignore: void class id\n    \"\"\"\n    if per_image:\n        loss = mean(lovasz_hinge_flat(*flatten_binary_scores(log.unsqueeze(0), lab.unsqueeze(0), ignore))\n                          for log, lab in zip(logits, labels))\n    else:\n        loss = lovasz_hinge_flat(*flatten_binary_scores(logits, labels, ignore))\n    return loss\n\ndef lovasz_hinge_flat(logits, labels):\n    \"\"\"\n    Binary Lovasz hinge loss\n      logits: [P] Variable, logits at each prediction (between -\\infty and +\\infty)\n      labels: [P] Tensor, binary ground truth labels (0 or 1)\n      ignore: label to ignore\n    \"\"\"\n    if len(labels) == 0:\n        # only void pixels, the gradients should be 0\n        return logits.sum() * 0.\n    signs = 2. * labels.float() - 1.\n    errors = (1. - logits * Variable(signs))\n    errors_sorted, perm = torch.sort(errors, dim=0, descending=True)\n    perm = perm.data\n    gt_sorted = labels[perm]\n    grad = lovasz_grad(gt_sorted)\n    loss = torch.dot(F.relu(errors_sorted), Variable(grad))\n    return loss\n\n#=====\n#Multi-class Lovasz loss\n#=====\n\ndef lovasz_softmax(probas, labels, classes='present', per_image=False, ignore=None):\n    \"\"\"\n    Multi-class Lovasz-Softmax loss\n      probas: [B, C, H, W] Variable, class probabilities at each prediction (between 0 and 1).\n              Interpreted as binary (sigmoid) output with outputs of size [B, H, W].\n      labels: [B, H, W] Tensor, ground truth labels (between 0 and C - 1)\n      classes: 'all' for all, 'present' for classes present in labels, or a list of classes to average.\n      per_image: compute the loss per image instead of per batch\n      ignore: void class labels\n    \"\"\"\n    if per_image:\n        loss = mean(lovasz_softmax_flat(*flatten_probas(prob.unsqueeze(0), lab.unsqueeze(0), ignore), classes=classes)\n                          for prob, lab in zip(probas, labels))\n    else:\n        loss = lovasz_softmax_flat(*flatten_probas(probas, labels, ignore), classes=classes)\n    return loss\n\n\ndef lovasz_softmax_flat(probas, labels, classes='present'):\n    \"\"\"\n    Multi-class Lovasz-Softmax loss\n      probas: [P, C] Variable, class probabilities at each prediction (between 0 and 1)\n      labels: [P] Tensor, ground truth labels (between 0 and C - 1)\n      classes: 'all' for all, 'present' for classes present in labels, or a list of classes to average.\n    \"\"\"\n    if probas.numel() == 0:\n        # only void pixels, the gradients should be 0\n        return probas * 0.\n    C = probas.size(1)\n    losses = []\n    class_to_sum = list(range(C)) if classes in ['all', 'present'] else classes\n    for c in class_to_sum:\n        fg = (labels == c).float() # foreground for class c\n        if (classes is 'present' and fg.sum() == 0):\n            continue\n        if C == 1:\n            if len(classes) > 1:\n                raise ValueError('Sigmoid output possible only with 1 class')\n            class_pred = probas[:, 0]\n        else:\n            class_pred = probas[:, c]\n        errors = (Variable(fg) - class_pred).abs()\n        errors_sorted, perm = torch.sort(errors, 0, descending=True)\n        perm = perm.data\n        fg_sorted = fg[perm]\n        losses.append(torch.dot(errors_sorted, Variable(lovasz_grad(fg_sorted))))\n    return mean(losses)","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#PyTorch\nclass LovaszHingeLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(LovaszHingeLoss, self).__init__()\n\n    def forward(self, inputs, targets):\n        inputs = F.sigmoid(inputs)    \n        Lovasz = lovasz_hinge(inputs, targets, per_image=False)                       \n        return Lovasz","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport numpy as np\n\n\ndef lovasz_grad(gt_sorted):\n    \"\"\"\n    Computes gradient of the Lovasz extension w.r.t sorted errors\n    See Alg. 1 in paper\n    \"\"\"\n    gts = tf.reduce_sum(gt_sorted)\n    intersection = gts - tf.cumsum(gt_sorted)\n    union = gts + tf.cumsum(1. - gt_sorted)\n    jaccard = 1. - intersection / union\n    jaccard = tf.concat((jaccard[0:1], jaccard[1:] - jaccard[:-1]), 0)\n    return jaccard\n\n\n# --------------------------- BINARY LOSSES ---------------------------\n\n\ndef lovasz_hinge(logits, labels, per_image=True, ignore=None):\n    \"\"\"\n    Binary Lovasz hinge loss\n      logits: [B, H, W] Variable, logits at each pixel (between -\\infty and +\\infty)\n      labels: [B, H, W] Tensor, binary ground truth masks (0 or 1)\n      per_image: compute the loss per image instead of per batch\n      ignore: void class id\n    \"\"\"\n    if per_image:\n        def treat_image(log_lab):\n            log, lab = log_lab\n            log, lab = tf.expand_dims(log, 0), tf.expand_dims(lab, 0)\n            log, lab = flatten_binary_scores(log, lab, ignore)\n            return lovasz_hinge_flat(log, lab)\n        losses = tf.map_fn(treat_image, (logits, labels), dtype=tf.float32)\n        loss = tf.reduce_mean(losses)\n    else:\n        loss = lovasz_hinge_flat(*flatten_binary_scores(logits, labels, ignore))\n    return loss\n\n\ndef lovasz_hinge_flat(logits, labels):\n    \"\"\"\n    Binary Lovasz hinge loss\n      logits: [P] Variable, logits at each prediction (between -\\infty and +\\infty)\n      labels: [P] Tensor, binary ground truth labels (0 or 1)\n      ignore: label to ignore\n    \"\"\"\n\n    def compute_loss():\n        labelsf = tf.cast(labels, logits.dtype)\n        signs = 2. * labelsf - 1.\n        errors = 1. - logits * tf.stop_gradient(signs)\n        errors_sorted, perm = tf.nn.top_k(errors, k=tf.shape(errors)[0], name=\"descending_sort\")\n        gt_sorted = tf.gather(labelsf, perm)\n        grad = lovasz_grad(gt_sorted)\n        loss = tf.tensordot(tf.nn.relu(errors_sorted), grad, 1, name=\"loss_non_void\")\n        return loss\n\n    # deal with the void prediction case (only void pixels)\n    loss = tf.cond(tf.equal(tf.shape(logits)[0], 0),\n                   lambda: tf.reduce_sum(logits) * 0.,\n                   compute_loss,\n                   strict=True,\n                   name=\"loss\"\n                   )\n    return loss\n\n\ndef flatten_binary_scores(scores, labels, ignore=None):\n    \"\"\"\n    Flattens predictions in the batch (binary case)\n    Remove labels equal to 'ignore'\n    \"\"\"\n    scores = tf.reshape(scores, (-1,))\n    labels = tf.reshape(labels, (-1,))\n    if ignore is None:\n        return scores, labels\n    valid = tf.not_equal(labels, ignore)\n    vscores = tf.boolean_mask(scores, valid, name='valid_scores')\n    vlabels = tf.boolean_mask(labels, valid, name='valid_labels')\n    return vscores, vlabels","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\n# not working yet\n# def LovaszHingeLoss(inputs, targets):\n#     return lovasz_hinge_loss(inputs, targets)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Combo Loss ( 콤보 손실 함수 )\n---\nThis loss was introduced by Taghanaki et al in their paper \"Combo loss: Handling input and output imbalance in multi-organ segmentation\", retrievable here: https://arxiv.org/abs/1805.02798. Combo loss is a combination of Dice Loss and a modified Cross-Entropy function that, like Tversky loss, has additional constants which penalise either false positives or false negatives more respectively.\n\nSince my GPU quota has run out this week, as of V16 these functions have not been tested, so please leave any debugging notes in the comments section below.","metadata":{}},{"cell_type":"markdown","source":"콤보 손실은 Taghanaki 등의 논문 \"Combo loss: Handling input and output 불균형 in multi-organ segmentation\"에서 소개되었으며 여기에서 검색할 수 있습니다: https://arxiv.org/abs/1805.02798. \n\n콤보 손실은 Tversky 손실과 마찬가지로 거짓 긍정 또는 거짓 부정에 각각 페널티를 주는 추가 상수가 있는 수정된 교차 엔트로피 함수와 주사위 손실의 조합입니다.( 결합 함수 )","metadata":{}},{"cell_type":"code","source":"#PyTorch\nALPHA = 0.5 # < 0.5 penalises FP more, > 0.5 penalises FN more\nCE_RATIO = 0.5 #weighted contribution of modified CE loss compared to Dice loss\n\nclass ComboLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(ComboLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1, alpha=ALPHA, beta=BETA, eps=1e-9):\n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #True Positives, False Positives & False Negatives\n        intersection = (inputs * targets).sum()    \n        dice = (2. * intersection + smooth) / (inputs.sum() + targets.sum() + smooth)\n        \n        inputs = torch.clamp(inputs, eps, 1.0 - eps)       \n        out = - (ALPHA * ((targets * torch.log(inputs)) + ((1 - ALPHA) * (1.0 - targets) * torch.log(1.0 - inputs))))\n        weighted_ce = out.mean(-1)\n        combo = (CE_RATIO * weighted_ce) - ((1 - CE_RATIO) * dice)\n        \n        return combo","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keras\nALPHA = 0.5 # < 0.5 penalises FP more, > 0.5 penalises FN more\nCE_RATIO = 0.5 #weighted contribution of modified CE loss compared to Dice loss\n\ndef Combo_loss(targets, inputs, eps=1e-9):\n    targets = K.flatten(targets)\n    inputs = K.flatten(inputs)\n    \n    intersection = K.sum(targets * inputs)\n    dice = (2. * intersection + smooth) / (K.sum(targets) + K.sum(inputs) + smooth)\n    inputs = K.clip(inputs, eps, 1.0 - eps)\n    out = - (ALPHA * ((targets * K.log(inputs)) + ((1 - ALPHA) * (1.0 - targets) * K.log(1.0 - inputs))))\n    weighted_ce = K.mean(out, axis=-1)\n    combo = (CE_RATIO * weighted_ce) - ((1 - CE_RATIO) * dice)\n    \n    return combo","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Boundary-based loss ( 경계기반 손실 함수 )\n---\nBoundary-based loss, a recent new type of loss function, aims to minimize the distance between ground truth and predicted segmentation. Usually, to make the training more robust, boundary-based loss functions are used with region-based loss.\n\n\n최근 새로운 유형의 손실 함수인 경계 기반 손실은 정답과 예측 간의 거리를 최소화하는 것을 목표로 합니다. \n일반적으로 더 나은 결과를 위해서 경계 기반 손실 함수가 영역 기반 손실과 함께 사용됩니다.","metadata":{}},{"cell_type":"markdown","source":"# Compound loss ( 혼합 손실 함수 )\n---\nBy summing over different types of loss functions, we can obtain several compound loss functions, such as Dice+CE, Dice+TopK, Dice+Focal and so on.\nAll the methioned loss functions can be usd in a plug-and-play way<BR>\n\n\n여러종류의 손실함수를 결합함으로서, 우리는 혼합된(결합된) 손실 함수를 얻을 수 있습니다. (ex Dice+CE, Dice + Focal, Dice + IOU )...\n모든 종류의 손실 함수들은 플러그 방식으로(조립 및 결합) 사용할 수 있습니다.","metadata":{}},{"cell_type":"markdown","source":"# Usage Tips ( 유용한 팁 )\n---\n\nIn my experience testing and debugging these losses, I have some observations that may be useful to beginners experimenting with different loss functions. These are not rules that are set in stone; they are simply my findings and your results may vary.\n\n* Tversky and Focal-Tversky loss benefit from very low learning rates, of the order 5e-5 to 1e-4. They would not see much improvement in my kernels until around 7-10 epochs, upon which performance would improve significantly.\n\n* In general, if a loss function does not appear to be working well (or at all), experiment with modifying the learning rate before moving on to other options.\n\n* You can easily create your own loss functions by combining any of the above with Binary Cross-Entropy or any combination of other losses. Bear in mind that loss is calculated for every batch, so more complex losses will increase runtime.\n\n* Care must be taken when writing loss functions for PyTorch. If you call a function to modify the inputs that doesn't entirely use PyTorch's numerical methods, the tensor will 'detach' from the the graph that maps it back through the neural network for the purposes of backpropagation, making the loss function unusable. \n\nI hope this kernel is of use to you. Good luck with your work!","metadata":{}},{"cell_type":"markdown","source":"이러한 손실을 테스트하고 디버깅한 경험에서 다양한 손실 함수를 실험하는 초보자에게 유용할 수 있는 몇 가지 관찰 사항이 있습니다. 이 내용들은 단순히 나의 발견이며 당신의 결과는 다를 수 있습니다.\n\n* Tversky 및 Focal-Tversky 손실은 5e-5에서 1e-4까지의 매우 낮은 학습률의 이점이 있습니다. 그들은 성능이 크게 향상되는 약 7-10 epoch가 될 때까지 내 커널에서 많은 개선을 보지 못할 것입니다.\n\n* 일반적으로 손실 함수가 제대로 작동하지 않거나 전혀 작동하지 않는 것으로 보이면 다른 옵션으로 이동하기 전에 학습률을 수정하는 실험을 하십시오.\n\n* 위의 항목을 이진 교차 엔트로피 또는 다른 손실의 조합과 결합하여 자신의 손실 함수를 쉽게 만들 수 있습니다. 손실은 모든 배치에 대해 계산되므로 손실이 더 복잡하면 런타임이 늘어납니다.\n\n* PyTorch에 대한 손실 함수를 작성할 때 주의해야 합니다. PyTorch의 수치적 방법을 완전히 사용하지 않는 입력을 수정하는 함수를 호출하면 텐서는 역전파 목적으로 신경망을 통해 다시 매핑하는 그래프에서 '분리'되어 손실 함수를 사용할 수 없게 만듭니다. 이에 대한 논의는 여기에서 볼 수 있습니다.\n\n이 커널이 당신에게 유용하기를 바랍니다. 작업에 행운을 빕니다!","metadata":{}}]}