{
  "id": 208325,
  "title": "[TF.Keras]: Implementation of Joint Optimization Framework",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/208325",
  "author_name": "",
  "post_date": "2021-01-02T23:29:05.009653600Z",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Paper: <a href=\"https://arxiv.org/abs/1803.11364\" target=\"_blank\">Joint Optimization Framework for Learning with Noisy Labels</a></p>\n<p>Here is the implementation of this loss function, hope it helps.</p>\n<pre><code>import tensorflow as tf \nfrom tensorflow.keras import backend as K\n\nclass JointOptimization(tf.losses.Loss):\n    def __init__(self, num_classes=5):\n        '''\n        Paper: https://arxiv.org/abs/1803.11364\n        '''\n        super(JointOptimization, self).__init__()\n        self.num_classes = num_classes\n        self.prob = np.ones(self.num_classes, \n                           dtype=np.float32)/self.num_classes\n\n    def call(self, y_true, y_pred):\n        y_pred_avg = K.mean(y_pred, axis=0)\n        lp = - K.sum(K.log(y_pred_avg)*self.prob)\n        le = K.categorical_crossentropy(y_pred, y_pred)\n        return K.categorical_crossentropy(y_true, y_pred) + \\\n                  1.2 * lp + 0.8 * le\n</code></pre>\n<p>Abstract:</p>\n<pre><code>Deep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance degradation because DNNs easily overfit to noisy labels. To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can **correct labels during training** by **alternating update of network parameters and labels**. We conduct experiments on the noisy CIFAR-10 datasets and the Clothing1M dataset. The results indicate that our approach significantly outperforms other state-of-the-art methods.\n</code></pre>",
  "messages": [
    {
      "id": "1136289",
      "postDate": "01/02/2021 23:29:05",
      "content": "<p>Paper: <a href=\"https://arxiv.org/abs/1803.11364\" target=\"_blank\">Joint Optimization Framework for Learning with Noisy Labels</a></p>\n<p>Here is the implementation of this loss function, hope it helps.</p>\n<pre><code>import tensorflow as tf \nfrom tensorflow.keras import backend as K\n\nclass JointOptimization(tf.losses.Loss):\n    def __init__(self, num_classes=5):\n        '''\n        Paper: https://arxiv.org/abs/1803.11364\n        '''\n        super(JointOptimization, self).__init__()\n        self.num_classes = num_classes\n        self.prob = np.ones(self.num_classes, \n                           dtype=np.float32)/self.num_classes\n\n    def call(self, y_true, y_pred):\n        y_pred_avg = K.mean(y_pred, axis=0)\n        lp = - K.sum(K.log(y_pred_avg)*self.prob)\n        le = K.categorical_crossentropy(y_pred, y_pred)\n        return K.categorical_crossentropy(y_true, y_pred) + \\\n                  1.2 * lp + 0.8 * le\n</code></pre>\n<p>Abstract:</p>\n<pre><code>Deep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance degradation because DNNs easily overfit to noisy labels. To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can **correct labels during training** by **alternating update of network parameters and labels**. We conduct experiments on the noisy CIFAR-10 datasets and the Clothing1M dataset. The results indicate that our approach significantly outperforms other state-of-the-art methods.\n</code></pre>",
      "rawMarkdown": "Paper: [Joint Optimization Framework for Learning with Noisy Labels](https://arxiv.org/abs/1803.11364)\n\nHere is the implementation of this loss function, hope it helps.\n\n```python\nimport tensorflow as tf \nfrom tensorflow.keras import backend as K\n\nclass JointOptimization(tf.losses.Loss):\n    def __init__(self, num_classes=5):\n        '''\n        Paper: https://arxiv.org/abs/1803.11364\n        '''\n        super(JointOptimization, self).__init__()\n        self.num_classes = num_classes\n        self.prob = np.ones(self.num_classes, \n                           dtype=np.float32)/self.num_classes\n\n    def call(self, y_true, y_pred):\n        y_pred_avg = K.mean(y_pred, axis=0)\n        lp = - K.sum(K.log(y_pred_avg)*self.prob)\n        le = K.categorical_crossentropy(y_pred, y_pred)\n        return K.categorical_crossentropy(y_true, y_pred) + \\\n                  1.2 * lp + 0.8 * le\n```\n\nAbstract:\n\n```\nDeep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance degradation because DNNs easily overfit to noisy labels. To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can **correct labels during training** by **alternating update of network parameters and labels**. We conduct experiments on the noisy CIFAR-10 datasets and the Clothing1M dataset. The results indicate that our approach significantly outperforms other state-of-the-art methods.\n```",
      "votes": null
    },
    {
      "id": "1136329",
      "postDate": "01/03/2021 00:45:31",
      "content": "<blockquote>\n  <p>inaccurate labels that are termed as noisy labels.</p>\n</blockquote>\n<p>Finally, i understand what <strong>noisy labels</strong> means, thanks for share.</p>",
      "rawMarkdown": "> inaccurate labels that are termed as noisy labels.\n\nFinally, i understand what **noisy labels** means, thanks for share.",
      "votes": null
    },
    {
      "id": "1136351",
      "postDate": "01/03/2021 01:51:49",
      "content": "<p>Wait, so how are we supposed to use this in our code? Take JointOptimization as a loss?</p>",
      "rawMarkdown": "Wait, so how are we supposed to use this in our code? Take JointOptimization as a loss?",
      "votes": null
    },
    {
      "id": "1136359",
      "postDate": "01/03/2021 02:07:12",
      "content": "<p>I only tried <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324\" target=\"_blank\">this</a> one so far. However, as per your concern </p>\n<pre><code>jot = JointOptimization()\nmodel.compile(loss=jot, ..)\n</code></pre>",
      "rawMarkdown": "I only tried [this](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324) one so far. However, as per your concern \n\n```\njot = JointOptimization()\nmodel.compile(loss=jot, ..)\n```",
      "votes": null
    },
    {
      "id": "1136363",
      "postDate": "01/03/2021 02:14:20",
      "content": "<p>Thanks for sharing! How were your results so far? </p>",
      "rawMarkdown": "Thanks for sharing! How were your results so far?",
      "votes": null
    },
    {
      "id": "1136372",
      "postDate": "01/03/2021 02:30:24",
      "content": "<p>As I mentioned, I didn't try this loss function yet. But using Symmetric Cross-Entropy Loss, implementation <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324\" target=\"_blank\">here</a>, did help to boost my local cv, from 0.90 to 0.901</p>",
      "rawMarkdown": "As I mentioned, I didn't try this loss function yet. But using Symmetric Cross-Entropy Loss, implementation [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324), did help to boost my local cv, from 0.90 to 0.901",
      "votes": null
    },
    {
      "id": "1136385",
      "postDate": "01/03/2021 03:04:27",
      "content": "<p>Sorry, I misunderstood. Thanks a lot for clarifying!</p>",
      "rawMarkdown": "Sorry, I misunderstood. Thanks a lot for clarifying!",
      "votes": null
    },
    {
      "id": "1136924",
      "postDate": "01/03/2021 14:36:05",
      "content": "<p>Beware that Joint Optimization is End to End Framework and not just a loss function.   Because there is an epoch wise update phase where the (wrong) labels need to be \"corrected\" in semi-supervised manner<br>\nIt's even said in the abstract :</p>\n<blockquote>\n  <p>To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can <strong>correct labels during training</strong> by <strong>alternating update of network parameters and labels</strong>.</p>\n</blockquote>\n<p>To implement this in TF, you will need custom training loop or may be at least highly customized callback. </p>",
      "rawMarkdown": "Beware that Joint Optimization is End to End Framework and not just a loss function.   Because there is an epoch wise update phase where the (wrong) labels need to be \"corrected\" in semi-supervised manner\nIt's even said in the abstract :\n\n>  To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can **correct labels during training** by **alternating update of network parameters and labels**.\n\n\nTo implement this in TF, you will need custom training loop or may be at least highly customized callback.",
      "votes": null
    },
    {
      "id": "1137024",
      "postDate": "01/03/2021 15:55:16",
      "content": "<p>Thanks for sharing! I was wondering why I couldn't see any improvement in the model after using just the loss function.</p>",
      "rawMarkdown": "Thanks for sharing! I was wondering why I couldn't see any improvement in the model after using just the loss function.",
      "votes": null
    },
    {
      "id": "1139139",
      "postDate": "01/05/2021 07:40:29",
      "content": "<p>Thanks for sharing!<br>\n<code>\nle = K.categorical_crossentropy(y_pred, y_pred)\n</code><br>\nI'm confused by this line. What does it do?<br>\nAnd also does \"self.prob\" need to be update after each iteration as well?</p>",
      "rawMarkdown": "Thanks for sharing!\n`\nle = K.categorical_crossentropy(y_pred, y_pred)\n`\nI'm confused by this line. What does it do?\nAnd also does \"self.prob\" need to be update after each iteration as well?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1136329,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "01/03/2021 00:45:31",
      "content": "<blockquote>\n  <p>inaccurate labels that are termed as noisy labels.</p>\n</blockquote>\n<p>Finally, i understand what <strong>noisy labels</strong> means, thanks for share.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1136351,
      "author_name": "junyingsg",
      "author_url": "",
      "post_date": "01/03/2021 01:51:49",
      "content": "<p>Wait, so how are we supposed to use this in our code? Take JointOptimization as a loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1136359,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "01/03/2021 02:07:12",
          "content": "<p>I only tried <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324\" target=\"_blank\">this</a> one so far. However, as per your concern </p>\n<pre><code>jot = JointOptimization()\nmodel.compile(loss=jot, ..)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136363,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/03/2021 02:14:20",
          "content": "<p>Thanks for sharing! How were your results so far? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136372,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "01/03/2021 02:30:24",
          "content": "<p>As I mentioned, I didn't try this loss function yet. But using Symmetric Cross-Entropy Loss, implementation <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324\" target=\"_blank\">here</a>, did help to boost my local cv, from 0.90 to 0.901</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136385,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/03/2021 03:04:27",
          "content": "<p>Sorry, I misunderstood. Thanks a lot for clarifying!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136924,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/03/2021 14:36:05",
      "content": "<p>Beware that Joint Optimization is End to End Framework and not just a loss function.   Because there is an epoch wise update phase where the (wrong) labels need to be \"corrected\" in semi-supervised manner<br>\nIt's even said in the abstract :</p>\n<blockquote>\n  <p>To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can <strong>correct labels during training</strong> by <strong>alternating update of network parameters and labels</strong>.</p>\n</blockquote>\n<p>To implement this in TF, you will need custom training loop or may be at least highly customized callback. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1137024,
          "author_name": "yerramvarun",
          "author_url": "",
          "post_date": "01/03/2021 15:55:16",
          "content": "<p>Thanks for sharing! I was wondering why I couldn't see any improvement in the model after using just the loss function.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139139,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "01/05/2021 07:40:29",
      "content": "<p>Thanks for sharing!<br>\n<code>\nle = K.categorical_crossentropy(y_pred, y_pred)\n</code><br>\nI'm confused by this line. What does it do?<br>\nAnd also does \"self.prob\" need to be update after each iteration as well?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1136289": "Paper: [Joint Optimization Framework for Learning with Noisy Labels](https://arxiv.org/abs/1803.11364)\n\nHere is the implementation of this loss function, hope it helps.\n\n```python\nimport tensorflow as tf \nfrom tensorflow.keras import backend as K\n\nclass JointOptimization(tf.losses.Loss):\n    def __init__(self, num_classes=5):\n        '''\n        Paper: https://arxiv.org/abs/1803.11364\n        '''\n        super(JointOptimization, self).__init__()\n        self.num_classes = num_classes\n        self.prob = np.ones(self.num_classes, \n                           dtype=np.float32)/self.num_classes\n\n    def call(self, y_true, y_pred):\n        y_pred_avg = K.mean(y_pred, axis=0)\n        lp = - K.sum(K.log(y_pred_avg)*self.prob)\n        le = K.categorical_crossentropy(y_pred, y_pred)\n        return K.categorical_crossentropy(y_true, y_pred) + \\\n                  1.2 * lp + 0.8 * le\n```\n\nAbstract:\n\n```\nDeep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance degradation because DNNs easily overfit to noisy labels. To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can **correct labels during training** by **alternating update of network parameters and labels**. We conduct experiments on the noisy CIFAR-10 datasets and the Clothing1M dataset. The results indicate that our approach significantly outperforms other state-of-the-art methods.\n```",
    "1136329": "> inaccurate labels that are termed as noisy labels.\n\nFinally, i understand what **noisy labels** means, thanks for share.",
    "1136351": "Wait, so how are we supposed to use this in our code? Take JointOptimization as a loss?",
    "1136359": "I only tried [this](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324) one so far. However, as per your concern \n\n```\njot = JointOptimization()\nmodel.compile(loss=jot, ..)\n```",
    "1136363": "Thanks for sharing! How were your results so far?",
    "1136372": "As I mentioned, I didn't try this loss function yet. But using Symmetric Cross-Entropy Loss, implementation [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/208324), did help to boost my local cv, from 0.90 to 0.901",
    "1136385": "Sorry, I misunderstood. Thanks a lot for clarifying!",
    "1136924": "Beware that Joint Optimization is End to End Framework and not just a loss function.   Because there is an epoch wise update phase where the (wrong) labels need to be \"corrected\" in semi-supervised manner\nIt's even said in the abstract :\n\n>  To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can **correct labels during training** by **alternating update of network parameters and labels**.\n\n\nTo implement this in TF, you will need custom training loop or may be at least highly customized callback.",
    "1137024": "Thanks for sharing! I was wondering why I couldn't see any improvement in the model after using just the loss function.",
    "1139139": "Thanks for sharing!\n`\nle = K.categorical_crossentropy(y_pred, y_pred)\n`\nI'm confused by this line. What does it do?\nAnd also does \"self.prob\" need to be update after each iteration as well?"
  },
  "source": "meta"
}