{
  "id": 215911,
  "title": "BiTemperedLogisticLoss Usage in any model",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215911",
  "author_name": "",
  "post_date": "2021-01-31T19:05:04.210521400Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>To use BiTemperedLogisticLoss class in any model use the below code. </p>\n<pre><code>loss=BiTemperedLogisticLoss(t1=1.0, t2=1.0)\n\nmodel.compile(\n        optimizer='adam',\n        loss = loss ,\n        metrics=['categorical_accuracy']\n    )\n</code></pre>\n<p>where BiTemperedLogisticLoss class is defined below:</p>\n<pre><code>import tensorflow as tf\n\ndef log_t(u, t):\n  \"\"\"Compute log_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.log(u)\n  else:\n    return (u**(1.0 - t) - 1.0) / (1.0 - t)\n\ndef exp_t(u, t):\n  \"\"\"Compute exp_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.exp(u)\n  else:\n    return tf.math.maximum(0.0, 1.0 + (1.0 - t) * u) ** (1.0 / (1.0 - t))\n\ndef compute_normalization_fixed_point(y_pred, t, num_iters=5):\n    \"\"\"Returns the normalization value for each example (t &gt; 1.0).\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (&gt; 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as y_pred with the last dimension being 1.\n    \"\"\"\n    mu = tf.math.reduce_max(y_pred, -1, keepdims=True)\n    normalized_y_pred_step_0 = y_pred - mu\n    normalized_y_pred = normalized_y_pred_step_0\n    i = 0\n    while i &lt; num_iters:\n        i += 1\n        logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t),-1, keepdims=True)\n        normalized_y_pred = normalized_y_pred_step_0 * (logt_partition ** (1.0 - t))\n\n    logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t), -1, keepdims=True)\n    return -log_t(1.0 / logt_partition, t) + mu\n\ndef compute_normalization(y_pred, t, num_iters=5):\n  \"\"\"Returns the normalization value for each example.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (&lt; 1.0 for finite support, &gt; 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as activation with the last dimension being 1.\n  \"\"\"\n  if t &lt; 1.0:\n    return None # not implemented as these values do not occur in the authors experiments...\n  else:\n    return compute_normalization_fixed_point(y_pred, t, num_iters)\n\ndef tempered_softmax(y_pred, t, num_iters=5):\n    \"\"\"Tempered softmax function.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature tensor &gt; 0.0.\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A probabilities tensor.\n    \"\"\"\n    if t == 1.0:\n        normalization_constants = tf.math.log(tf.math.reduce_sum(tf.math.exp(y_pred), -1, keepdims=True))\n    else:\n        normalization_constants = compute_normalization(y_pred, t, num_iters)\n\n    return exp_t(y_pred - normalization_constants, t)\n\ndef bi_tempered_logistic_loss(y_pred, y_true, t1, t2, num_iters=5, label_smoothing=0.0):\n    \"\"\"Bi-Tempered Logistic Loss with custom gradient.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    y_true: A tensor with shape and dtype as y_pred.\n    t1: Temperature 1 (&lt; 1.0 for boundedness).\n    t2: Temperature 2 (&gt; 1.0 for tail heaviness, &lt; 1.0 for finite support).\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A loss tensor.\n    \"\"\"\n    y_pred = tf.cast(y_pred, tf.float32)\n    y_true = tf.cast(y_true, tf.float32)\n\n    if label_smoothing &gt; 0.0:\n        num_classes = tf.cast(tf.shape(y_true)[-1], tf.float32)\n        y_true = (1 - num_classes /(num_classes - 1) * label_smoothing) * y_true + label_smoothing / (num_classes - 1)\n\n    probabilities = tempered_softmax(y_pred, t2, num_iters)\n\n    temp1 = (log_t(y_true + 1e-10, t1) - log_t(probabilities, t1)) * y_true\n    temp2 = (1 / (2 - t1)) * (tf.math.pow(y_true, 2 - t1) - tf.math.pow(probabilities, 2 - t1))\n    loss_values = temp1 - temp2\n\n    return tf.math.reduce_sum(loss_values, -1)\n\nclass BiTemperedLogisticLoss(tf.keras.losses.Loss):\n    def __init__(self, t1, t2, n_iter=5, label_smoothing=0.0):\n        super(BiTemperedLogisticLoss, self).__init__()\n        self.t1 = t1\n        self.t2 = t2\n        self.n_iter = n_iter\n        self.label_smoothing = label_smoothing\n\n    def call(self, y_true, y_pred):\n        return bi_tempered_logistic_loss(y_pred, y_true, self.t1, self.t2, self.n_iter, self.label_smoothing)\n</code></pre>\n<p>The BiTemperedLogisticLoss class is taken from this <a href=\"https://github.com/Diulhio/bitemperedloss-tf/blob/68b3f7e9ee0d66e836c0ec487720a442edc483bc/tf_bi_tempered_loss.py#L52\" target=\"_blank\">url</a></p>",
  "messages": [
    {
      "id": "1179731",
      "postDate": "01/31/2021 19:05:04",
      "content": "<p>To use BiTemperedLogisticLoss class in any model use the below code. </p>\n<pre><code>loss=BiTemperedLogisticLoss(t1=1.0, t2=1.0)\n\nmodel.compile(\n        optimizer='adam',\n        loss = loss ,\n        metrics=['categorical_accuracy']\n    )\n</code></pre>\n<p>where BiTemperedLogisticLoss class is defined below:</p>\n<pre><code>import tensorflow as tf\n\ndef log_t(u, t):\n  \"\"\"Compute log_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.log(u)\n  else:\n    return (u**(1.0 - t) - 1.0) / (1.0 - t)\n\ndef exp_t(u, t):\n  \"\"\"Compute exp_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.exp(u)\n  else:\n    return tf.math.maximum(0.0, 1.0 + (1.0 - t) * u) ** (1.0 / (1.0 - t))\n\ndef compute_normalization_fixed_point(y_pred, t, num_iters=5):\n    \"\"\"Returns the normalization value for each example (t &gt; 1.0).\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (&gt; 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as y_pred with the last dimension being 1.\n    \"\"\"\n    mu = tf.math.reduce_max(y_pred, -1, keepdims=True)\n    normalized_y_pred_step_0 = y_pred - mu\n    normalized_y_pred = normalized_y_pred_step_0\n    i = 0\n    while i &lt; num_iters:\n        i += 1\n        logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t),-1, keepdims=True)\n        normalized_y_pred = normalized_y_pred_step_0 * (logt_partition ** (1.0 - t))\n\n    logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t), -1, keepdims=True)\n    return -log_t(1.0 / logt_partition, t) + mu\n\ndef compute_normalization(y_pred, t, num_iters=5):\n  \"\"\"Returns the normalization value for each example.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (&lt; 1.0 for finite support, &gt; 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as activation with the last dimension being 1.\n  \"\"\"\n  if t &lt; 1.0:\n    return None # not implemented as these values do not occur in the authors experiments...\n  else:\n    return compute_normalization_fixed_point(y_pred, t, num_iters)\n\ndef tempered_softmax(y_pred, t, num_iters=5):\n    \"\"\"Tempered softmax function.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature tensor &gt; 0.0.\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A probabilities tensor.\n    \"\"\"\n    if t == 1.0:\n        normalization_constants = tf.math.log(tf.math.reduce_sum(tf.math.exp(y_pred), -1, keepdims=True))\n    else:\n        normalization_constants = compute_normalization(y_pred, t, num_iters)\n\n    return exp_t(y_pred - normalization_constants, t)\n\ndef bi_tempered_logistic_loss(y_pred, y_true, t1, t2, num_iters=5, label_smoothing=0.0):\n    \"\"\"Bi-Tempered Logistic Loss with custom gradient.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    y_true: A tensor with shape and dtype as y_pred.\n    t1: Temperature 1 (&lt; 1.0 for boundedness).\n    t2: Temperature 2 (&gt; 1.0 for tail heaviness, &lt; 1.0 for finite support).\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A loss tensor.\n    \"\"\"\n    y_pred = tf.cast(y_pred, tf.float32)\n    y_true = tf.cast(y_true, tf.float32)\n\n    if label_smoothing &gt; 0.0:\n        num_classes = tf.cast(tf.shape(y_true)[-1], tf.float32)\n        y_true = (1 - num_classes /(num_classes - 1) * label_smoothing) * y_true + label_smoothing / (num_classes - 1)\n\n    probabilities = tempered_softmax(y_pred, t2, num_iters)\n\n    temp1 = (log_t(y_true + 1e-10, t1) - log_t(probabilities, t1)) * y_true\n    temp2 = (1 / (2 - t1)) * (tf.math.pow(y_true, 2 - t1) - tf.math.pow(probabilities, 2 - t1))\n    loss_values = temp1 - temp2\n\n    return tf.math.reduce_sum(loss_values, -1)\n\nclass BiTemperedLogisticLoss(tf.keras.losses.Loss):\n    def __init__(self, t1, t2, n_iter=5, label_smoothing=0.0):\n        super(BiTemperedLogisticLoss, self).__init__()\n        self.t1 = t1\n        self.t2 = t2\n        self.n_iter = n_iter\n        self.label_smoothing = label_smoothing\n\n    def call(self, y_true, y_pred):\n        return bi_tempered_logistic_loss(y_pred, y_true, self.t1, self.t2, self.n_iter, self.label_smoothing)\n</code></pre>\n<p>The BiTemperedLogisticLoss class is taken from this <a href=\"https://github.com/Diulhio/bitemperedloss-tf/blob/68b3f7e9ee0d66e836c0ec487720a442edc483bc/tf_bi_tempered_loss.py#L52\" target=\"_blank\">url</a></p>",
      "rawMarkdown": "To use BiTemperedLogisticLoss class in any model use the below code. \n```\nloss=BiTemperedLogisticLoss(t1=1.0, t2=1.0)\n\nmodel.compile(\n        optimizer='adam',\n        loss = loss ,\n        metrics=['categorical_accuracy']\n    )\n```\nwhere BiTemperedLogisticLoss class is defined below:\n\n```\nimport tensorflow as tf\n\ndef log_t(u, t):\n  \"\"\"Compute log_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.log(u)\n  else:\n    return (u**(1.0 - t) - 1.0) / (1.0 - t)\n\ndef exp_t(u, t):\n  \"\"\"Compute exp_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.exp(u)\n  else:\n    return tf.math.maximum(0.0, 1.0 + (1.0 - t) * u) ** (1.0 / (1.0 - t))\n\ndef compute_normalization_fixed_point(y_pred, t, num_iters=5):\n    \"\"\"Returns the normalization value for each example (t > 1.0).\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (> 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as y_pred with the last dimension being 1.\n    \"\"\"\n    mu = tf.math.reduce_max(y_pred, -1, keepdims=True)\n    normalized_y_pred_step_0 = y_pred - mu\n    normalized_y_pred = normalized_y_pred_step_0\n    i = 0\n    while i < num_iters:\n        i += 1\n        logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t),-1, keepdims=True)\n        normalized_y_pred = normalized_y_pred_step_0 * (logt_partition ** (1.0 - t))\n\n    logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t), -1, keepdims=True)\n    return -log_t(1.0 / logt_partition, t) + mu\n\ndef compute_normalization(y_pred, t, num_iters=5):\n  \"\"\"Returns the normalization value for each example.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (< 1.0 for finite support, > 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as activation with the last dimension being 1.\n  \"\"\"\n  if t < 1.0:\n    return None # not implemented as these values do not occur in the authors experiments...\n  else:\n    return compute_normalization_fixed_point(y_pred, t, num_iters)\n\ndef tempered_softmax(y_pred, t, num_iters=5):\n    \"\"\"Tempered softmax function.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature tensor > 0.0.\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A probabilities tensor.\n    \"\"\"\n    if t == 1.0:\n        normalization_constants = tf.math.log(tf.math.reduce_sum(tf.math.exp(y_pred), -1, keepdims=True))\n    else:\n        normalization_constants = compute_normalization(y_pred, t, num_iters)\n\n    return exp_t(y_pred - normalization_constants, t)\n\ndef bi_tempered_logistic_loss(y_pred, y_true, t1, t2, num_iters=5, label_smoothing=0.0):\n    \"\"\"Bi-Tempered Logistic Loss with custom gradient.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    y_true: A tensor with shape and dtype as y_pred.\n    t1: Temperature 1 (< 1.0 for boundedness).\n    t2: Temperature 2 (> 1.0 for tail heaviness, < 1.0 for finite support).\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A loss tensor.\n    \"\"\"\n    y_pred = tf.cast(y_pred, tf.float32)\n    y_true = tf.cast(y_true, tf.float32)\n\n    if label_smoothing > 0.0:\n        num_classes = tf.cast(tf.shape(y_true)[-1], tf.float32)\n        y_true = (1 - num_classes /(num_classes - 1) * label_smoothing) * y_true + label_smoothing / (num_classes - 1)\n\n    probabilities = tempered_softmax(y_pred, t2, num_iters)\n\n    temp1 = (log_t(y_true + 1e-10, t1) - log_t(probabilities, t1)) * y_true\n    temp2 = (1 / (2 - t1)) * (tf.math.pow(y_true, 2 - t1) - tf.math.pow(probabilities, 2 - t1))\n    loss_values = temp1 - temp2\n\n    return tf.math.reduce_sum(loss_values, -1)\n\nclass BiTemperedLogisticLoss(tf.keras.losses.Loss):\n    def __init__(self, t1, t2, n_iter=5, label_smoothing=0.0):\n        super(BiTemperedLogisticLoss, self).__init__()\n        self.t1 = t1\n        self.t2 = t2\n        self.n_iter = n_iter\n        self.label_smoothing = label_smoothing\n\n    def call(self, y_true, y_pred):\n        return bi_tempered_logistic_loss(y_pred, y_true, self.t1, self.t2, self.n_iter, self.label_smoothing)\n\n```\n\nThe BiTemperedLogisticLoss class is taken from this [url](https://github.com/Diulhio/bitemperedloss-tf/blob/68b3f7e9ee0d66e836c0ec487720a442edc483bc/tf_bi_tempered_loss.py#L52)",
      "votes": null
    },
    {
      "id": "1179732",
      "postDate": "01/31/2021 19:09:00",
      "content": "<p>was it useful as compared to conventional cross entropy loss , in any of your experiments?</p>",
      "rawMarkdown": "was it useful as compared to conventional cross entropy loss , in any of your experiments?",
      "votes": null
    },
    {
      "id": "1179741",
      "postDate": "01/31/2021 19:18:06",
      "content": "<p>It was little helpful but did not make much of a difference for me. Since everyone is talking about it so shared it. It might be helpful for others depending on model and methodology used and will act as reference for me in future competitions.</p>",
      "rawMarkdown": "It was little helpful but did not make much of a difference for me. Since everyone is talking about it so shared it. It might be helpful for others depending on model and methodology used and will act as reference for me in future competitions.",
      "votes": null
    },
    {
      "id": "1179743",
      "postDate": "01/31/2021 19:22:30",
      "content": "<p>Ohhh kkkkk…<br>\nthank You for sharing though.</p>\n<p>Can you please also provide some other tips to improve CV . It would be a great help.<br>\nThank You</p>",
      "rawMarkdown": "Ohhh kkkkk...\nthank You for sharing though.\n\nCan you please also provide some other tips to improve CV . It would be a great help.\nThank You",
      "votes": null
    },
    {
      "id": "1179756",
      "postDate": "01/31/2021 19:35:36",
      "content": "<p>U can see my previous posts: </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490\" target=\"_blank\">Old Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215359\" target=\"_blank\">What works</a></li>\n</ol>",
      "rawMarkdown": "U can see my previous posts: \n1. [Old Dataset](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490)\n2. [What works](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215359)",
      "votes": null
    },
    {
      "id": "1181374",
      "postDate": "02/01/2021 20:44:16",
      "content": "<p>I compared both too just a couple hours ago, bi_tempered_loss was better in CV</p>",
      "rawMarkdown": "I compared both too just a couple hours ago, bi_tempered_loss was better in CV",
      "votes": null
    },
    {
      "id": "1181387",
      "postDate": "02/01/2021 21:00:56",
      "content": "<p>Good to see it worked for you</p>",
      "rawMarkdown": "Good to see it worked for you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1179732,
      "author_name": "prashantarorat",
      "author_url": "",
      "post_date": "01/31/2021 19:09:00",
      "content": "<p>was it useful as compared to conventional cross entropy loss , in any of your experiments?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1179741,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "01/31/2021 19:18:06",
          "content": "<p>It was little helpful but did not make much of a difference for me. Since everyone is talking about it so shared it. It might be helpful for others depending on model and methodology used and will act as reference for me in future competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1179743,
          "author_name": "prashantarorat",
          "author_url": "",
          "post_date": "01/31/2021 19:22:30",
          "content": "<p>Ohhh kkkkk…<br>\nthank You for sharing though.</p>\n<p>Can you please also provide some other tips to improve CV . It would be a great help.<br>\nThank You</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1179756,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "01/31/2021 19:35:36",
          "content": "<p>U can see my previous posts: </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490\" target=\"_blank\">Old Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215359\" target=\"_blank\">What works</a></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1181374,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "02/01/2021 20:44:16",
          "content": "<p>I compared both too just a couple hours ago, bi_tempered_loss was better in CV</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1181387,
          "author_name": "vickygoyal",
          "author_url": "",
          "post_date": "02/01/2021 21:00:56",
          "content": "<p>Good to see it worked for you</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1179731": "To use BiTemperedLogisticLoss class in any model use the below code. \n```\nloss=BiTemperedLogisticLoss(t1=1.0, t2=1.0)\n\nmodel.compile(\n        optimizer='adam',\n        loss = loss ,\n        metrics=['categorical_accuracy']\n    )\n```\nwhere BiTemperedLogisticLoss class is defined below:\n\n```\nimport tensorflow as tf\n\ndef log_t(u, t):\n  \"\"\"Compute log_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.log(u)\n  else:\n    return (u**(1.0 - t) - 1.0) / (1.0 - t)\n\ndef exp_t(u, t):\n  \"\"\"Compute exp_t for `u`.\"\"\"\n  if t == 1.0:\n    return tf.math.exp(u)\n  else:\n    return tf.math.maximum(0.0, 1.0 + (1.0 - t) * u) ** (1.0 / (1.0 - t))\n\ndef compute_normalization_fixed_point(y_pred, t, num_iters=5):\n    \"\"\"Returns the normalization value for each example (t > 1.0).\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (> 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as y_pred with the last dimension being 1.\n    \"\"\"\n    mu = tf.math.reduce_max(y_pred, -1, keepdims=True)\n    normalized_y_pred_step_0 = y_pred - mu\n    normalized_y_pred = normalized_y_pred_step_0\n    i = 0\n    while i < num_iters:\n        i += 1\n        logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t),-1, keepdims=True)\n        normalized_y_pred = normalized_y_pred_step_0 * (logt_partition ** (1.0 - t))\n\n    logt_partition = tf.math.reduce_sum(exp_t(normalized_y_pred, t), -1, keepdims=True)\n    return -log_t(1.0 / logt_partition, t) + mu\n\ndef compute_normalization(y_pred, t, num_iters=5):\n  \"\"\"Returns the normalization value for each example.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature 2 (< 1.0 for finite support, > 1.0 for tail heaviness).\n    num_iters: Number of iterations to run the method.\n    Return: A tensor of same rank as activation with the last dimension being 1.\n  \"\"\"\n  if t < 1.0:\n    return None # not implemented as these values do not occur in the authors experiments...\n  else:\n    return compute_normalization_fixed_point(y_pred, t, num_iters)\n\ndef tempered_softmax(y_pred, t, num_iters=5):\n    \"\"\"Tempered softmax function.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    t: Temperature tensor > 0.0.\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A probabilities tensor.\n    \"\"\"\n    if t == 1.0:\n        normalization_constants = tf.math.log(tf.math.reduce_sum(tf.math.exp(y_pred), -1, keepdims=True))\n    else:\n        normalization_constants = compute_normalization(y_pred, t, num_iters)\n\n    return exp_t(y_pred - normalization_constants, t)\n\ndef bi_tempered_logistic_loss(y_pred, y_true, t1, t2, num_iters=5, label_smoothing=0.0):\n    \"\"\"Bi-Tempered Logistic Loss with custom gradient.\n    Args:\n    y_pred: A multi-dimensional tensor with last dimension `num_classes`.\n    y_true: A tensor with shape and dtype as y_pred.\n    t1: Temperature 1 (< 1.0 for boundedness).\n    t2: Temperature 2 (> 1.0 for tail heaviness, < 1.0 for finite support).\n    num_iters: Number of iterations to run the method.\n    Returns:\n    A loss tensor.\n    \"\"\"\n    y_pred = tf.cast(y_pred, tf.float32)\n    y_true = tf.cast(y_true, tf.float32)\n\n    if label_smoothing > 0.0:\n        num_classes = tf.cast(tf.shape(y_true)[-1], tf.float32)\n        y_true = (1 - num_classes /(num_classes - 1) * label_smoothing) * y_true + label_smoothing / (num_classes - 1)\n\n    probabilities = tempered_softmax(y_pred, t2, num_iters)\n\n    temp1 = (log_t(y_true + 1e-10, t1) - log_t(probabilities, t1)) * y_true\n    temp2 = (1 / (2 - t1)) * (tf.math.pow(y_true, 2 - t1) - tf.math.pow(probabilities, 2 - t1))\n    loss_values = temp1 - temp2\n\n    return tf.math.reduce_sum(loss_values, -1)\n\nclass BiTemperedLogisticLoss(tf.keras.losses.Loss):\n    def __init__(self, t1, t2, n_iter=5, label_smoothing=0.0):\n        super(BiTemperedLogisticLoss, self).__init__()\n        self.t1 = t1\n        self.t2 = t2\n        self.n_iter = n_iter\n        self.label_smoothing = label_smoothing\n\n    def call(self, y_true, y_pred):\n        return bi_tempered_logistic_loss(y_pred, y_true, self.t1, self.t2, self.n_iter, self.label_smoothing)\n\n```\n\nThe BiTemperedLogisticLoss class is taken from this [url](https://github.com/Diulhio/bitemperedloss-tf/blob/68b3f7e9ee0d66e836c0ec487720a442edc483bc/tf_bi_tempered_loss.py#L52)",
    "1179732": "was it useful as compared to conventional cross entropy loss , in any of your experiments?",
    "1179741": "It was little helpful but did not make much of a difference for me. Since everyone is talking about it so shared it. It might be helpful for others depending on model and methodology used and will act as reference for me in future competitions.",
    "1179743": "Ohhh kkkkk...\nthank You for sharing though.\n\nCan you please also provide some other tips to improve CV . It would be a great help.\nThank You",
    "1179756": "U can see my previous posts: \n1. [Old Dataset](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215490)\n2. [What works](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215359)",
    "1181374": "I compared both too just a couple hours ago, bi_tempered_loss was better in CV",
    "1181387": "Good to see it worked for you"
  },
  "source": "meta"
}