{
  "id": 556955,
  "title": "Is F-Beta a good metric?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/556955",
  "author_name": "Success Moses",
  "post_date": "2025-01-15T23:29:31.450000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This is the metric I use to train <code>res_unet</code> model. This F-Beta score, beta=4. I am wondering if this is representative of the competition metrics. </p>\n<pre><code> tensorflow  tf\n\n (tf.keras.metrics.Metric):\n     ():\n        ().__init__(name=name, **kwargs)\n        .beta = beta  \n        \n        .true_labels = []\n        .predicted_labels = []\n\n     ():\n        \n        y_true = tf.reshape(y_true, [-])  \n\n        \n        y_pred = tf.argmax(y_pred, axis=-)  \n        y_pred = tf.reshape(y_pred, [-])  \n\n        .true_labels.append(y_true)\n        .predicted_labels.append(y_pred)\n\n     ():\n        \n        true_labels = tf.concat(.true_labels, axis=)  \n        predicted_labels = tf.concat(.predicted_labels, axis=)  \n\n        \n        num_classes = tf.reduce_max(true_labels) +   \n\n        f1_scores = []\n\n         i  (num_classes):\n            \n            tp = tf.reduce_sum(tf.cast((true_labels == i) &amp; (predicted_labels == i), tf.float32))\n            fp = tf.reduce_sum(tf.cast((true_labels != i) &amp; (predicted_labels == i), tf.float32))\n            fn = tf.reduce_sum(tf.cast((true_labels == i) &amp; (predicted_labels != i), tf.float32))\n\n            \n            precision = tf.math.divide_no_nan(tp, tp + fp)\n            recall = tf.math.divide_no_nan(tp, tp + fn)\n\n            \n            fbeta = tf.math.divide_no_nan(\n                ( + .beta**) * precision * recall,\n                (.beta** * precision) + recall\n            )\n\n            f1_scores.append(fbeta)\n\n         tf.convert_to_tensor(f1_scores, dtype=tf.float32)\n\n     ():\n        \n        .true_labels = []\n        .predicted_labels = []\n</code></pre>",
  "messages": [
    {
      "id": 3097974,
      "postDate": "2025-01-15T23:29:31.450Z",
      "content": "<p>This is the metric I use to train <code>res_unet</code> model. This F-Beta score, beta=4. I am wondering if this is representative of the competition metrics. </p>\n<pre><code> tensorflow  tf\n\n (tf.keras.metrics.Metric):\n     ():\n        ().__init__(name=name, **kwargs)\n        .beta = beta  \n        \n        .true_labels = []\n        .predicted_labels = []\n\n     ():\n        \n        y_true = tf.reshape(y_true, [-])  \n\n        \n        y_pred = tf.argmax(y_pred, axis=-)  \n        y_pred = tf.reshape(y_pred, [-])  \n\n        .true_labels.append(y_true)\n        .predicted_labels.append(y_pred)\n\n     ():\n        \n        true_labels = tf.concat(.true_labels, axis=)  \n        predicted_labels = tf.concat(.predicted_labels, axis=)  \n\n        \n        num_classes = tf.reduce_max(true_labels) +   \n\n        f1_scores = []\n\n         i  (num_classes):\n            \n            tp = tf.reduce_sum(tf.cast((true_labels == i) &amp; (predicted_labels == i), tf.float32))\n            fp = tf.reduce_sum(tf.cast((true_labels != i) &amp; (predicted_labels == i), tf.float32))\n            fn = tf.reduce_sum(tf.cast((true_labels == i) &amp; (predicted_labels != i), tf.float32))\n\n            \n            precision = tf.math.divide_no_nan(tp, tp + fp)\n            recall = tf.math.divide_no_nan(tp, tp + fn)\n\n            \n            fbeta = tf.math.divide_no_nan(\n                ( + .beta**) * precision * recall,\n                (.beta** * precision) + recall\n            )\n\n            f1_scores.append(fbeta)\n\n         tf.convert_to_tensor(f1_scores, dtype=tf.float32)\n\n     ():\n        \n        .true_labels = []\n        .predicted_labels = []\n</code></pre>",
      "rawMarkdown": "This is the metric I use to train `res_unet` model. This F-Beta score, beta=4. I am wondering if this is representative of the competition metrics. \n\n```\nimport tensorflow as tf\n\nclass F1ScorePerClass(tf.keras.metrics.Metric):\n    def __init__(self, beta=4, name=\"f1_score_per_class\", **kwargs):\n        super().__init__(name=name, **kwargs)\n        self.beta = beta  # Set beta to 4 or any other value\n        # Accumulate true labels and predictions across the epoch\n        self.true_labels = []\n        self.predicted_labels = []\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        # Flatten the true labels (y_true is of shape (batch_size, 72, 72, 72))\n        y_true = tf.reshape(y_true, [-1])  # Shape: (batch_size * 72 * 72 * 72,)\n        \n        # Convert predictions to class indices and flatten them (y_pred is of shape (batch_size, 72, 72, 72, n_classes))\n        y_pred = tf.argmax(y_pred, axis=-1)  # Shape: (batch_size, 72, 72, 72)\n        y_pred = tf.reshape(y_pred, [-1])  # Shape: (batch_size * 72 * 72 * 72,)\n        \n        self.true_labels.append(y_true)\n        self.predicted_labels.append(y_pred)\n\n    def result(self):\n        # Concatenate all batches into one array\n        true_labels = tf.concat(self.true_labels, axis=0)  # Shape: (total_samples,)\n        predicted_labels = tf.concat(self.predicted_labels, axis=0)  # Shape: (total_samples,)\n\n        # Calculate the number of classes\n        num_classes = tf.reduce_max(true_labels) + 1  # Assuming class labels are 0-based\n        \n        f1_scores = []\n\n        for i in range(num_classes):\n            # Calculate TP, FP, FN for each class\n            tp = tf.reduce_sum(tf.cast((true_labels == i) & (predicted_labels == i), tf.float32))\n            fp = tf.reduce_sum(tf.cast((true_labels != i) & (predicted_labels == i), tf.float32))\n            fn = tf.reduce_sum(tf.cast((true_labels == i) & (predicted_labels != i), tf.float32))\n\n            # Calculate precision, recall, and F-beta score with beta=4\n            precision = tf.math.divide_no_nan(tp, tp + fp)\n            recall = tf.math.divide_no_nan(tp, tp + fn)\n\n            # Calculate F-beta score\n            fbeta = tf.math.divide_no_nan(\n                (1 + self.beta**2) * precision * recall,\n                (self.beta**2 * precision) + recall\n            )\n\n            f1_scores.append(fbeta)\n\n        return tf.convert_to_tensor(f1_scores, dtype=tf.float32)\n\n    def reset_states(self):\n        # Clear the lists at the end of each epoch\n        self.true_labels = []\n        self.predicted_labels = []\n\n```",
      "votes": 1
    },
    {
      "id": 3098442,
      "postDate": "2025-01-16T13:35:54.723Z",
      "content": "<p>I recommend you to incorporate official metric  in your training pipeline, because your proposed solution will measure metric per pixel for segmentation that is not the same that the metric used for the competition</p>",
      "rawMarkdown": "I recommend you to incorporate official metric  in your training pipeline, because your proposed solution will measure metric per pixel for segmentation that is not the same that the metric used for the competition",
      "votes": 2,
      "replies": [
        {
          "id": 3098556,
          "postDate": "2025-01-16T16:04:40.593Z",
          "content": "<p>I am trying to do that, but computing the official metric involves connected component analysis on predicted masks and other post processing. I don't know if it will be a good idea to do this post processing every epoch just to evaluate the model. I don't know any better way</p>",
          "rawMarkdown": "I am trying to do that, but computing the official metric involves connected component analysis on predicted masks and other post processing. I don't know if it will be a good idea to do this post processing every epoch just to evaluate the model. I don't know any better way",
          "replies": [
            {
              "id": 3098587,
              "postDate": "2025-01-16T16:55:57.710Z",
              "content": "<p>\"I think computing the official metric will slow down my training speed, and my CV-LB relationship is terrible. 😮‍💨 For example: CV 0.80-&gt;LB 0.75 and CV 0.75-&gt;LB 0.763😂. I know the reason, but I can't change it. 🙂‍↕️</p>",
              "rawMarkdown": "\"I think computing the official metric will slow down my training speed, and my CV-LB relationship is terrible. 😮‍💨 For example: CV 0.80->LB 0.75 and CV 0.75->LB 0.763😂. I know the reason, but I can't change it. 🙂‍↕️",
              "votes": 2
            },
            {
              "id": 3098819,
              "postDate": "2025-01-17T00:39:56.533Z",
              "content": "<p>I calculate a pixel wise lb score for my model.  It <strong>might</strong> have a slight advantage over the dice metric (which I also use.)  I probably wouldn't go through the full process of calculating the official metric during training though.  The real problem here is we're trying to use 7 samples to do inference on 500.  A better official metric value may very well mean you over-fitted the 7 samples more.</p>",
              "rawMarkdown": "I calculate a pixel wise lb score for my model.  It **might** have a slight advantage over the dice metric (which I also use.)  I probably wouldn't go through the full process of calculating the official metric during training though.  The real problem here is we're trying to use 7 samples to do inference on 500.  A better official metric value may very well mean you over-fitted the 7 samples more."
            },
            {
              "id": 3109511,
              "postDate": "2025-01-29T00:33:23.280Z",
              "content": "<p>hi <a href=\"https://www.kaggle.com/luoziqian\" target=\"_blank\">@luoziqian</a> <br>\nThanks for sharing. How do you validate ? cross validation with 7 experiments(7fold) or just one experiments ? </p>",
              "rawMarkdown": "hi @luoziqian \nThanks for sharing. How do you validate ? cross validation with 7 experiments(7fold) or just one experiments ? "
            },
            {
              "id": 3109526,
              "postDate": "2025-01-29T01:40:25.350Z",
              "content": "<p>f you have tried to use CV as a metric for training, you will find that CV actually has huge fluctuations during training. Moreover, our experiments found that the LB scores of the 7 samples had a large gap. The reason may be that the number of samples in each category in the 7 samples is not as balanced as expected, leading to large fluctuations. However, this conclusion may not be correct due to the fluctuation of CV during training😂. I tend to use TS_6_4 in order to unify the standards of CV when comparing (although my partner tried others). However, it is well known that TS_6_4 has the only 2-column VLP in the dataset, so the recall of VLP is &lt;= 0.8 if TS_6_4 is selected. I have a guess that calculating CV together with the training set will be more accurate, but I have no theoretical basis except a lot of statistical data😂. (P.S. My English is not really well.😂すみません)</p>",
              "rawMarkdown": "f you have tried to use CV as a metric for training, you will find that CV actually has huge fluctuations during training. Moreover, our experiments found that the LB scores of the 7 samples had a large gap. The reason may be that the number of samples in each category in the 7 samples is not as balanced as expected, leading to large fluctuations. However, this conclusion may not be correct due to the fluctuation of CV during training😂. I tend to use TS_6_4 in order to unify the standards of CV when comparing (although my partner tried others). However, it is well known that TS_6_4 has the only 2-column VLP in the dataset, so the recall of VLP is <= 0.8 if TS_6_4 is selected. I have a guess that calculating CV together with the training set will be more accurate, but I have no theoretical basis except a lot of statistical data😂. (P.S. My English is not really well.😂すみません)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3097987,
      "postDate": "2025-01-16T00:17:30.513Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3098442,
      "author_name": "Kostiantyn Maksymov",
      "author_url": "",
      "post_date": "2025-01-16T13:35:54.723000",
      "content": "<p>I recommend you to incorporate official metric  in your training pipeline, because your proposed solution will measure metric per pixel for segmentation that is not the same that the metric used for the competition</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3098556,
          "author_name": "Success Moses",
          "author_url": "",
          "post_date": "2025-01-16T16:04:40.593000",
          "content": "<p>I am trying to do that, but computing the official metric involves connected component analysis on predicted masks and other post processing. I don't know if it will be a good idea to do this post processing every epoch just to evaluate the model. I don't know any better way</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3098587,
              "author_name": "LuoZiqian",
              "author_url": "",
              "post_date": "2025-01-16T16:55:57.710000",
              "content": "<p>\"I think computing the official metric will slow down my training speed, and my CV-LB relationship is terrible. 😮‍💨 For example: CV 0.80-&gt;LB 0.75 and CV 0.75-&gt;LB 0.763😂. I know the reason, but I can't change it. 🙂‍↕️</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3098819,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-17T00:39:56.533000",
              "content": "<p>I calculate a pixel wise lb score for my model.  It <strong>might</strong> have a slight advantage over the dice metric (which I also use.)  I probably wouldn't go through the full process of calculating the official metric during training though.  The real problem here is we're trying to use 7 samples to do inference on 500.  A better official metric value may very well mean you over-fitted the 7 samples more.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3109511,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2025-01-29T00:33:23.280000",
              "content": "<p>hi <a href=\"https://www.kaggle.com/luoziqian\" target=\"_blank\">@luoziqian</a> <br>\nThanks for sharing. How do you validate ? cross validation with 7 experiments(7fold) or just one experiments ? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3109526,
              "author_name": "LuoZiqian",
              "author_url": "",
              "post_date": "2025-01-29T01:40:25.350000",
              "content": "<p>f you have tried to use CV as a metric for training, you will find that CV actually has huge fluctuations during training. Moreover, our experiments found that the LB scores of the 7 samples had a large gap. The reason may be that the number of samples in each category in the 7 samples is not as balanced as expected, leading to large fluctuations. However, this conclusion may not be correct due to the fluctuation of CV during training😂. I tend to use TS_6_4 in order to unify the standards of CV when comparing (although my partner tried others). However, it is well known that TS_6_4 has the only 2-column VLP in the dataset, so the recall of VLP is &lt;= 0.8 if TS_6_4 is selected. I have a guess that calculating CV together with the training set will be more accurate, but I have no theoretical basis except a lot of statistical data😂. (P.S. My English is not really well.😂すみません)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3097987,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-16T00:17:30.513000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3097974": "This is the metric I use to train `res_unet` model. This F-Beta score, beta=4. I am wondering if this is representative of the competition metrics. \n\n```\nimport tensorflow as tf\n\nclass F1ScorePerClass(tf.keras.metrics.Metric):\n    def __init__(self, beta=4, name=\"f1_score_per_class\", **kwargs):\n        super().__init__(name=name, **kwargs)\n        self.beta = beta  # Set beta to 4 or any other value\n        # Accumulate true labels and predictions across the epoch\n        self.true_labels = []\n        self.predicted_labels = []\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        # Flatten the true labels (y_true is of shape (batch_size, 72, 72, 72))\n        y_true = tf.reshape(y_true, [-1])  # Shape: (batch_size * 72 * 72 * 72,)\n        \n        # Convert predictions to class indices and flatten them (y_pred is of shape (batch_size, 72, 72, 72, n_classes))\n        y_pred = tf.argmax(y_pred, axis=-1)  # Shape: (batch_size, 72, 72, 72)\n        y_pred = tf.reshape(y_pred, [-1])  # Shape: (batch_size * 72 * 72 * 72,)\n        \n        self.true_labels.append(y_true)\n        self.predicted_labels.append(y_pred)\n\n    def result(self):\n        # Concatenate all batches into one array\n        true_labels = tf.concat(self.true_labels, axis=0)  # Shape: (total_samples,)\n        predicted_labels = tf.concat(self.predicted_labels, axis=0)  # Shape: (total_samples,)\n\n        # Calculate the number of classes\n        num_classes = tf.reduce_max(true_labels) + 1  # Assuming class labels are 0-based\n        \n        f1_scores = []\n\n        for i in range(num_classes):\n            # Calculate TP, FP, FN for each class\n            tp = tf.reduce_sum(tf.cast((true_labels == i) & (predicted_labels == i), tf.float32))\n            fp = tf.reduce_sum(tf.cast((true_labels != i) & (predicted_labels == i), tf.float32))\n            fn = tf.reduce_sum(tf.cast((true_labels == i) & (predicted_labels != i), tf.float32))\n\n            # Calculate precision, recall, and F-beta score with beta=4\n            precision = tf.math.divide_no_nan(tp, tp + fp)\n            recall = tf.math.divide_no_nan(tp, tp + fn)\n\n            # Calculate F-beta score\n            fbeta = tf.math.divide_no_nan(\n                (1 + self.beta**2) * precision * recall,\n                (self.beta**2 * precision) + recall\n            )\n\n            f1_scores.append(fbeta)\n\n        return tf.convert_to_tensor(f1_scores, dtype=tf.float32)\n\n    def reset_states(self):\n        # Clear the lists at the end of each epoch\n        self.true_labels = []\n        self.predicted_labels = []\n\n```",
    "3098442": "I recommend you to incorporate official metric  in your training pipeline, because your proposed solution will measure metric per pixel for segmentation that is not the same that the metric used for the competition",
    "3097987": ""
  }
}