{
  "id": 130272,
  "title": "Data balancing with TPU ?",
  "url": "/competitions/flower-classification-with-tpus/discussion/130272",
  "author_name": "",
  "post_date": "2020-02-13T06:25:21.170181200Z",
  "votes": 4,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Has someone any remarks about balancing the data using the given format ?</p>\n\n<p>UPDATE:</p>\n\n<p>Following @Ian's suggestion, and adding some lines of code, I found a solution for this. </p>\n\n<p>Here you are the code:\n<a href=\"https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\">https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu</a></p>\n\n<p>I could generate the weights, (as indicated in the comments bellow), and I now have to apply weights to the CNN Kernel. I shall give feedback about the results.  </p>\n\n<p>We are all here to improve ourselves, so if someone has remarks or suggestion, he/she are welcome ! </p>\n\n<p>UPDATE 2:</p>\n\n<p>I tested the solution :</p>\n\n<p>** Generating weights takes a lot of time. I generated them only once and saved them for further parsing.\n** Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.</p>",
  "messages": [
    {
      "id": "744791",
      "postDate": "02/13/2020 06:25:21",
      "content": "<p>Has someone any remarks about balancing the data using the given format ?</p>\n\n<p>UPDATE:</p>\n\n<p>Following @Ian's suggestion, and adding some lines of code, I found a solution for this. </p>\n\n<p>Here you are the code:\n<a href=\"https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\">https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu</a></p>\n\n<p>I could generate the weights, (as indicated in the comments bellow), and I now have to apply weights to the CNN Kernel. I shall give feedback about the results.  </p>\n\n<p>We are all here to improve ourselves, so if someone has remarks or suggestion, he/she are welcome ! </p>\n\n<p>UPDATE 2:</p>\n\n<p>I tested the solution :</p>\n\n<p>** Generating weights takes a lot of time. I generated them only once and saved them for further parsing.\n** Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.</p>",
      "rawMarkdown": "Has someone any remarks about balancing the data using the given format ?\n\nUPDATE:\n\nFollowing @Ian's suggestion, and adding some lines of code, I found a solution for this. \n\nHere you are the code:\nhttps://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\n\nI could generate the weights, (as indicated in the comments bellow), and I now have to apply weights to the CNN Kernel. I shall give feedback about the results.  \n\nWe are all here to improve ourselves, so if someone has remarks or suggestion, he/she are welcome ! \n\nUPDATE 2:\n\nI tested the solution :\n\n** Generating weights takes a lot of time. I generated them only once and saved them for further parsing.\n** Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.",
      "votes": null
    },
    {
      "id": "744871",
      "postDate": "02/13/2020 08:32:23",
      "content": "<p>For unblanacing data the best solution (in deep learning) always is use class weights.  Class wegihts in tpu models isn't avialable, but you can make your own loss function with the wights of the class.\nThis is an example of wieghted categorical loss function:</p>\n\n<p>from keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"\n    A weighted version of keras.objectives.categorical_crossentropy</p>\n\n<pre><code>Variables:\n    weights: numpy array of shape (C,) where C is the number of classes\n\nUsage:\n    weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n    loss = weighted_categorical_crossentropy(weights)\n    model.compile(loss=loss,optimizer='adam')\n\"\"\"\n\nweights = K.variable(weights)\n\ndef loss(y_true, y_pred):\n    # scale predictions so that the class probas of each sample sum to 1\n    y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n    # clip to prevent NaN's and Inf's\n    y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n    # calc\n    loss = y_true * K.log(y_pred) * weights\n    loss = -K.sum(loss, -1)\n    return loss\n\nreturn loss  \n</code></pre>",
      "rawMarkdown": "For unblanacing data the best solution (in deep learning) always is use class weights.  Class wegihts in tpu models isn't avialable, but you can make your own loss function with the wights of the class.\nThis is an example of wieghted categorical loss function:\n\nfrom keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"\n    A weighted version of keras.objectives.categorical_crossentropy\n    \n    Variables:\n        weights: numpy array of shape (C,) where C is the number of classes\n    \n    Usage:\n        weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n        loss = weighted_categorical_crossentropy(weights)\n        model.compile(loss=loss,optimizer='adam')\n    \"\"\"\n    \n    weights = K.variable(weights)\n        \n    def loss(y_true, y_pred):\n        # scale predictions so that the class probas of each sample sum to 1\n        y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n        # clip to prevent NaN's and Inf's\n        y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n        # calc\n        loss = y_true * K.log(y_pred) * weights\n        loss = -K.sum(loss, -1)\n        return loss\n    \n    return loss",
      "votes": null
    },
    {
      "id": "744886",
      "postDate": "02/13/2020 08:47:11",
      "content": "<p>Thank you for your answer! So if I understand well one have to apply the weights to the loss function one declares at the compilation stage. That is practical rather than re-sampling the training set (I do not even want to think about that, given the huge amount of data -- even on TPU).</p>\n\n<p>I shall test that.</p>",
      "rawMarkdown": "Thank you for your answer! So if I understand well one have to apply the weights to the loss function one declares at the compilation stage. That is practical rather than re-sampling the training set (I do not even want to think about that, given the huge amount of data -- even on TPU).\n\nI shall test that.",
      "votes": null
    },
    {
      "id": "744894",
      "postDate": "02/13/2020 09:26:36",
      "content": "<p>Yes. You can apply that loss function (or other wighted loss function) when you compile the model. In my experience this is more easy than change the distrubution of the data, and it's effective.</p>",
      "rawMarkdown": "Yes. You can apply that loss function (or other wighted loss function) when you compile the model. In my experience this is more easy than change the distrubution of the data, and it's effective.",
      "votes": null
    },
    {
      "id": "744895",
      "postDate": "02/13/2020 09:28:48",
      "content": "<p>It is certainly better to do what you say rather than re-sampling! Very good idea! </p>",
      "rawMarkdown": "It is certainly better to do what you say rather than re-sampling! Very good idea!",
      "votes": null
    },
    {
      "id": "744900",
      "postDate": "02/13/2020 09:33:53",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice",
      "votes": null
    },
    {
      "id": "745000",
      "postDate": "02/13/2020 11:34:24",
      "content": "<p>Yes. Now I must find a way and write this using the format we have here (i.e. TF structures). That is the main issue. </p>",
      "rawMarkdown": "Yes. Now I must find a way and write this using the format we have here (i.e. TF structures). That is the main issue.",
      "votes": null
    },
    {
      "id": "745018",
      "postDate": "02/13/2020 11:51:15",
      "content": "<p>For the data o for the loss function?\nIf is for the data i have the same problem.\nIf is for the loss function i dont think you have to make big changes, just import keras backend from tensorflow.keras</p>",
      "rawMarkdown": "For the data o for the loss function?\nIf is for the data i have the same problem.\nIf is for the loss function i dont think you have to make big changes, just import keras backend from tensorflow.keras",
      "votes": null
    },
    {
      "id": "745152",
      "postDate": "02/13/2020 14:59:51",
      "content": "<p>@Ian, you should format correctly your code, look, it seems that a part of your code appears as plain text. I found the function on github (btw, was it you who wrote it ?).</p>",
      "rawMarkdown": "Ian, you should format correctly your code, look, it seems that a part of your code appears as plain text. I found the function on github (btw, was it you who wrote it ?).",
      "votes": null
    },
    {
      "id": "745269",
      "postDate": "02/13/2020 16:54:36",
      "content": "<p>These are great ideas! If you are defining additional variables, make sure they are defined in the <code>strategy.scope()</code>, otherwise they will not be placed on the TPU.</p>",
      "rawMarkdown": "These are great ideas! If you are defining additional variables, make sure they are defined in the `strategy.scope()`, otherwise they will not be placed on the TPU.",
      "votes": null
    },
    {
      "id": "745680",
      "postDate": "02/14/2020 04:30:10",
      "content": "<p>Hallo, just tested the solution : </p>\n\n<ol>\n<li>Generating weights takes a lot of time. I generated them only once and saved them for further parsing.</li>\n<li>Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.</li>\n</ol>",
      "rawMarkdown": "Hallo, just tested the solution : \n\n1. Generating weights takes a lot of time. I generated them only once and saved them for further parsing.\n2. Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.",
      "votes": null
    },
    {
      "id": "746176",
      "postDate": "02/14/2020 17:53:07",
      "content": "<p>Why are there more steps per epoch ?</p>",
      "rawMarkdown": "Why are there more steps per epoch ?",
      "votes": null
    },
    {
      "id": "746184",
      "postDate": "02/14/2020 18:03:14",
      "content": "<p>I do not know why. I suppose the loss function works as such. I should check how the loss functions work in TF/Keras.</p>\n\n<p>I create a new loss function, as indicated in my notebook. It is stored in a variable called, let's say, <code>home_made_loss</code>.</p>\n\n<p>I apply it thus : </p>\n\n<p><code>\nmodel.compile(\n    optimizer='adam',\n    loss = home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n</code></p>\n\n<p>The I run the model : </p>\n\n<p>```\nhistory = model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=10, callbacks=[lr_callback], validation_data=get_validation_dataset())</p>\n\n<p>```</p>\n\n<p>And here, for each iteration, I have NOT 99 steps as usual, but 747 steps, and 1 step per second!\nI can't afford spending so much time, given that we have limited time for the use of TPU.</p>\n\n<p>By the way I have several questions : </p>\n\n<ol>\n<li><p>How can I do in order to add samples to the training batch ? As the loss function is so slow I would like to try to re-sample the training dataset.</p></li>\n<li><p>Is there a way to deal with the batch set without looping on it? I generally try to avoid loops ... I looped on the batch in order to count the nb of samples per class, but I hope there must be a better way to do that. </p></li>\n</ol>",
      "rawMarkdown": "I do not know why. I suppose the loss function works as such. I should check how the loss functions work in TF/Keras.\n\nI create a new loss function, as indicated in my notebook. It is stored in a variable called, let's say, `home_made_loss`.\n\nI apply it thus : \n\n```\nmodel.compile(\n    optimizer='adam',\n    loss = home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n```\n\nThe I run the model : \n\n\n```\nhistory = model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=10, callbacks=[lr_callback], validation_data=get_validation_dataset())\n\n```\n\nAnd here, for each iteration, I have NOT 99 steps as usual, but 747 steps, and 1 step per second!\nI can't afford spending so much time, given that we have limited time for the use of TPU.\n\nBy the way I have several questions : \n\n1. How can I do in order to add samples to the training batch ? As the loss function is so slow I would like to try to re-sample the training dataset.\n\n2. Is there a way to deal with the batch set without looping on it? I generally try to avoid loops ... I looped on the batch in order to count the nb of samples per class, but I hope there must be a better way to do that.",
      "votes": null
    },
    {
      "id": "746294",
      "postDate": "02/14/2020 20:12:29",
      "content": "<p>For counting all elements in the dataset, yes you need to loop on it.</p>",
      "rawMarkdown": "For counting all elements in the dataset, yes you need to loop on it.",
      "votes": null
    },
    {
      "id": "746300",
      "postDate": "02/14/2020 20:15:46",
      "content": "<p>Not sure how exactly you want to add samples, but you might find these two functions useful:</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map\">tf.data.Dataset.map</a>: Applies the same transformation to each element in the dataset. This is a 1:1 operation.</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#flat_map\">tf.data.Dataset.flat_map</a>: Same as above but expected to return multiple elements. This is what you use for a 1:many transformation.</p>",
      "rawMarkdown": "Not sure how exactly you want to add samples, but you might find these two functions useful:\n\n[tf.data.Dataset.map](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map): Applies the same transformation to each element in the dataset. This is a 1:1 operation.\n\n[tf.data.Dataset.flat_map](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#flat_map): Same as above but expected to return multiple elements. This is what you use for a 1:many transformation.",
      "votes": null
    },
    {
      "id": "746305",
      "postDate": "02/14/2020 20:23:23",
      "content": "<p>Well, I want to duplicate (multiply the number of) the samples in the classes with less samples. For example, if I have 100 samples class 1 and 50 samples in class 2, I want to add once again the 50 samples of class 2 to the dataset in order to have 100 samples in each class.</p>\n\n<p>On the TPU, we have a lot of memory but not much time. So I shall try to save the processing time and use more memory. That is why I shall try to increase the  amount of samples in the training set rather than setting weights in the loss function, which is time consuming (I think it would spend 1 hour on each epoch!!!). </p>",
      "rawMarkdown": "Well, I want to duplicate (multiply the number of) the samples in the classes with less samples. For example, if I have 100 samples class 1 and 50 samples in class 2, I want to add once again the 50 samples of class 2 to the dataset in order to have 100 samples in each class.\n\nOn the TPU, we have a lot of memory but not much time. So I shall try to save the processing time and use more memory. That is why I shall try to increase the  amount of samples in the training set rather than setting weights in the loss function, which is time consuming (I think it would spend 1 hour on each epoch!!!).",
      "votes": null
    },
    {
      "id": "746346",
      "postDate": "02/14/2020 21:51:23",
      "content": "<p>Quick idea here: have you tried putting a @tf.function before your custom loss function?</p>\n\n<p>Good idea also on replicating poorly represented samples. I'd even replicate them with transformations (data augmentation). If you want to do that based on class statistics, another function that can help you is:</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#range\">tf.data.Dataset.range</a></p>",
      "rawMarkdown": "Quick idea here: have you tried putting a @tf.function before your custom loss function?\n\nGood idea also on replicating poorly represented samples. I'd even replicate them with transformations (data augmentation). If you want to do that based on class statistics, another function that can help you is:\n\n[tf.data.Dataset.range](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#range)",
      "votes": null
    },
    {
      "id": "746348",
      "postDate": "02/14/2020 21:56:25",
      "content": "<p>What do you mean, adding a tf loss function before the other one? Can we add two loss functions at the compilation stage? </p>",
      "rawMarkdown": "What do you mean, adding a tf loss function before the other one? Can we add two loss functions at the compilation stage?",
      "votes": null
    },
    {
      "id": "746354",
      "postDate": "02/14/2020 22:09:04",
      "content": "<p>No, I meant annotating your loss function with @tf.function. TPUs only execute your Tensorflow graph, not your Python code. @tf.function transforms your Python code into a graph of TF operations. Docs here: <a href=\"https://www.tensorflow.org/api_docs/python/tf/function\">https://www.tensorflow.org/api_docs/python/tf/function</a></p>\n\n<p>I have already ran into cases where my model executed (thanks to TF 2 eager execution) but was extremely slow on TPU without @tf.function.</p>\n\n<p>Just an idea, I have not tested this.</p>",
      "rawMarkdown": "No, I meant annotating your loss function with @tf.function. TPUs only execute your Tensorflow graph, not your Python code. @tf.function transforms your Python code into a graph of TF operations. Docs here: https://www.tensorflow.org/api_docs/python/tf/function\n\nI have already ran into cases where my model executed (thanks to TF 2 eager execution) but was extremely slow on TPU without @tf.function.\n\nJust an idea, I have not tested this.",
      "votes": null
    },
    {
      "id": "746356",
      "postDate": "02/14/2020 22:13:45",
      "content": "<p>That is very good! I noticed that TPU executes only TF (so I use tf.keras.whatever rather than keras.whatever in TPU). I shall search &amp; test that. Thank you!</p>",
      "rawMarkdown": "That is very good! I noticed that TPU executes only TF (so I use tf.keras.whatever rather than keras.whatever in TPU). I shall search &amp; test that. Thank you!",
      "votes": null
    },
    {
      "id": "752320",
      "postDate": "02/20/2020 21:48:46",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> :  Hallo,I tested the 'home made loss' and then I tried to use one of Keras losses, but it seems that the other one is used in the model, even if I change the loss. For example, if in <code>model.compile</code> I use Keras loss, and then I run the training with <code>model.fit</code>, the behavior is similar to the one of my loss (long and 747 steps) and not of that of the Keras functions. I mean as if the loss I created was recorded somewhere and any kernel I use points on it systematically. Can anyone help here ? </p>",
      "rawMarkdown": "mgornergoogle :  Hallo,I tested the 'home made loss' and then I tried to use one of Keras losses, but it seems that the other one is used in the model, even if I change the loss. For example, if in `model.compile` I use Keras loss, and then I run the training with `model.fit`, the behavior is similar to the one of my loss (long and 747 steps) and not of that of the Keras functions. I mean as if the loss I created was recorded somewhere and any kernel I use points on it systematically. Can anyone help here ?",
      "votes": null
    },
    {
      "id": "752548",
      "postDate": "02/21/2020 06:27:50",
      "content": "<p>TensorFlow published a guide to help with imbalanced data <a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\">here</a>. They show examples of various solutions including setting class weights and upsampling the tensorflow dataset.</p>",
      "rawMarkdown": "TensorFlow published a guide to help with imbalanced data [here][1]. They show examples of various solutions including setting class weights and upsampling the tensorflow dataset.\n\n[1]: https://www.tensorflow.org/tutorials/structured_data/imbalanced_data",
      "votes": null
    },
    {
      "id": "752590",
      "postDate": "02/21/2020 07:33:30",
      "content": "<p>Thank you! The problem here is that I cannot change the loss function I created and I used once. Even if I change it, it seems that the algorithm still points on the loss I created and not on the one I set. </p>",
      "rawMarkdown": "Thank you! The problem here is that I cannot change the loss function I created and I used once. Even if I change it, it seems that the algorithm still points on the loss I created and not on the one I set.",
      "votes": null
    },
    {
      "id": "752604",
      "postDate": "02/21/2020 07:52:31",
      "content": "<p>Modifying loss is one way to deal with imbalance. Another way is to leave the loss as it is, and resample the dataset so you have an equal proportion of each type of flower. And there are other ways described in the link I posted above.</p>",
      "rawMarkdown": "Modifying loss is one way to deal with imbalance. Another way is to leave the loss as it is, and resample the dataset so you have an equal proportion of each type of flower. And there are other ways described in the link I posted above.",
      "votes": null
    },
    {
      "id": "752613",
      "postDate": "02/21/2020 08:05:15",
      "content": "<p>Thank you for your answer. The problem is not that, but the fact that, even if I do set a different loss, it seems that the kernel will always point on the old one. I did at first, using my loss :</p>\n\n<p><code>\nmodel.compile(\n    optimizer='adam',\n    loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n</code></p>\n\n<p>Then, <code>model.fit</code> and run.</p>\n\n<p>Afterwards, I changed : </p>\n\n<p><code>\nmodel.compile(\n    optimizer='adam',\n    loss = 'sparse_categorical_crossentropy',\n    #loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n</code></p>\n\n<p>Then <code>model.fit</code> and run. But even now, it seems that, when running, the kernel still points on <code>home_made_loss</code>, as the behaviour is the one of <code>home_made_loss</code> and not the one of <code>sparse_categorical_crossentropy</code>.</p>\n\n<p>My question is : What shall I do in order to use <code>sparse_categorical_crossentropy</code> again ? </p>",
      "rawMarkdown": "Thank you for your answer. The problem is not that, but the fact that, even if I do set a different loss, it seems that the kernel will always point on the old one. I did at first, using my loss :\n\n```\nmodel.compile(\n    optimizer='adam',\n    loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n```\n\nThen, `model.fit` and run.\n\nAfterwards, I changed : \n\n```\nmodel.compile(\n    optimizer='adam',\n    loss = 'sparse_categorical_crossentropy',\n    #loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n```\n\nThen `model.fit` and run. But even now, it seems that, when running, the kernel still points on `home_made_loss`, as the behaviour is the one of `home_made_loss` and not the one of `sparse_categorical_crossentropy`.\n\nMy question is : What shall I do in order to use `sparse_categorical_crossentropy` again ?",
      "votes": null
    },
    {
      "id": "755279",
      "postDate": "02/24/2020 16:12:41",
      "content": "<p>Thanks for the example!  I'm  going to run through this myself.</p>",
      "rawMarkdown": "Thanks for the example!  I'm  going to run through this myself.",
      "votes": null
    },
    {
      "id": "755422",
      "postDate": "02/24/2020 19:03:59",
      "content": "<p>It is easy to lose track of what is defined and used where in a Notebook. Have you tried restarting the notebook and running it from the beginning again ? If you are looking for a less extreme solution, I usually re-run all model creation cells when I change anything to my model.</p>",
      "rawMarkdown": "It is easy to lose track of what is defined and used where in a Notebook. Have you tried restarting the notebook and running it from the beginning again ? If you are looking for a less extreme solution, I usually re-run all model creation cells when I change anything to my model.",
      "votes": null
    },
    {
      "id": "755428",
      "postDate": "02/24/2020 19:08:58",
      "content": "<p>I tried having a look at your notebook but it has errors in many cell outputs. Ping me when you get this in a more runnable shape and I can have a look again. As for the code using the wrong loss, your best bet is to re-run the notebook from scratch. The ability to execute notebooks cells out of order can sometimes be confusing.</p>",
      "rawMarkdown": "I tried having a look at your notebook but it has errors in many cell outputs. Ping me when you get this in a more runnable shape and I can have a look again. As for the code using the wrong loss, your best bet is to re-run the notebook from scratch. The ability to execute notebooks cells out of order can sometimes be confusing.",
      "votes": null
    },
    {
      "id": "755443",
      "postDate": "02/24/2020 19:25:39",
      "content": "<p>I shall make my notebook runnable, at the beginning I just wanted to present a function in it, and not run it. \nConcerning the wrong loss : I always run the notebook from scratch (run all cells), I do not like running cell per cell, I find it bad practice.\nI also stopped the notebook and started it again. </p>",
      "rawMarkdown": "I shall make my notebook runnable, at the beginning I just wanted to present a function in it, and not run it. \nConcerning the wrong loss : I always run the notebook from scratch (run all cells), I do not like running cell per cell, I find it bad practice.\nI also stopped the notebook and started it again.",
      "votes": null
    },
    {
      "id": "755509",
      "postDate": "02/24/2020 20:43:47",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "756123",
      "postDate": "02/25/2020 12:41:19",
      "content": "<p>Here, my script is runnable, I added all the required settings. Can you have a look please ? I've just done a commit. </p>",
      "rawMarkdown": "Here, my script is runnable, I added all the required settings. Can you have a look please ? I've just done a commit.",
      "votes": null
    },
    {
      "id": "756642",
      "postDate": "02/25/2020 23:08:44",
      "content": "<p>I tried to run it but <code>load_dataset</code> is not defined in your code.\nThe script I tried: <a href=\"https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\">https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu</a></p>\n\n<p>Also, the <code>class_weight</code> parameter in <code>model.fit()</code> can be used to assigned different weights to different classes. On TPU, this has to be a flat list of weights (not a dictionary as written in the Keras docs). No need to define a custom loss.</p>",
      "rawMarkdown": "I tried to run it but `load_dataset` is not defined in your code.\nThe script I tried: https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\n\nAlso, the `class_weight` parameter in `model.fit()` can be used to assigned different weights to different classes. On TPU, this has to be a flat list of weights (not a dictionary as written in the Keras docs). No need to define a custom loss.",
      "votes": null
    },
    {
      "id": "756981",
      "postDate": "02/26/2020 09:49:08",
      "content": "<p>I added the function. I could not test yesterday, I had an error related to a connection problem (not mine, Google perhaps). I see what you mean, it is explained <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131876#754366\">HERE</a>. </p>\n\n<p>The problem is that even now, after having changed the loss, I have the same behaviour or my loss. I mean, as if I did not change it. I tried to create another notebook but I cannot have two notebooks with TPU. </p>",
      "rawMarkdown": "I added the function. I could not test yesterday, I had an error related to a connection problem (not mine, Google perhaps). I see what you mean, it is explained [HERE](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131876#754366). \n\nThe problem is that even now, after having changed the loss, I have the same behaviour or my loss. I mean, as if I did not change it. I tried to create another notebook but I cannot have two notebooks with TPU.",
      "votes": null
    },
    {
      "id": "799465",
      "postDate": "04/06/2020 13:08:35",
      "content": "<p>class_weight parameter in model.fit() can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more simple for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )</p>\n\n<p>on keras you have to pass a dictionary to class_weight parameter in model.fit()\nexample class_weights ={0:50, 1:0.5, 2:20}\n               model.fit(....     ....  class_weight=class_weights)</p>\n\n<p>but on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)\nso you have just to convert your dictionary to a flat list </p>\n\n<p>*<strong>*class_weights =[item for k in class_weights for item in (k, class_weights[k])]**</strong></p>\n\n<p><strong>No need to define a custom loss.</strong></p>",
      "rawMarkdown": "class_weight parameter in model.fit() can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more simple for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )\n\non keras you have to pass a dictionary to class_weight parameter in model.fit()\nexample class_weights ={0:50, 1:0.5, 2:20}\n               model.fit(....     ....  class_weight=class_weights)\n\nbut on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)\nso you have just to convert your dictionary to a flat list \n\n****class_weights =[item for k in class_weights for item in (k, class_weights[k])]****\n\n**No need to define a custom loss.**",
      "votes": null
    },
    {
      "id": "799472",
      "postDate": "04/06/2020 13:14:41",
      "content": "<p>Indeed. Someone opened a discussion on this matter. Too late for me, I run my program with the new loss and even when removing it, it seems that I am still blocked on the loss I created. I mean, whatever I do, I set the original settings, including what you suggested (it was already mentioned in a discussion as I wrote before), the training has about 750 steps and takes a lot of time to parse one step, it takes 4 hours to parse 100 steps. I could not solve this and I am not allowed to open a new notebook (with TPU &amp;co) on this competition.</p>",
      "rawMarkdown": "Indeed. Someone opened a discussion on this matter. Too late for me, I run my program with the new loss and even when removing it, it seems that I am still blocked on the loss I created. I mean, whatever I do, I set the original settings, including what you suggested (it was already mentioned in a discussion as I wrote before), the training has about 750 steps and takes a lot of time to parse one step, it takes 4 hours to parse 100 steps. I could not solve this and I am not allowed to open a new notebook (with TPU &amp;co) on this competition.",
      "votes": null
    },
    {
      "id": "799744",
      "postDate": "04/06/2020 17:56:34",
      "content": "<p>pouvez vous partager votre kernel afin que je puisse jeter un coup d'oeil , peut etre je peux aider et par la meme aucasion aprendre quelque chose de nouveau </p>",
      "rawMarkdown": "pouvez vous partager votre kernel afin que je puisse jeter un coup d'oeil , peut etre je peux aider et par la meme aucasion aprendre quelque chose de nouveau",
      "votes": null
    },
    {
      "id": "824496",
      "postDate": "04/28/2020 12:14:47",
      "content": "<p>And did it work..?</p>",
      "rawMarkdown": "And did it work..?",
      "votes": null
    },
    {
      "id": "825174",
      "postDate": "04/28/2020 20:39:38",
      "content": "<p>No, couldn't debug yet.</p>",
      "rawMarkdown": "No, couldn't debug yet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 744871,
      "author_name": "ianwilkinson",
      "author_url": "",
      "post_date": "02/13/2020 08:32:23",
      "content": "<p>For unblanacing data the best solution (in deep learning) always is use class weights.  Class wegihts in tpu models isn't avialable, but you can make your own loss function with the wights of the class.\nThis is an example of wieghted categorical loss function:</p>\n\n<p>from keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"\n    A weighted version of keras.objectives.categorical_crossentropy</p>\n\n<pre><code>Variables:\n    weights: numpy array of shape (C,) where C is the number of classes\n\nUsage:\n    weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n    loss = weighted_categorical_crossentropy(weights)\n    model.compile(loss=loss,optimizer='adam')\n\"\"\"\n\nweights = K.variable(weights)\n\ndef loss(y_true, y_pred):\n    # scale predictions so that the class probas of each sample sum to 1\n    y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n    # clip to prevent NaN's and Inf's\n    y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n    # calc\n    loss = y_true * K.log(y_pred) * weights\n    loss = -K.sum(loss, -1)\n    return loss\n\nreturn loss  \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 744886,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/13/2020 08:47:11",
          "content": "<p>Thank you for your answer! So if I understand well one have to apply the weights to the loss function one declares at the compilation stage. That is practical rather than re-sampling the training set (I do not even want to think about that, given the huge amount of data -- even on TPU).</p>\n\n<p>I shall test that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 744894,
          "author_name": "ianwilkinson",
          "author_url": "",
          "post_date": "02/13/2020 09:26:36",
          "content": "<p>Yes. You can apply that loss function (or other wighted loss function) when you compile the model. In my experience this is more easy than change the distrubution of the data, and it's effective.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745152,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/13/2020 14:59:51",
          "content": "<p>@Ian, you should format correctly your code, look, it seems that a part of your code appears as plain text. I found the function on github (btw, was it you who wrote it ?).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745269,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/13/2020 16:54:36",
          "content": "<p>These are great ideas! If you are defining additional variables, make sure they are defined in the <code>strategy.scope()</code>, otherwise they will not be placed on the TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 744895,
      "author_name": "catadanna",
      "author_url": "",
      "post_date": "02/13/2020 09:28:48",
      "content": "<p>It is certainly better to do what you say rather than re-sampling! Very good idea! </p>",
      "votes": null,
      "replies": [
        {
          "id": 744900,
          "author_name": "ianwilkinson",
          "author_url": "",
          "post_date": "02/13/2020 09:33:53",
          "content": "<p>Nice</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745000,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/13/2020 11:34:24",
          "content": "<p>Yes. Now I must find a way and write this using the format we have here (i.e. TF structures). That is the main issue. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745018,
          "author_name": "ianwilkinson",
          "author_url": "",
          "post_date": "02/13/2020 11:51:15",
          "content": "<p>For the data o for the loss function?\nIf is for the data i have the same problem.\nIf is for the loss function i dont think you have to make big changes, just import keras backend from tensorflow.keras</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745680,
      "author_name": "catadanna",
      "author_url": "",
      "post_date": "02/14/2020 04:30:10",
      "content": "<p>Hallo, just tested the solution : </p>\n\n<ol>\n<li>Generating weights takes a lot of time. I generated them only once and saved them for further parsing.</li>\n<li>Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 746176,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 17:53:07",
          "content": "<p>Why are there more steps per epoch ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746184,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/14/2020 18:03:14",
          "content": "<p>I do not know why. I suppose the loss function works as such. I should check how the loss functions work in TF/Keras.</p>\n\n<p>I create a new loss function, as indicated in my notebook. It is stored in a variable called, let's say, <code>home_made_loss</code>.</p>\n\n<p>I apply it thus : </p>\n\n<p><code>\nmodel.compile(\n    optimizer='adam',\n    loss = home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n</code></p>\n\n<p>The I run the model : </p>\n\n<p>```\nhistory = model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=10, callbacks=[lr_callback], validation_data=get_validation_dataset())</p>\n\n<p>```</p>\n\n<p>And here, for each iteration, I have NOT 99 steps as usual, but 747 steps, and 1 step per second!\nI can't afford spending so much time, given that we have limited time for the use of TPU.</p>\n\n<p>By the way I have several questions : </p>\n\n<ol>\n<li><p>How can I do in order to add samples to the training batch ? As the loss function is so slow I would like to try to re-sample the training dataset.</p></li>\n<li><p>Is there a way to deal with the batch set without looping on it? I generally try to avoid loops ... I looped on the batch in order to count the nb of samples per class, but I hope there must be a better way to do that. </p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746294,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 20:12:29",
          "content": "<p>For counting all elements in the dataset, yes you need to loop on it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746300,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 20:15:46",
          "content": "<p>Not sure how exactly you want to add samples, but you might find these two functions useful:</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map\">tf.data.Dataset.map</a>: Applies the same transformation to each element in the dataset. This is a 1:1 operation.</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#flat_map\">tf.data.Dataset.flat_map</a>: Same as above but expected to return multiple elements. This is what you use for a 1:many transformation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746305,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/14/2020 20:23:23",
          "content": "<p>Well, I want to duplicate (multiply the number of) the samples in the classes with less samples. For example, if I have 100 samples class 1 and 50 samples in class 2, I want to add once again the 50 samples of class 2 to the dataset in order to have 100 samples in each class.</p>\n\n<p>On the TPU, we have a lot of memory but not much time. So I shall try to save the processing time and use more memory. That is why I shall try to increase the  amount of samples in the training set rather than setting weights in the loss function, which is time consuming (I think it would spend 1 hour on each epoch!!!). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746346,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 21:51:23",
          "content": "<p>Quick idea here: have you tried putting a @tf.function before your custom loss function?</p>\n\n<p>Good idea also on replicating poorly represented samples. I'd even replicate them with transformations (data augmentation). If you want to do that based on class statistics, another function that can help you is:</p>\n\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#range\">tf.data.Dataset.range</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746348,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/14/2020 21:56:25",
          "content": "<p>What do you mean, adding a tf loss function before the other one? Can we add two loss functions at the compilation stage? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746354,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 22:09:04",
          "content": "<p>No, I meant annotating your loss function with @tf.function. TPUs only execute your Tensorflow graph, not your Python code. @tf.function transforms your Python code into a graph of TF operations. Docs here: <a href=\"https://www.tensorflow.org/api_docs/python/tf/function\">https://www.tensorflow.org/api_docs/python/tf/function</a></p>\n\n<p>I have already ran into cases where my model executed (thanks to TF 2 eager execution) but was extremely slow on TPU without @tf.function.</p>\n\n<p>Just an idea, I have not tested this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746356,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/14/2020 22:13:45",
          "content": "<p>That is very good! I noticed that TPU executes only TF (so I use tf.keras.whatever rather than keras.whatever in TPU). I shall search &amp; test that. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824496,
          "author_name": "ronyroy",
          "author_url": "",
          "post_date": "04/28/2020 12:14:47",
          "content": "<p>And did it work..?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 825174,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "04/28/2020 20:39:38",
          "content": "<p>No, couldn't debug yet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 752320,
      "author_name": "catadanna",
      "author_url": "",
      "post_date": "02/20/2020 21:48:46",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> :  Hallo,I tested the 'home made loss' and then I tried to use one of Keras losses, but it seems that the other one is used in the model, even if I change the loss. For example, if in <code>model.compile</code> I use Keras loss, and then I run the training with <code>model.fit</code>, the behavior is similar to the one of my loss (long and 747 steps) and not of that of the Keras functions. I mean as if the loss I created was recorded somewhere and any kernel I use points on it systematically. Can anyone help here ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 755428,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/24/2020 19:08:58",
          "content": "<p>I tried having a look at your notebook but it has errors in many cell outputs. Ping me when you get this in a more runnable shape and I can have a look again. As for the code using the wrong loss, your best bet is to re-run the notebook from scratch. The ability to execute notebooks cells out of order can sometimes be confusing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755443,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/24/2020 19:25:39",
          "content": "<p>I shall make my notebook runnable, at the beginning I just wanted to present a function in it, and not run it. \nConcerning the wrong loss : I always run the notebook from scratch (run all cells), I do not like running cell per cell, I find it bad practice.\nI also stopped the notebook and started it again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756123,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/25/2020 12:41:19",
          "content": "<p>Here, my script is runnable, I added all the required settings. Can you have a look please ? I've just done a commit. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756642,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/25/2020 23:08:44",
          "content": "<p>I tried to run it but <code>load_dataset</code> is not defined in your code.\nThe script I tried: <a href=\"https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\">https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu</a></p>\n\n<p>Also, the <code>class_weight</code> parameter in <code>model.fit()</code> can be used to assigned different weights to different classes. On TPU, this has to be a flat list of weights (not a dictionary as written in the Keras docs). No need to define a custom loss.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756981,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/26/2020 09:49:08",
          "content": "<p>I added the function. I could not test yesterday, I had an error related to a connection problem (not mine, Google perhaps). I see what you mean, it is explained <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131876#754366\">HERE</a>. </p>\n\n<p>The problem is that even now, after having changed the loss, I have the same behaviour or my loss. I mean, as if I did not change it. I tried to create another notebook but I cannot have two notebooks with TPU. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 752548,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/21/2020 06:27:50",
      "content": "<p>TensorFlow published a guide to help with imbalanced data <a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\">here</a>. They show examples of various solutions including setting class weights and upsampling the tensorflow dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 752590,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/21/2020 07:33:30",
          "content": "<p>Thank you! The problem here is that I cannot change the loss function I created and I used once. Even if I change it, it seems that the algorithm still points on the loss I created and not on the one I set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752604,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/21/2020 07:52:31",
          "content": "<p>Modifying loss is one way to deal with imbalance. Another way is to leave the loss as it is, and resample the dataset so you have an equal proportion of each type of flower. And there are other ways described in the link I posted above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752613,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "02/21/2020 08:05:15",
          "content": "<p>Thank you for your answer. The problem is not that, but the fact that, even if I do set a different loss, it seems that the kernel will always point on the old one. I did at first, using my loss :</p>\n\n<p><code>\nmodel.compile(\n    optimizer='adam',\n    loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n</code></p>\n\n<p>Then, <code>model.fit</code> and run.</p>\n\n<p>Afterwards, I changed : </p>\n\n<p><code>\nmodel.compile(\n    optimizer='adam',\n    loss = 'sparse_categorical_crossentropy',\n    #loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n</code></p>\n\n<p>Then <code>model.fit</code> and run. But even now, it seems that, when running, the kernel still points on <code>home_made_loss</code>, as the behaviour is the one of <code>home_made_loss</code> and not the one of <code>sparse_categorical_crossentropy</code>.</p>\n\n<p>My question is : What shall I do in order to use <code>sparse_categorical_crossentropy</code> again ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755279,
          "author_name": "larsen0966",
          "author_url": "",
          "post_date": "02/24/2020 16:12:41",
          "content": "<p>Thanks for the example!  I'm  going to run through this myself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755422,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/24/2020 19:03:59",
          "content": "<p>It is easy to lose track of what is defined and used where in a Notebook. Have you tried restarting the notebook and running it from the beginning again ? If you are looking for a less extreme solution, I usually re-run all model creation cells when I change anything to my model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755509,
          "author_name": "larsen0966",
          "author_url": "",
          "post_date": "02/24/2020 20:43:47",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799465,
      "author_name": "hatemamine",
      "author_url": "",
      "post_date": "04/06/2020 13:08:35",
      "content": "<p>class_weight parameter in model.fit() can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more simple for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )</p>\n\n<p>on keras you have to pass a dictionary to class_weight parameter in model.fit()\nexample class_weights ={0:50, 1:0.5, 2:20}\n               model.fit(....     ....  class_weight=class_weights)</p>\n\n<p>but on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)\nso you have just to convert your dictionary to a flat list </p>\n\n<p>*<strong>*class_weights =[item for k in class_weights for item in (k, class_weights[k])]**</strong></p>\n\n<p><strong>No need to define a custom loss.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 799472,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "04/06/2020 13:14:41",
          "content": "<p>Indeed. Someone opened a discussion on this matter. Too late for me, I run my program with the new loss and even when removing it, it seems that I am still blocked on the loss I created. I mean, whatever I do, I set the original settings, including what you suggested (it was already mentioned in a discussion as I wrote before), the training has about 750 steps and takes a lot of time to parse one step, it takes 4 hours to parse 100 steps. I could not solve this and I am not allowed to open a new notebook (with TPU &amp;co) on this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799744,
          "author_name": "hatemamine",
          "author_url": "",
          "post_date": "04/06/2020 17:56:34",
          "content": "<p>pouvez vous partager votre kernel afin que je puisse jeter un coup d'oeil , peut etre je peux aider et par la meme aucasion aprendre quelque chose de nouveau </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "744791": "Has someone any remarks about balancing the data using the given format ?\n\nUPDATE:\n\nFollowing @Ian's suggestion, and adding some lines of code, I found a solution for this. \n\nHere you are the code:\nhttps://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\n\nI could generate the weights, (as indicated in the comments bellow), and I now have to apply weights to the CNN Kernel. I shall give feedback about the results.  \n\nWe are all here to improve ourselves, so if someone has remarks or suggestion, he/she are welcome ! \n\nUPDATE 2:\n\nI tested the solution :\n\n** Generating weights takes a lot of time. I generated them only once and saved them for further parsing.\n** Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.",
    "744871": "For unblanacing data the best solution (in deep learning) always is use class weights.  Class wegihts in tpu models isn't avialable, but you can make your own loss function with the wights of the class.\nThis is an example of wieghted categorical loss function:\n\nfrom keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"\n    A weighted version of keras.objectives.categorical_crossentropy\n    \n    Variables:\n        weights: numpy array of shape (C,) where C is the number of classes\n    \n    Usage:\n        weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n        loss = weighted_categorical_crossentropy(weights)\n        model.compile(loss=loss,optimizer='adam')\n    \"\"\"\n    \n    weights = K.variable(weights)\n        \n    def loss(y_true, y_pred):\n        # scale predictions so that the class probas of each sample sum to 1\n        y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n        # clip to prevent NaN's and Inf's\n        y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n        # calc\n        loss = y_true * K.log(y_pred) * weights\n        loss = -K.sum(loss, -1)\n        return loss\n    \n    return loss",
    "744886": "Thank you for your answer! So if I understand well one have to apply the weights to the loss function one declares at the compilation stage. That is practical rather than re-sampling the training set (I do not even want to think about that, given the huge amount of data -- even on TPU).\n\nI shall test that.",
    "744894": "Yes. You can apply that loss function (or other wighted loss function) when you compile the model. In my experience this is more easy than change the distrubution of the data, and it's effective.",
    "744895": "It is certainly better to do what you say rather than re-sampling! Very good idea!",
    "744900": "Nice",
    "745000": "Yes. Now I must find a way and write this using the format we have here (i.e. TF structures). That is the main issue.",
    "745018": "For the data o for the loss function?\nIf is for the data i have the same problem.\nIf is for the loss function i dont think you have to make big changes, just import keras backend from tensorflow.keras",
    "745152": "Ian, you should format correctly your code, look, it seems that a part of your code appears as plain text. I found the function on github (btw, was it you who wrote it ?).",
    "745269": "These are great ideas! If you are defining additional variables, make sure they are defined in the `strategy.scope()`, otherwise they will not be placed on the TPU.",
    "745680": "Hallo, just tested the solution : \n\n1. Generating weights takes a lot of time. I generated them only once and saved them for further parsing.\n2. Fitting the model with the new loss is also very slow, have 747 steps per epoch rather than 100. I am not sure if it is the best solution for this particular case.",
    "746176": "Why are there more steps per epoch ?",
    "746184": "I do not know why. I suppose the loss function works as such. I should check how the loss functions work in TF/Keras.\n\nI create a new loss function, as indicated in my notebook. It is stored in a variable called, let's say, `home_made_loss`.\n\nI apply it thus : \n\n```\nmodel.compile(\n    optimizer='adam',\n    loss = home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n```\n\nThe I run the model : \n\n\n```\nhistory = model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=10, callbacks=[lr_callback], validation_data=get_validation_dataset())\n\n```\n\nAnd here, for each iteration, I have NOT 99 steps as usual, but 747 steps, and 1 step per second!\nI can't afford spending so much time, given that we have limited time for the use of TPU.\n\nBy the way I have several questions : \n\n1. How can I do in order to add samples to the training batch ? As the loss function is so slow I would like to try to re-sample the training dataset.\n\n2. Is there a way to deal with the batch set without looping on it? I generally try to avoid loops ... I looped on the batch in order to count the nb of samples per class, but I hope there must be a better way to do that.",
    "746294": "For counting all elements in the dataset, yes you need to loop on it.",
    "746300": "Not sure how exactly you want to add samples, but you might find these two functions useful:\n\n[tf.data.Dataset.map](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map): Applies the same transformation to each element in the dataset. This is a 1:1 operation.\n\n[tf.data.Dataset.flat_map](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#flat_map): Same as above but expected to return multiple elements. This is what you use for a 1:many transformation.",
    "746305": "Well, I want to duplicate (multiply the number of) the samples in the classes with less samples. For example, if I have 100 samples class 1 and 50 samples in class 2, I want to add once again the 50 samples of class 2 to the dataset in order to have 100 samples in each class.\n\nOn the TPU, we have a lot of memory but not much time. So I shall try to save the processing time and use more memory. That is why I shall try to increase the  amount of samples in the training set rather than setting weights in the loss function, which is time consuming (I think it would spend 1 hour on each epoch!!!).",
    "746346": "Quick idea here: have you tried putting a @tf.function before your custom loss function?\n\nGood idea also on replicating poorly represented samples. I'd even replicate them with transformations (data augmentation). If you want to do that based on class statistics, another function that can help you is:\n\n[tf.data.Dataset.range](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#range)",
    "746348": "What do you mean, adding a tf loss function before the other one? Can we add two loss functions at the compilation stage?",
    "746354": "No, I meant annotating your loss function with @tf.function. TPUs only execute your Tensorflow graph, not your Python code. @tf.function transforms your Python code into a graph of TF operations. Docs here: https://www.tensorflow.org/api_docs/python/tf/function\n\nI have already ran into cases where my model executed (thanks to TF 2 eager execution) but was extremely slow on TPU without @tf.function.\n\nJust an idea, I have not tested this.",
    "746356": "That is very good! I noticed that TPU executes only TF (so I use tf.keras.whatever rather than keras.whatever in TPU). I shall search &amp; test that. Thank you!",
    "752320": "mgornergoogle :  Hallo,I tested the 'home made loss' and then I tried to use one of Keras losses, but it seems that the other one is used in the model, even if I change the loss. For example, if in `model.compile` I use Keras loss, and then I run the training with `model.fit`, the behavior is similar to the one of my loss (long and 747 steps) and not of that of the Keras functions. I mean as if the loss I created was recorded somewhere and any kernel I use points on it systematically. Can anyone help here ?",
    "752548": "TensorFlow published a guide to help with imbalanced data [here][1]. They show examples of various solutions including setting class weights and upsampling the tensorflow dataset.\n\n[1]: https://www.tensorflow.org/tutorials/structured_data/imbalanced_data",
    "752590": "Thank you! The problem here is that I cannot change the loss function I created and I used once. Even if I change it, it seems that the algorithm still points on the loss I created and not on the one I set.",
    "752604": "Modifying loss is one way to deal with imbalance. Another way is to leave the loss as it is, and resample the dataset so you have an equal proportion of each type of flower. And there are other ways described in the link I posted above.",
    "752613": "Thank you for your answer. The problem is not that, but the fact that, even if I do set a different loss, it seems that the kernel will always point on the old one. I did at first, using my loss :\n\n```\nmodel.compile(\n    optimizer='adam',\n    loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n```\n\nThen, `model.fit` and run.\n\nAfterwards, I changed : \n\n```\nmodel.compile(\n    optimizer='adam',\n    loss = 'sparse_categorical_crossentropy',\n    #loss=home_made_loss,\n    metrics=['sparse_categorical_accuracy']\n)\n```\n\nThen `model.fit` and run. But even now, it seems that, when running, the kernel still points on `home_made_loss`, as the behaviour is the one of `home_made_loss` and not the one of `sparse_categorical_crossentropy`.\n\nMy question is : What shall I do in order to use `sparse_categorical_crossentropy` again ?",
    "755279": "Thanks for the example!  I'm  going to run through this myself.",
    "755422": "It is easy to lose track of what is defined and used where in a Notebook. Have you tried restarting the notebook and running it from the beginning again ? If you are looking for a less extreme solution, I usually re-run all model creation cells when I change anything to my model.",
    "755428": "I tried having a look at your notebook but it has errors in many cell outputs. Ping me when you get this in a more runnable shape and I can have a look again. As for the code using the wrong loss, your best bet is to re-run the notebook from scratch. The ability to execute notebooks cells out of order can sometimes be confusing.",
    "755443": "I shall make my notebook runnable, at the beginning I just wanted to present a function in it, and not run it. \nConcerning the wrong loss : I always run the notebook from scratch (run all cells), I do not like running cell per cell, I find it bad practice.\nI also stopped the notebook and started it again.",
    "755509": "",
    "756123": "Here, my script is runnable, I added all the required settings. Can you have a look please ? I've just done a commit.",
    "756642": "I tried to run it but `load_dataset` is not defined in your code.\nThe script I tried: https://www.kaggle.com/catadanna/data-balancing-solution-with-tpu\n\nAlso, the `class_weight` parameter in `model.fit()` can be used to assigned different weights to different classes. On TPU, this has to be a flat list of weights (not a dictionary as written in the Keras docs). No need to define a custom loss.",
    "756981": "I added the function. I could not test yesterday, I had an error related to a connection problem (not mine, Google perhaps). I see what you mean, it is explained [HERE](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131876#754366). \n\nThe problem is that even now, after having changed the loss, I have the same behaviour or my loss. I mean, as if I did not change it. I tried to create another notebook but I cannot have two notebooks with TPU.",
    "799465": "class_weight parameter in model.fit() can be used to assigned different weights to different classes, to deal with the problem of imbalanced data , when you have more simple for one class then another (exemple you have 400 image of class 1 and 20 image of class 2 and 10 image of class3 etc )\n\non keras you have to pass a dictionary to class_weight parameter in model.fit()\nexample class_weights ={0:50, 1:0.5, 2:20}\n               model.fit(....     ....  class_weight=class_weights)\n\nbut on TPU your have to pass flat list of weights (not a dictionary as written in the Keras docs)\nso you have just to convert your dictionary to a flat list \n\n****class_weights =[item for k in class_weights for item in (k, class_weights[k])]****\n\n**No need to define a custom loss.**",
    "799472": "Indeed. Someone opened a discussion on this matter. Too late for me, I run my program with the new loss and even when removing it, it seems that I am still blocked on the loss I created. I mean, whatever I do, I set the original settings, including what you suggested (it was already mentioned in a discussion as I wrote before), the training has about 750 steps and takes a lot of time to parse one step, it takes 4 hours to parse 100 steps. I could not solve this and I am not allowed to open a new notebook (with TPU &amp;co) on this competition.",
    "799744": "pouvez vous partager votre kernel afin que je puisse jeter un coup d'oeil , peut etre je peux aider et par la meme aucasion aprendre quelque chose de nouveau",
    "824496": "And did it work..?",
    "825174": "No, couldn't debug yet."
  },
  "source": "meta"
}