{
  "id": 232804,
  "title": "How To Calculate Metrics Accurately?",
  "url": "/competitions/bms-molecular-translation/discussion/232804",
  "author_name": "Darien Schettler",
  "post_date": "2021-04-15T13:45:16.152000",
  "votes": 5,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Hi there. I had a dumb question. I'm using a simple CNN-&gt;LSTM pipeline w/ Attention.</p>\n<hr>\n<p>In my training/validation loop I'm performing the decoding character by character. After each predicted character I update my accuracy.</p>\n<p>The problem is that I've <strong>defined a loss function that disregards padding</strong>. This means that my model will correctly predict the string ONLY UP UNTIL IT PREDICTS THE END TOKEN. At which point it will just predict the remaining tokens as gibberish. This is completely fine from a prediction standpoint as I would remove everything after the END token anyway.</p>\n<p>However, when I calculate accuracy, all the gibberish lowers my score (same for LD calculation) as it scores negatively when comparing my gibberish to the PAD token.</p>\n<p>i.e. Let's assume I am doing evaluation along the string all at once… even though during training I do it one character at a time.</p>\n<pre><code># '2' is the &lt;END&gt; token and '0' is the &lt;PAD&gt; token\nground_truth = [15, 11, 16, 26, 88, 2, 0, 0, 0, 0]\nprediction   = [15, 11, 16, 26, 89, 2, 44, 50, 2, 15]\n\n# Oversimplified\nraw_accuracy = [1.0, 1.0, 1.0, 1.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0]\n\n# This includes the &lt;END&gt; token\nmean_accuracy   = 0.50\nmean_relevant_accuracy = 83.33 # only up to and including &lt;END&gt; token\n</code></pre>\n<p>Is there some way to calculate the accuracy using masking? I would need to probably wrap the existing metric <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/metrics/SparseCategoricalAccuracy\" target=\"_blank\">(<strong><code>SparseCategoricalAccuracy</code></strong>)</a>.</p>\n<p>But the same problem arises w.r.t. calculating Levenshtein distance. My current workaround is to use the ground truth padding tokens as a mask over irrelevant data… this works but if my model predicts sequences that are longer than the GT this error will be masked.</p>\n<hr>\n<p>Does anyone have any suggestions? </p>\n<hr>\n<p><em>I'll be sharing a notebook in the next couple of days (once my TPU quota resets). It will show my training with the broken metrics. Hopefully, we can fix it together!</em></p>",
  "messages": [
    {
      "id": 1274642,
      "postDate": "2021-04-15T13:45:16.153Z",
      "content": "<p>Hi there. I had a dumb question. I'm using a simple CNN-&gt;LSTM pipeline w/ Attention.</p>\n<hr>\n<p>In my training/validation loop I'm performing the decoding character by character. After each predicted character I update my accuracy.</p>\n<p>The problem is that I've <strong>defined a loss function that disregards padding</strong>. This means that my model will correctly predict the string ONLY UP UNTIL IT PREDICTS THE END TOKEN. At which point it will just predict the remaining tokens as gibberish. This is completely fine from a prediction standpoint as I would remove everything after the END token anyway.</p>\n<p>However, when I calculate accuracy, all the gibberish lowers my score (same for LD calculation) as it scores negatively when comparing my gibberish to the PAD token.</p>\n<p>i.e. Let's assume I am doing evaluation along the string all at once… even though during training I do it one character at a time.</p>\n<pre><code># '2' is the &lt;END&gt; token and '0' is the &lt;PAD&gt; token\nground_truth = [15, 11, 16, 26, 88, 2, 0, 0, 0, 0]\nprediction   = [15, 11, 16, 26, 89, 2, 44, 50, 2, 15]\n\n# Oversimplified\nraw_accuracy = [1.0, 1.0, 1.0, 1.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0]\n\n# This includes the &lt;END&gt; token\nmean_accuracy   = 0.50\nmean_relevant_accuracy = 83.33 # only up to and including &lt;END&gt; token\n</code></pre>\n<p>Is there some way to calculate the accuracy using masking? I would need to probably wrap the existing metric <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/metrics/SparseCategoricalAccuracy\" target=\"_blank\">(<strong><code>SparseCategoricalAccuracy</code></strong>)</a>.</p>\n<p>But the same problem arises w.r.t. calculating Levenshtein distance. My current workaround is to use the ground truth padding tokens as a mask over irrelevant data… this works but if my model predicts sequences that are longer than the GT this error will be masked.</p>\n<hr>\n<p>Does anyone have any suggestions? </p>\n<hr>\n<p><em>I'll be sharing a notebook in the next couple of days (once my TPU quota resets). It will show my training with the broken metrics. Hopefully, we can fix it together!</em></p>",
      "rawMarkdown": "Hi there. I had a dumb question. I'm using a simple CNN->LSTM pipeline w/ Attention.\n\n---\n\nIn my training/validation loop I'm performing the decoding character by character. After each predicted character I update my accuracy.\n\nThe problem is that I've **defined a loss function that disregards padding**. This means that my model will correctly predict the string ONLY UP UNTIL IT PREDICTS THE END TOKEN. At which point it will just predict the remaining tokens as gibberish. This is completely fine from a prediction standpoint as I would remove everything after the END token anyway.\n\nHowever, when I calculate accuracy, all the gibberish lowers my score (same for LD calculation) as it scores negatively when comparing my gibberish to the PAD token.\n\ni.e. Let's assume I am doing evaluation along the string all at once... even though during training I do it one character at a time.\n\n```python\n\n# '2' is the <END> token and '0' is the <PAD> token\nground_truth = [15, 11, 16, 26, 88, 2, 0, 0, 0, 0]\nprediction   = [15, 11, 16, 26, 89, 2, 44, 50, 2, 15]\n\n# Oversimplified\nraw_accuracy = [1.0, 1.0, 1.0, 1.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0]\n\n# This includes the <END> token\nmean_accuracy   = 0.50\nmean_relevant_accuracy = 83.33 # only up to and including <END> token\n```\n\nIs there some way to calculate the accuracy using masking? I would need to probably wrap the existing metric [(**`SparseCategoricalAccuracy`**)](https://www.tensorflow.org/api_docs/python/tf/keras/metrics/SparseCategoricalAccuracy).\n\nBut the same problem arises w.r.t. calculating Levenshtein distance. My current workaround is to use the ground truth padding tokens as a mask over irrelevant data... this works but if my model predicts sequences that are longer than the GT this error will be masked.\n\n---\n\nDoes anyone have any suggestions? \n\n---\n\n*I'll be sharing a notebook in the next couple of days (once my TPU quota resets). It will show my training with the broken metrics. Hopefully, we can fix it together!*",
      "votes": 5
    },
    {
      "id": 1278177,
      "postDate": "2021-04-19T16:08:11.760Z",
      "content": "<p>Is this closer to what you are looking for?</p>\n<pre><code>preds = tensor([[5, 6, 2, 2, 3, 7, 4, 5],\n                [6, 7, 7, 5, 7, 5, 2, 6],\n                [6, 5, 6, 3, 4, 3, 6, 6]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -&gt; tensor([2, 6, 0])\nidxs[idxs==0] = preds.size(1)\n# idxs -&gt; tensor([2, 6, 8])\npred_mask = (torch.arange(preds.size(1)).unsqueeze(0) &lt;= idxs.unsqueeze(1)) * 1.\n# pred_mask -&gt; tensor([[1., 1., 1., 0., 0., 0., 0., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 1.]])\nacc_mask = ((ground_truth!=0) + pred_mask).clamp(max=1.)\n</code></pre>",
      "rawMarkdown": "Is this closer to what you are looking for?\n```\npreds = tensor([[5, 6, 2, 2, 3, 7, 4, 5],\n                [6, 7, 7, 5, 7, 5, 2, 6],\n                [6, 5, 6, 3, 4, 3, 6, 6]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -> tensor([2, 6, 0])\nidxs[idxs==0] = preds.size(1)\n# idxs -> tensor([2, 6, 8])\npred_mask = (torch.arange(preds.size(1)).unsqueeze(0) <= idxs.unsqueeze(1)) * 1.\n# pred_mask -> tensor([[1., 1., 1., 0., 0., 0., 0., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 1.]])\nacc_mask = ((ground_truth!=0) + pred_mask).clamp(max=1.)\n```",
      "votes": 1,
      "replies": [
        {
          "id": 1278182,
          "postDate": "2021-04-19T16:11:48.387Z",
          "content": "<p>This looks VERY good. Would this be able to handle multiple predicted end tokens? </p>\n<p>i.e. I notice sometimes my preds tensor has the proper END token… but in and amongst the gibberish it may also predict end tokens.</p>\n<p>Thank you very much for your response though! I will probably use something like this!</p>",
          "rawMarkdown": "This looks VERY good. Would this be able to handle multiple predicted end tokens? \n\ni.e. I notice sometimes my preds tensor has the proper END token... but in and amongst the gibberish it may also predict end tokens.\n\nThank you very much for your response though! I will probably use something like this!",
          "votes": 1
        },
        {
          "id": 1278194,
          "postDate": "2021-04-19T16:21:20.333Z",
          "content": "<p>It seems that for multiple end tokens, max returns the first occurence. I guess that's fortunate :)</p>\n<pre><code>preds = tensor([[5, 6, 2, 2, 3, 7, 4, 5]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -&gt; tensor([2])\n</code></pre>\n<p>One thing is, how should cases like this be handled:</p>\n<pre><code>ground_truth = [3,4,7,2,0,0]\nprediction   = [3,2,7,2,9,7]\n</code></pre>",
          "rawMarkdown": "It seems that for multiple end tokens, max returns the first occurence. I guess that's fortunate :)\n\n```\npreds = tensor([[5, 6, 2, 2, 3, 7, 4, 5]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -> tensor([2])\n```\n\nOne thing is, how should cases like this be handled:\n```\nground_truth = [3,4,7,2,0,0]\nprediction   = [3,2,7,2,9,7]\n```",
          "votes": 1
        },
        {
          "id": 1278199,
          "postDate": "2021-04-19T16:26:49.597Z",
          "content": "<p>Thank you so much! I think this is it! I will implement it and ensure that I give you credit in the notebook.</p>",
          "rawMarkdown": "Thank you so much! I think this is it! I will implement it and ensure that I give you credit in the notebook.",
          "votes": 1
        },
        {
          "id": 1278201,
          "postDate": "2021-04-19T16:28:36.197Z",
          "content": "<p>Thanks and welcome! :)</p>",
          "rawMarkdown": "Thanks and welcome! :)"
        },
        {
          "id": 1278207,
          "postDate": "2021-04-19T16:37:08.147Z",
          "content": "<p>No need for <code>idxs[idxs!=0] = idxs[idxs!=0] + 1</code> if <code>&lt;=</code> is used to create <code>pred_mask</code>.<br>\nI updated above code accordingly.</p>",
          "rawMarkdown": "No need for `idxs[idxs!=0] = idxs[idxs!=0] + 1` if `<=` is used to create `pred_mask`.\nI updated above code accordingly.",
          "votes": 1
        },
        {
          "id": 1278223,
          "postDate": "2021-04-19T16:58:15.830Z",
          "content": "<p>I would still like to know why you need that. Accuracy isn't a good metric for predictions of different lengths.</p>",
          "rawMarkdown": "I would still like to know why you need that. Accuracy isn't a good metric for predictions of different lengths."
        },
        {
          "id": 1278229,
          "postDate": "2021-04-19T17:05:54.610Z",
          "content": "<p>I would argue if it's a bad idea on different lengths. Is it a bad idea on autoregressive task though, that's another question.</p>\n<p>If it correlates well with levenshtein, maybe better than the loss then no, it's not a bad idea. Why? Because it is much faster calculate than it is to figure out the levenshtein distance.</p>\n<p>Every epoch I only inspect the validation loss, and calculate levenshtein at every N epochs.</p>",
          "rawMarkdown": "I would argue if it's a bad idea on different lengths. Is it a bad idea on autoregressive task though, that's another question.\n\nIf it correlates well with levenshtein, maybe better than the loss then no, it's not a bad idea. Why? Because it is much faster calculate than it is to figure out the levenshtein distance.\n\nEvery epoch I only inspect the validation loss, and calculate levenshtein at every N epochs.",
          "votes": 1
        },
        {
          "id": 1278252,
          "postDate": "2021-04-19T17:28:55.747Z",
          "content": "<p><code>maybe better than the loss</code><br>\nWell, then why not just use top-k accuracy? Normally you train with parallel teacher forcing. So your prediction and ground truth have the same length. And then use top-1 and e.g. top-4 accuracy and you have a approximate feedback for greedy &amp; beam-search performance.</p>\n<p>I don't understand why you would go the slow training way and predict tokens step by step. Because this is the only way you will get different lengths outside of inference.</p>\n<p>As for accuracy for different lengths: It doesn't tell you if your mistake was due to wrong predictions or the missing length. Especially if the length difference variance is high that isn't good feedback. You also can't use top-k accuracy where k &gt; 1. <br>\nYes, Levenshtein distance also doesn't give you a good feedback, but it at least tells you the absolute difference. So a better top-1 accuracy. Best would be to take the intermediate parts calculated for Levenshtein distance and display them. Meaning, how many tokens had to be deleted/added and based on them you also can calculate the amount of correct ones. And then normalize those numbers by ground-truth length. </p>",
          "rawMarkdown": "`maybe better than the loss`\nWell, then why not just use top-k accuracy? Normally you train with parallel teacher forcing. So your prediction and ground truth have the same length. And then use top-1 and e.g. top-4 accuracy and you have a approximate feedback for greedy & beam-search performance.\n\nI don't understand why you would go the slow training way and predict tokens step by step. Because this is the only way you will get different lengths outside of inference.\n\nAs for accuracy for different lengths: It doesn't tell you if your mistake was due to wrong predictions or the missing length. Especially if the length difference variance is high that isn't good feedback. You also can't use top-k accuracy where k > 1. \nYes, Levenshtein distance also doesn't give you a good feedback, but it at least tells you the absolute difference. So a better top-1 accuracy. Best would be to take the intermediate parts calculated for Levenshtein distance and display them. Meaning, how many tokens had to be deleted/added and based on them you also can calculate the amount of correct ones. And then normalize those numbers by ground-truth length. "
        },
        {
          "id": 1278266,
          "postDate": "2021-04-19T17:47:11.047Z",
          "content": "<blockquote>\n  <p>this is the only way you will get different lengths outside of inference</p>\n</blockquote>\n<p>No. You can argmax your predictions.</p>",
          "rawMarkdown": "> this is the only way you will get different lengths outside of inference\n\nNo. You can argmax your predictions."
        },
        {
          "id": 1278268,
          "postDate": "2021-04-19T17:50:54.673Z",
          "content": "<p><code>argmax your predictions.</code><br>\nWhat do you mean by that? Just taking the most probable token prediction (index). And using that as your predicted sequence?<br>\nBut why would you want to do that vs. just using accuracy on the given predictions. I makes more sense. Because then you actually have valid feedback that takes your length into account anyway.<br>\nAnd you can get more than top-1 accuracy.</p>\n<p>TLDR: Yes, your are correct, you can do that. But what is the benefit of it for train/valid feedback?</p>",
          "rawMarkdown": "`argmax your predictions.`\nWhat do you mean by that? Just taking the most probable token prediction (index). And using that as your predicted sequence?\nBut why would you want to do that vs. just using accuracy on the given predictions. I makes more sense. Because then you actually have valid feedback that takes your length into account anyway.\nAnd you can get more than top-1 accuracy.\n\nTLDR: Yes, your are correct, you can do that. But what is the benefit of it for train/valid feedback?"
        },
        {
          "id": 1278279,
          "postDate": "2021-04-19T17:56:50.043Z",
          "content": "<p>Teacher forcing output -&gt; calculate loss; calculate accuracy (&lt;- here you argmax for \"accuracy\")</p>",
          "rawMarkdown": "Teacher forcing output -> calculate loss; calculate accuracy (<- here you argmax for \"accuracy\")",
          "votes": 1
        },
        {
          "id": 1278292,
          "postDate": "2021-04-19T18:03:06.217Z",
          "content": "<p>Yes, for top-1 accuracy you (can) do that. <br>\nBut that doesn't answer the question why you want to go to these lengths where you want a special function for your predictions of different lengths, where my original proposal of masked accuracy works just fine.<br>\nI have the feeling we are talking past each other and don't understand what the other truly means.</p>",
          "rawMarkdown": "Yes, for top-1 accuracy you (can) do that. \nBut that doesn't answer the question why you want to go to these lengths where you want a special function for your predictions of different lengths, where my original proposal of masked accuracy works just fine.\nI have the feeling we are talking past each other and don't understand what the other truly means."
        },
        {
          "id": 1278299,
          "postDate": "2021-04-19T18:07:05.783Z",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> - I'm using this for both accuracy and tokenwise Levenshtein distance (using <code>tf.edit_distance</code>).</p>\n<p></p>",
          "rawMarkdown": "@cepheidq - I'm using this for both accuracy and tokenwise Levenshtein distance (using `tf.edit_distance`).\n\n~~On a side note, as indicated by **nofreewill**, I am currently using argmax for accuracy. During inference, I will most likely use beam search. (although argmax would suffice for a rudimentary inference solution)~~"
        },
        {
          "id": 1278303,
          "postDate": "2021-04-19T18:12:41.557Z",
          "content": "<p>So you do generate the output, not just doing a teacher forcing? Then I was mistaken.</p>\n<p>But I feel like there is so much potential for misunderstanding around this topic : D</p>",
          "rawMarkdown": "So you do generate the output, not just doing a teacher forcing? Then I was mistaken.\n\nBut I feel like there is so much potential for misunderstanding around this topic : D"
        },
        {
          "id": 1278306,
          "postDate": "2021-04-19T18:16:29.657Z",
          "content": "<p>haha no I am using teacher forcing. I'm a dummy. I only use argmax during inference currently. I updated my previous comment.</p>\n<pre><code>def train_step(_image_batch, _inchi_batch):\n    \"\"\" Forward pass (calculate and update gradients)\"\"\"\n\n    batch_loss = tf.constant(0.0, tf.float32)   \n    with tf.GradientTape() as tape:\n        # image_batch_embedding has shape --&gt; (REPLICA_BATCH_SIZE, IMG_EMB_DIM)\n        image_batch_embedding = encoder(_image_batch, training=True)\n\n        # hidden and memory both have the shape --&gt; (REPLICA_BATCH_SIZE, N_RNN_UNITS)\n        hidden_batch, memory_batch = decoder.init_hidden_state(image_batch_embedding, training=True)\n\n        # decoder_input has a shape --&gt; (REPLICA_BATCH_SIZE, 1)\n        decoder_input_batch = tf.ones((REPLICA_BATCH_SIZE, 1), dtype=tf.uint8)\n\n        # Teacher forcing - feeding the target as the next input\n        for c_idx in range(1, MAX_LEN):\n            gt_batch = _inchi_batch[:, c_idx]\n\n            # passing enc_output to the decoder\n            prediction_batch, hidden_batch, memory_batch = \\\n                decoder(decoder_input_batch, hidden_batch, memory_batch, image_batch_embedding, training=True)\n\n            # Update Loss Accumulator\n            batch_loss += loss_fn(gt_batch, prediction_batch)\n\n            # Update Accuracy Metric\n            metrics[\"train_acc\"].update_state(gt_batch, prediction_batch, \n                                              sample_weight=tf.where(tf.not_equal(gt_batch, PAD_TOKEN), 1.0, 0.0))\n\n            # teacher forcing, use correct character as next input to LSTMCell\n            decoder_input_batch = tf.expand_dims(gt_batch, 1)\n\n    # backpropagation using variables, gradients and loss\n    #    - split this into two seperate optimizers/lrs/etc in the future\n    #    - we use the batch loss accumulation to update gradients\n    gradients = tape.gradient(batch_loss, encoder.trainable_variables + decoder.trainable_variables)\n\n    # Normalize loss across all characters    \n    batch_loss = batch_loss/(MAX_LEN-1)\n\n    metrics[\"batch_loss\"].update_state(batch_loss)\n    metrics[\"train_loss\"].update_state(batch_loss)\n\n    optimizer.apply_gradients(zip(gradients, encoder.trainable_variables+decoder.trainable_variables))\n\n@tf.function\ndef dist_train_step(_image_batch, _inchi_batch):\n    strategy.run(train_step, args=(_image_batch, _inchi_batch))\n</code></pre>",
          "rawMarkdown": "haha no I am using teacher forcing. I'm a dummy. I only use argmax during inference currently. I updated my previous comment.\n\n```python\n\ndef train_step(_image_batch, _inchi_batch):\n    \"\"\" Forward pass (calculate and update gradients)\"\"\"\n    \n    batch_loss = tf.constant(0.0, tf.float32)   \n    with tf.GradientTape() as tape:\n        # image_batch_embedding has shape --> (REPLICA_BATCH_SIZE, IMG_EMB_DIM)\n        image_batch_embedding = encoder(_image_batch, training=True)\n        \n        # hidden and memory both have the shape --> (REPLICA_BATCH_SIZE, N_RNN_UNITS)\n        hidden_batch, memory_batch = decoder.init_hidden_state(image_batch_embedding, training=True)\n        \n        # decoder_input has a shape --> (REPLICA_BATCH_SIZE, 1)\n        decoder_input_batch = tf.ones((REPLICA_BATCH_SIZE, 1), dtype=tf.uint8)\n        \n        # Teacher forcing - feeding the target as the next input\n        for c_idx in range(1, MAX_LEN):\n            gt_batch = _inchi_batch[:, c_idx]\n            \n            # passing enc_output to the decoder\n            prediction_batch, hidden_batch, memory_batch = \\\n                decoder(decoder_input_batch, hidden_batch, memory_batch, image_batch_embedding, training=True)\n            \n            # Update Loss Accumulator\n            batch_loss += loss_fn(gt_batch, prediction_batch)\n            \n            # Update Accuracy Metric\n            metrics[\"train_acc\"].update_state(gt_batch, prediction_batch, \n                                              sample_weight=tf.where(tf.not_equal(gt_batch, PAD_TOKEN), 1.0, 0.0))\n            \n            # teacher forcing, use correct character as next input to LSTMCell\n            decoder_input_batch = tf.expand_dims(gt_batch, 1)\n\n    # backpropagation using variables, gradients and loss\n    #    - split this into two seperate optimizers/lrs/etc in the future\n    #    - we use the batch loss accumulation to update gradients\n    gradients = tape.gradient(batch_loss, encoder.trainable_variables + decoder.trainable_variables)\n    \n    # Normalize loss across all characters    \n    batch_loss = batch_loss/(MAX_LEN-1)\n    \n    metrics[\"batch_loss\"].update_state(batch_loss)\n    metrics[\"train_loss\"].update_state(batch_loss)\n    \n    optimizer.apply_gradients(zip(gradients, encoder.trainable_variables+decoder.trainable_variables))\n\n@tf.function\ndef dist_train_step(_image_batch, _inchi_batch):\n    strategy.run(train_step, args=(_image_batch, _inchi_batch))\n```"
        },
        {
          "id": 1278327,
          "postDate": "2021-04-19T18:47:07.477Z",
          "content": "<p> Ah, I see you are using sequential teacher forcing. But shouldn't you then abort/mask after the ground truth end? I don't understand the TF code well enough to determine the latter.</p>\n<p>I don't see how sequential teacher forcing can be really better then parallel one. Is it? Your get far more epochs per time with parallel one so you can make up for maybe a bit worse performance.</p>\n<p>Keep in mind that if you use Levenshtein distance directly on the token ids, you will get a wrong distance iff you have tokens that span more than one character. E.g. token id for 'Cl' = 13. So converting to string first would be better.</p>",
          "rawMarkdown": "~~If you are using parallel teacher forcing during training/validation I suggest using (Tensorflow equivalent of) my masked accuracy function I posted. That does the job as intended for all top-k accuracy. Not just top-1.~~ Ah, I see you are using sequential teacher forcing. But shouldn't you then abort/mask after the ground truth end? I don't understand the TF code well enough to determine the latter.\n\nI don't see how sequential teacher forcing can be really better then parallel one. Is it? Your get far more epochs per time with parallel one so you can make up for maybe a bit worse performance.\n\nKeep in mind that if you use Levenshtein distance directly on the token ids, you will get a wrong distance iff you have tokens that span more than one character. E.g. token id for 'Cl' = 13. So converting to string first would be better.",
          "votes": 1
        },
        {
          "id": 1278348,
          "postDate": "2021-04-19T19:25:14.887Z",
          "content": "<p><code>Every epoch I only inspect the validation loss</code><br>\n<a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> : I have observed higher top-1 &amp; top-4 accuracy despite higher loss . So you should be careful with that and prefer the Levenshtein distance every n epochs or also add accuracy to your metrics.</p>",
          "rawMarkdown": "`Every epoch I only inspect the validation loss`\n@nofreewill : I have observed higher top-1 & top-4 accuracy despite higher loss . So you should be careful with that and prefer the Levenshtein distance every n epochs or also add accuracy to your metrics."
        }
      ]
    },
    {
      "id": 1275988,
      "postDate": "2021-04-17T00:37:06.993Z",
      "content": "<p>You can just pad your sequences as given and then count the number of unpadded token-id positions. <br>\nThen you calculate how many of your predictions match the ground-truth. This matching count you divide by the unpadded id count. This will give you the true accuracy. <br>\nYou should mask the padding ids in your loss function, otherwise your model will spend some of its predictive power to learn useless patterns. Try the difference in accuracy after you have implemented the masked accuracy function.<br>\nI have here one for Pytorch, maybe you still understand it. Could also be helpful for other people. I have added comments.</p>\n<p>Adapted from <a href=\"https://github.com/catalyst-team/catalyst/blob/master/catalyst/metrics/functional/_accuracy.py\" target=\"_blank\">https://github.com/catalyst-team/catalyst/blob/master/catalyst/metrics/functional/_accuracy.py</a><br>\nTherefore the licence notice. I added the masked accuracy functionality:<br>\nCopyright 2018 Sergey Kolesnikov</p>\n<p>Licensed under the Apache License, Version 2.0 (the \"License\");<br>\n   you may not use this file except in compliance with the License.<br>\n   You may obtain a copy of the License at</p>\n<pre><code>   http://www.apache.org/licenses/LICENSE-2.0\n</code></pre>\n<p>Unless required by applicable law or agreed to in writing, software<br>\n   distributed under the License is distributed on an \"AS IS\" BASIS,<br>\n   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.<br>\n   See the License for the specific language governing permissions and<br>\n   limitations under the License.</p>\n<p>`mask_idx = int(-sys.maxsize) if mask_idx is None else mask_idx   # Your padding index id</p>\n<pre><code>    max_k = max(topk)\n    # Here you count the number of unpadded token positions:\n    batch_size = (targets != mask_idx).to(dtype=torch.uint8).sum(dim=0).item()\n\n    if len(outputs.shape) == 1 or outputs.shape[1] == 1:\n        # binary accuracy\n        pred = outputs.t()\n    else:\n        # multiclass accuracy\n        _, pred = outputs.topk(max_k, 1, True, True)  # noqa: WPS425\n        pred = pred.t()\n    # Here you compute which token predictions where correct and which not:\n    correct = pred.eq(targets.long().view(1, -1).expand_as(pred))\n    # Calculate mask, which shows all unpadded positions (as True):\n    mask = (targets.long() != mask_idx)\n    mask = mask.view(1, -1)\n\n    output = []\n    for k in topk:\n        # Here you only return those predictions which correspond to a unpadded position: \n        correct_k = torch.masked_select(correct[:k], mask)\n        correct_k = correct_k.float().sum(0, keepdim=True)\n        output.append(correct_k.mul_(1.0 / batch_size))`   # Normalization by unpadded position count\n       # output contains the accuracies for our batch &amp; top-k \n       # Note, that you shouldn't place the function at a position where a gradient might be calculated\n       # OR use with 'torch.no_grad():' and outputs.detach() to be safe.`\n</code></pre>\n<p>(Maybe one day I will get the code blocks to work properly)</p>",
      "rawMarkdown": "You can just pad your sequences as given and then count the number of unpadded token-id positions. \nThen you calculate how many of your predictions match the ground-truth. This matching count you divide by the unpadded id count. This will give you the true accuracy. \nYou should mask the padding ids in your loss function, otherwise your model will spend some of its predictive power to learn useless patterns. Try the difference in accuracy after you have implemented the masked accuracy function.\nI have here one for Pytorch, maybe you still understand it. Could also be helpful for other people. I have added comments.\n\nAdapted from https://github.com/catalyst-team/catalyst/blob/master/catalyst/metrics/functional/_accuracy.py\nTherefore the licence notice. I added the masked accuracy functionality:\nCopyright 2018 Sergey Kolesnikov\n\n   Licensed under the Apache License, Version 2.0 (the \"License\");\n   you may not use this file except in compliance with the License.\n   You may obtain a copy of the License at\n\n       http://www.apache.org/licenses/LICENSE-2.0\n\n   Unless required by applicable law or agreed to in writing, software\n   distributed under the License is distributed on an \"AS IS\" BASIS,\n   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n   See the License for the specific language governing permissions and\n   limitations under the License.\n\n`mask_idx = int(-sys.maxsize) if mask_idx is None else mask_idx   # Your padding index id\n\n        max_k = max(topk)\n        # Here you count the number of unpadded token positions:\n        batch_size = (targets != mask_idx).to(dtype=torch.uint8).sum(dim=0).item()\n\n        if len(outputs.shape) == 1 or outputs.shape[1] == 1:\n            # binary accuracy\n            pred = outputs.t()\n        else:\n            # multiclass accuracy\n            _, pred = outputs.topk(max_k, 1, True, True)  # noqa: WPS425\n            pred = pred.t()\n        # Here you compute which token predictions where correct and which not:\n        correct = pred.eq(targets.long().view(1, -1).expand_as(pred))\n        # Calculate mask, which shows all unpadded positions (as True):\n        mask = (targets.long() != mask_idx)\n        mask = mask.view(1, -1)\n\n        output = []\n        for k in topk:\n            # Here you only return those predictions which correspond to a unpadded position: \n            correct_k = torch.masked_select(correct[:k], mask)\n            correct_k = correct_k.float().sum(0, keepdim=True)\n            output.append(correct_k.mul_(1.0 / batch_size))`   # Normalization by unpadded position count\n           # output contains the accuracies for our batch & top-k \n           # Note, that you shouldn't place the function at a position where a gradient might be calculated\n           # OR use with 'torch.no_grad():' and outputs.detach() to be safe.`\n\n(Maybe one day I will get the code blocks to work properly)",
      "votes": 1,
      "replies": [
        {
          "id": 1277486,
          "postDate": "2021-04-18T19:31:54.363Z",
          "content": "<p>This makes sense. I was leaning towards a solution like this.</p>\n<hr>\n<p>However, it isn't quite accurate. Let me give an example for clarity.</p>\n<pre><code># '1' is the &lt;START&gt; token, '2' is the &lt;END&gt; token and '0' is the &lt;PAD&gt; token\nground_truth = [1, 15, 11, 16, 26, 88, 2, 0, 0, 0]\nprediction   = [1, 15, 11, 16, 26, 88, 44, 50, 2, 15]\n\n# Mask Function as Per Your Logic\nmask = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0]\nmasked_prediction = [1, 15, 11, 16, 26, 88, 44, 0, 0, 0]\nmask_accuracy = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0]\nave_accuracy = sum(mask_accuracy)/sum(mask) # 6/7=0.857\n\n# Real accuracy should compute the accuracy of all values up to the\n# &lt;END&gt; token in the prediction string (not the gt string)\nreal_mask = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0]\nreal_masked_prediction = [1, 15, 11, 16, 26, 88, 44, 50, 2, 0]\nreal_mask_accuracy = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0]\nreal_ave_accuracy = sum(real_mask_accuracy)/sum(real_mask) # 6/9=0.667\n</code></pre>\n<hr>\n<p><strong>As you can see, if we only mask up to the ground truth <em>END</em> token we won't be calculating the real accuracy when the predicted sequence is longer than the ground truth sequence</strong></p>\n<hr>\n<p>All this being said, as this is relatively easy to implement and will give a good feel for accuracy I will most likely go with something like this anyway.</p>",
          "rawMarkdown": "This makes sense. I was leaning towards a solution like this.\n\n---\n\nHowever, it isn't quite accurate. Let me give an example for clarity.\n\n```python\n# '1' is the <START> token, '2' is the <END> token and '0' is the <PAD> token\nground_truth = [1, 15, 11, 16, 26, 88, 2, 0, 0, 0]\nprediction   = [1, 15, 11, 16, 26, 88, 44, 50, 2, 15]\n\n# Mask Function as Per Your Logic\nmask = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0]\nmasked_prediction = [1, 15, 11, 16, 26, 88, 44, 0, 0, 0]\nmask_accuracy = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0]\nave_accuracy = sum(mask_accuracy)/sum(mask) # 6/7=0.857\n\n# Real accuracy should compute the accuracy of all values up to the\n# <END> token in the prediction string (not the gt string)\nreal_mask = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0]\nreal_masked_prediction = [1, 15, 11, 16, 26, 88, 44, 50, 2, 0]\nreal_mask_accuracy = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0]\nreal_ave_accuracy = sum(real_mask_accuracy)/sum(real_mask) # 6/9=0.667\n```\n\n---\n\n**As you can see, if we only mask up to the ground truth *END* token we won't be calculating the real accuracy when the predicted sequence is longer than the ground truth sequence**\n\n---\n\nAll this being said, as this is relatively easy to implement and will give a good feel for accuracy I will most likely go with something like this anyway."
        },
        {
          "id": 1277517,
          "postDate": "2021-04-18T20:53:51.200Z",
          "content": "<p>It wasn't meant to work for comparison with sequences of different lengths. Just top-k classification accuracy like in e.g. image classification. For different lengths I would use Levenshtein distance, makes way more sense.</p>",
          "rawMarkdown": "It wasn't meant to work for comparison with sequences of different lengths. Just top-k classification accuracy like in e.g. image classification. For different lengths I would use Levenshtein distance, makes way more sense."
        },
        {
          "id": 1277519,
          "postDate": "2021-04-18T20:59:04.623Z",
          "content": "<p>Levenshtein distance faces the same problems as the accuracy calculation above. I appreciate your help though! </p>\n<p>I think something like what you suggested is the best I can do right now.</p>\n<hr>\n<p><strong><em>[EDIT] - Did you downvote me because I tried to explain what I was looking for? Seems a bit harsh…</em></strong></p>",
          "rawMarkdown": "Levenshtein distance faces the same problems as the accuracy calculation above. I appreciate your help though! \n\nI think something like what you suggested is the best I can do right now.\n\n---\n\n***[EDIT] - Did you downvote me because I tried to explain what I was looking for? Seems a bit harsh...***"
        },
        {
          "id": 1277996,
          "postDate": "2021-04-19T13:18:53.170Z",
          "content": "<p>Well you can  [see comment below] as described for the accuracy. Although I am very confused why you have potentially two vectors of separate lengths. Do you generate all sequences from scratch during training, instead of parallelism? That would be really slow.<br>\nI mean during inference you (should) convert the outputs back to strings and then use Levenshtein distance. And your tokenizer obviously ignores the padding token. So there is no problem.</p>\n<p><code>Did you downvote me because I tried to explain what I was looking for? Seems a bit harsh…</code><br>\nNo. Was a mistake. I fixed that.</p>",
          "rawMarkdown": "Well you can ~~mask the vector ~~ [see comment below] as described for the accuracy. Although I am very confused why you have potentially two vectors of separate lengths. Do you generate all sequences from scratch during training, instead of parallelism? That would be really slow.\nI mean during inference you (should) convert the outputs back to strings and then use Levenshtein distance. And your tokenizer obviously ignores the padding token. So there is no problem.\n\n`Did you downvote me because I tried to explain what I was looking for? Seems a bit harsh…`\nNo. Was a mistake. I fixed that."
        },
        {
          "id": 1278094,
          "postDate": "2021-04-19T14:52:07.787Z",
          "content": "<p>If you want to calculate Levenshtein distance without transforming to strings first (keep in mind that some tokens may make up 2 characters or more) you could to as follows: <br>\nTensorflow already as an edit distance function, so you don't need to implement that part.<br>\nTherefore you replace all tokens in your prediction X after the end/eos token with the padding token id. If your ground truth Y is shorter than X, you pad it with the padding id, until it has the same length. Then you use the edit distance function.<br>\n<br>\nYou still have to correct it by length difference, -_-</p>\n<p>Example:<br>\nX = predictions; Y = ground truth; 0 = padding id; 1 = end/eos id</p>\n<ol>\n<li>Our vectors:<br>\nX = [ [2, 5, <strong>1</strong>,  7, 3], [2, 4, 3, 9, <strong>1</strong>] ]<br>\nY = [ [2, 5, 6, <strong>1</strong>], [2, 4, 3, <strong>1</strong>]]</li>\n<li>Replace token ids after eos with padding id:<br>\nX = [ [2, 5, 1,  <strong>0</strong>, <strong>0</strong>], [2, 4, 3, 9, <strong>1</strong>] ]</li>\n<li>Pad X &amp; Y to same length:<br>\nY = [ [2, 5, 6,  1, <strong>0</strong>], [2, 4, 3, 1, <strong>0</strong>]]</li>\n<li>Compute distance<br>\nX1/Y1 = 2<br>\nX2/Y2 = 2</li>\n<li>Correct by absolute length difference:<br>\nX1/Y1 = 2 -1 = 1<br>\nX2/Y2 = 2 -1 = 1</li>\n</ol>\n<p>How to replace ids after eos id? Match for eos id, and take index of first occurence in vector, then replace all positions afterward with padding id: X[0, (eos_id +1):] = pad_id. There might be better, inbuild ways to do that.</p>\n<p>I think due to the necessary length difference correction you can just compute edit distance directly, then find the first eos positions, compute true length of X, and then correct the distance calculation.</p>",
          "rawMarkdown": "If you want to calculate Levenshtein distance without transforming to strings first (keep in mind that some tokens may make up 2 characters or more) you could to as follows: \nTensorflow already as an edit distance function, so you don't need to implement that part.\nTherefore you replace all tokens in your prediction X after the end/eos token with the padding token id. If your ground truth Y is shorter than X, you pad it with the padding id, until it has the same length. Then you use the edit distance function.\n~~That should work, because levenshtein/edit distance doesn't measure relative changes (depending on length). So if you have to pad your Y to match length of X, where then X isn't padded, you get the same amount of distance as with two vectors of different lengths.~~\nYou still have to correct it by length difference, -_-\n\nExample:\nX = predictions; Y = ground truth; 0 = padding id; 1 = end/eos id\n0. Our vectors:\nX = [ [2, 5, **1**,  7, 3], [2, 4, 3, 9, **1**] ]\nY = [ [2, 5, 6, **1**], [2, 4, 3, **1**]]\n1. Replace token ids after eos with padding id:\nX = [ [2, 5, 1,  **0**, **0**], [2, 4, 3, 9, **1**] ]\n2. Pad X & Y to same length:\nY = [ [2, 5, 6,  1, **0**], [2, 4, 3, 1, **0**]]\n3. Compute distance\nX1/Y1 = 2\nX2/Y2 = 2\n4. Correct by absolute length difference:\nX1/Y1 = 2 -1 = 1\nX2/Y2 = 2 -1 = 1\n\nHow to replace ids after eos id? Match for eos id, and take index of first occurence in vector, then replace all positions afterward with padding id: X[0, (eos_id +1):] = pad_id. There might be better, inbuild ways to do that.\n\nI think due to the necessary length difference correction you can just compute edit distance directly, then find the first eos positions, compute true length of X, and then correct the distance calculation."
        }
      ]
    },
    {
      "id": 1275410,
      "postDate": "2021-04-16T09:26:53.960Z",
      "content": "<p>Your model should learn to always predict &lt;pad&gt; after an &lt;end&gt; token fast, this should therefore not cause any problems. I am using the same model architecture as you (CNN encoder and LSTM decoder with attention) and in my experience you should see you model consistently predict &lt;pad&gt; tokens after an &lt;end&gt; token in a few hundred training steps.</p>\n<p>Conceptually this also makes sense, predicting the correct InChI is a complex task, learning to always say &lt;pad&gt; after you see an &lt;end&gt; token is quite simple ;)</p>",
      "rawMarkdown": "Your model should learn to always predict <pad\\> after an <end\\> token fast, this should therefore not cause any problems. I am using the same model architecture as you (CNN encoder and LSTM decoder with attention) and in my experience you should see you model consistently predict <pad\\> tokens after an <end\\> token in a few hundred training steps.\n\nConceptually this also makes sense, predicting the correct InChI is a complex task, learning to always say <pad\\> after you see an <end\\> token is quite simple ;)",
      "votes": 1,
      "replies": [
        {
          "id": 1275581,
          "postDate": "2021-04-16T13:13:00.147Z",
          "content": "<p><strong>My current loss function masks out the padding token</strong>. This means my model never learns to predict the PAD token after the END token. See below:</p>\n<pre><code>loss_object = tf.keras.losses.SparseCategoricalCrossentropy(\n            from_logits=True, reduction=tf.keras.losses.Reduction.NONE\n        )\n\ndef loss_fn(real, pred):\n    mask = tf.math.not_equal(real, 0) # 0 is my pad token\n    loss_ = loss_object(real, pred)\n    loss_ *= tf.cast(mask, dtype=loss_.dtype)\n    loss_ = tf.nn.compute_average_loss(\n        loss_, global_batch_size=GLOBAL_BATCH_SIZE\n    )\n\n    return loss_\n</code></pre>\n<p>I initially used a loss function similar to yours. However, I'm predicting the full-lengths (max length is ~280 using my tokenization method) of the inchi… not a truncated version (as you are). The average inchi length (using my tokenization method) is around 90 tokens. This means that on average there are 190 PAD tokens in every INCHI string. Because of this massive class imbalance, the model will just start predicting PAD tokens all the time. I guess it would eventually break out of this local minima, however, I thought it made sense to just mask them out so the gradient is smoother.</p>\n<p>Hope this makes sense.</p>",
          "rawMarkdown": "**My current loss function masks out the padding token**. This means my model never learns to predict the PAD token after the END token. See below:\n\n```python\n\nloss_object = tf.keras.losses.SparseCategoricalCrossentropy(\n            from_logits=True, reduction=tf.keras.losses.Reduction.NONE\n        )\n        \ndef loss_fn(real, pred):\n    mask = tf.math.not_equal(real, 0) # 0 is my pad token\n    loss_ = loss_object(real, pred)\n    loss_ *= tf.cast(mask, dtype=loss_.dtype)\n    loss_ = tf.nn.compute_average_loss(\n        loss_, global_batch_size=GLOBAL_BATCH_SIZE\n    )\n\n    return loss_\n\n```\n\nI initially used a loss function similar to yours. However, I'm predicting the full-lengths (max length is ~280 using my tokenization method) of the inchi... not a truncated version (as you are). The average inchi length (using my tokenization method) is around 90 tokens. This means that on average there are 190 PAD tokens in every INCHI string. Because of this massive class imbalance, the model will just start predicting PAD tokens all the time. I guess it would eventually break out of this local minima, however, I thought it made sense to just mask them out so the gradient is smoother.\n\nHope this makes sense."
        }
      ]
    },
    {
      "id": 1278151,
      "postDate": "2021-04-19T15:40:47.640Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1278177,
      "author_name": "nofreewill42",
      "author_url": "",
      "post_date": "2021-04-19T16:08:11.760000",
      "content": "<p>Is this closer to what you are looking for?</p>\n<pre><code>preds = tensor([[5, 6, 2, 2, 3, 7, 4, 5],\n                [6, 7, 7, 5, 7, 5, 2, 6],\n                [6, 5, 6, 3, 4, 3, 6, 6]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -&gt; tensor([2, 6, 0])\nidxs[idxs==0] = preds.size(1)\n# idxs -&gt; tensor([2, 6, 8])\npred_mask = (torch.arange(preds.size(1)).unsqueeze(0) &lt;= idxs.unsqueeze(1)) * 1.\n# pred_mask -&gt; tensor([[1., 1., 1., 0., 0., 0., 0., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 1.]])\nacc_mask = ((ground_truth!=0) + pred_mask).clamp(max=1.)\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1278182,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-19T16:11:48.387000",
          "content": "<p>This looks VERY good. Would this be able to handle multiple predicted end tokens? </p>\n<p>i.e. I notice sometimes my preds tensor has the proper END token… but in and amongst the gibberish it may also predict end tokens.</p>\n<p>Thank you very much for your response though! I will probably use something like this!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278194,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T16:21:20.333000",
          "content": "<p>It seems that for multiple end tokens, max returns the first occurence. I guess that's fortunate :)</p>\n<pre><code>preds = tensor([[5, 6, 2, 2, 3, 7, 4, 5]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -&gt; tensor([2])\n</code></pre>\n<p>One thing is, how should cases like this be handled:</p>\n<pre><code>ground_truth = [3,4,7,2,0,0]\nprediction   = [3,2,7,2,9,7]\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278199,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-19T16:26:49.597000",
          "content": "<p>Thank you so much! I think this is it! I will implement it and ensure that I give you credit in the notebook.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278201,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T16:28:36.197000",
          "content": "<p>Thanks and welcome! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278207,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T16:37:08.147000",
          "content": "<p>No need for <code>idxs[idxs!=0] = idxs[idxs!=0] + 1</code> if <code>&lt;=</code> is used to create <code>pred_mask</code>.<br>\nI updated above code accordingly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278223,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T16:58:15.830000",
          "content": "<p>I would still like to know why you need that. Accuracy isn't a good metric for predictions of different lengths.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278229,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T17:05:54.610000",
          "content": "<p>I would argue if it's a bad idea on different lengths. Is it a bad idea on autoregressive task though, that's another question.</p>\n<p>If it correlates well with levenshtein, maybe better than the loss then no, it's not a bad idea. Why? Because it is much faster calculate than it is to figure out the levenshtein distance.</p>\n<p>Every epoch I only inspect the validation loss, and calculate levenshtein at every N epochs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278252,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T17:28:55.747000",
          "content": "<p><code>maybe better than the loss</code><br>\nWell, then why not just use top-k accuracy? Normally you train with parallel teacher forcing. So your prediction and ground truth have the same length. And then use top-1 and e.g. top-4 accuracy and you have a approximate feedback for greedy &amp; beam-search performance.</p>\n<p>I don't understand why you would go the slow training way and predict tokens step by step. Because this is the only way you will get different lengths outside of inference.</p>\n<p>As for accuracy for different lengths: It doesn't tell you if your mistake was due to wrong predictions or the missing length. Especially if the length difference variance is high that isn't good feedback. You also can't use top-k accuracy where k &gt; 1. <br>\nYes, Levenshtein distance also doesn't give you a good feedback, but it at least tells you the absolute difference. So a better top-1 accuracy. Best would be to take the intermediate parts calculated for Levenshtein distance and display them. Meaning, how many tokens had to be deleted/added and based on them you also can calculate the amount of correct ones. And then normalize those numbers by ground-truth length. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278266,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T17:47:11.047000",
          "content": "<blockquote>\n  <p>this is the only way you will get different lengths outside of inference</p>\n</blockquote>\n<p>No. You can argmax your predictions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278268,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T17:50:54.673000",
          "content": "<p><code>argmax your predictions.</code><br>\nWhat do you mean by that? Just taking the most probable token prediction (index). And using that as your predicted sequence?<br>\nBut why would you want to do that vs. just using accuracy on the given predictions. I makes more sense. Because then you actually have valid feedback that takes your length into account anyway.<br>\nAnd you can get more than top-1 accuracy.</p>\n<p>TLDR: Yes, your are correct, you can do that. But what is the benefit of it for train/valid feedback?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278279,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T17:56:50.043000",
          "content": "<p>Teacher forcing output -&gt; calculate loss; calculate accuracy (&lt;- here you argmax for \"accuracy\")</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278292,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T18:03:06.217000",
          "content": "<p>Yes, for top-1 accuracy you (can) do that. <br>\nBut that doesn't answer the question why you want to go to these lengths where you want a special function for your predictions of different lengths, where my original proposal of masked accuracy works just fine.<br>\nI have the feeling we are talking past each other and don't understand what the other truly means.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278299,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-19T18:07:05.783000",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> - I'm using this for both accuracy and tokenwise Levenshtein distance (using <code>tf.edit_distance</code>).</p>\n<p></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278303,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-19T18:12:41.557000",
          "content": "<p>So you do generate the output, not just doing a teacher forcing? Then I was mistaken.</p>\n<p>But I feel like there is so much potential for misunderstanding around this topic : D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278306,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-19T18:16:29.657000",
          "content": "<p>haha no I am using teacher forcing. I'm a dummy. I only use argmax during inference currently. I updated my previous comment.</p>\n<pre><code>def train_step(_image_batch, _inchi_batch):\n    \"\"\" Forward pass (calculate and update gradients)\"\"\"\n\n    batch_loss = tf.constant(0.0, tf.float32)   \n    with tf.GradientTape() as tape:\n        # image_batch_embedding has shape --&gt; (REPLICA_BATCH_SIZE, IMG_EMB_DIM)\n        image_batch_embedding = encoder(_image_batch, training=True)\n\n        # hidden and memory both have the shape --&gt; (REPLICA_BATCH_SIZE, N_RNN_UNITS)\n        hidden_batch, memory_batch = decoder.init_hidden_state(image_batch_embedding, training=True)\n\n        # decoder_input has a shape --&gt; (REPLICA_BATCH_SIZE, 1)\n        decoder_input_batch = tf.ones((REPLICA_BATCH_SIZE, 1), dtype=tf.uint8)\n\n        # Teacher forcing - feeding the target as the next input\n        for c_idx in range(1, MAX_LEN):\n            gt_batch = _inchi_batch[:, c_idx]\n\n            # passing enc_output to the decoder\n            prediction_batch, hidden_batch, memory_batch = \\\n                decoder(decoder_input_batch, hidden_batch, memory_batch, image_batch_embedding, training=True)\n\n            # Update Loss Accumulator\n            batch_loss += loss_fn(gt_batch, prediction_batch)\n\n            # Update Accuracy Metric\n            metrics[\"train_acc\"].update_state(gt_batch, prediction_batch, \n                                              sample_weight=tf.where(tf.not_equal(gt_batch, PAD_TOKEN), 1.0, 0.0))\n\n            # teacher forcing, use correct character as next input to LSTMCell\n            decoder_input_batch = tf.expand_dims(gt_batch, 1)\n\n    # backpropagation using variables, gradients and loss\n    #    - split this into two seperate optimizers/lrs/etc in the future\n    #    - we use the batch loss accumulation to update gradients\n    gradients = tape.gradient(batch_loss, encoder.trainable_variables + decoder.trainable_variables)\n\n    # Normalize loss across all characters    \n    batch_loss = batch_loss/(MAX_LEN-1)\n\n    metrics[\"batch_loss\"].update_state(batch_loss)\n    metrics[\"train_loss\"].update_state(batch_loss)\n\n    optimizer.apply_gradients(zip(gradients, encoder.trainable_variables+decoder.trainable_variables))\n\n@tf.function\ndef dist_train_step(_image_batch, _inchi_batch):\n    strategy.run(train_step, args=(_image_batch, _inchi_batch))\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278327,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T18:47:07.477000",
          "content": "<p> Ah, I see you are using sequential teacher forcing. But shouldn't you then abort/mask after the ground truth end? I don't understand the TF code well enough to determine the latter.</p>\n<p>I don't see how sequential teacher forcing can be really better then parallel one. Is it? Your get far more epochs per time with parallel one so you can make up for maybe a bit worse performance.</p>\n<p>Keep in mind that if you use Levenshtein distance directly on the token ids, you will get a wrong distance iff you have tokens that span more than one character. E.g. token id for 'Cl' = 13. So converting to string first would be better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278348,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T19:25:14.887000",
          "content": "<p><code>Every epoch I only inspect the validation loss</code><br>\n<a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> : I have observed higher top-1 &amp; top-4 accuracy despite higher loss . So you should be careful with that and prefer the Levenshtein distance every n epochs or also add accuracy to your metrics.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1275988,
      "author_name": "Gabriel Lindenmaier",
      "author_url": "",
      "post_date": "2021-04-17T00:37:06.993000",
      "content": "<p>You can just pad your sequences as given and then count the number of unpadded token-id positions. <br>\nThen you calculate how many of your predictions match the ground-truth. This matching count you divide by the unpadded id count. This will give you the true accuracy. <br>\nYou should mask the padding ids in your loss function, otherwise your model will spend some of its predictive power to learn useless patterns. Try the difference in accuracy after you have implemented the masked accuracy function.<br>\nI have here one for Pytorch, maybe you still understand it. Could also be helpful for other people. I have added comments.</p>\n<p>Adapted from <a href=\"https://github.com/catalyst-team/catalyst/blob/master/catalyst/metrics/functional/_accuracy.py\" target=\"_blank\">https://github.com/catalyst-team/catalyst/blob/master/catalyst/metrics/functional/_accuracy.py</a><br>\nTherefore the licence notice. I added the masked accuracy functionality:<br>\nCopyright 2018 Sergey Kolesnikov</p>\n<p>Licensed under the Apache License, Version 2.0 (the \"License\");<br>\n   you may not use this file except in compliance with the License.<br>\n   You may obtain a copy of the License at</p>\n<pre><code>   http://www.apache.org/licenses/LICENSE-2.0\n</code></pre>\n<p>Unless required by applicable law or agreed to in writing, software<br>\n   distributed under the License is distributed on an \"AS IS\" BASIS,<br>\n   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.<br>\n   See the License for the specific language governing permissions and<br>\n   limitations under the License.</p>\n<p>`mask_idx = int(-sys.maxsize) if mask_idx is None else mask_idx   # Your padding index id</p>\n<pre><code>    max_k = max(topk)\n    # Here you count the number of unpadded token positions:\n    batch_size = (targets != mask_idx).to(dtype=torch.uint8).sum(dim=0).item()\n\n    if len(outputs.shape) == 1 or outputs.shape[1] == 1:\n        # binary accuracy\n        pred = outputs.t()\n    else:\n        # multiclass accuracy\n        _, pred = outputs.topk(max_k, 1, True, True)  # noqa: WPS425\n        pred = pred.t()\n    # Here you compute which token predictions where correct and which not:\n    correct = pred.eq(targets.long().view(1, -1).expand_as(pred))\n    # Calculate mask, which shows all unpadded positions (as True):\n    mask = (targets.long() != mask_idx)\n    mask = mask.view(1, -1)\n\n    output = []\n    for k in topk:\n        # Here you only return those predictions which correspond to a unpadded position: \n        correct_k = torch.masked_select(correct[:k], mask)\n        correct_k = correct_k.float().sum(0, keepdim=True)\n        output.append(correct_k.mul_(1.0 / batch_size))`   # Normalization by unpadded position count\n       # output contains the accuracies for our batch &amp; top-k \n       # Note, that you shouldn't place the function at a position where a gradient might be calculated\n       # OR use with 'torch.no_grad():' and outputs.detach() to be safe.`\n</code></pre>\n<p>(Maybe one day I will get the code blocks to work properly)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1277486,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-18T19:31:54.363000",
          "content": "<p>This makes sense. I was leaning towards a solution like this.</p>\n<hr>\n<p>However, it isn't quite accurate. Let me give an example for clarity.</p>\n<pre><code># '1' is the &lt;START&gt; token, '2' is the &lt;END&gt; token and '0' is the &lt;PAD&gt; token\nground_truth = [1, 15, 11, 16, 26, 88, 2, 0, 0, 0]\nprediction   = [1, 15, 11, 16, 26, 88, 44, 50, 2, 15]\n\n# Mask Function as Per Your Logic\nmask = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0]\nmasked_prediction = [1, 15, 11, 16, 26, 88, 44, 0, 0, 0]\nmask_accuracy = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0]\nave_accuracy = sum(mask_accuracy)/sum(mask) # 6/7=0.857\n\n# Real accuracy should compute the accuracy of all values up to the\n# &lt;END&gt; token in the prediction string (not the gt string)\nreal_mask = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0]\nreal_masked_prediction = [1, 15, 11, 16, 26, 88, 44, 50, 2, 0]\nreal_mask_accuracy = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0]\nreal_ave_accuracy = sum(real_mask_accuracy)/sum(real_mask) # 6/9=0.667\n</code></pre>\n<hr>\n<p><strong>As you can see, if we only mask up to the ground truth <em>END</em> token we won't be calculating the real accuracy when the predicted sequence is longer than the ground truth sequence</strong></p>\n<hr>\n<p>All this being said, as this is relatively easy to implement and will give a good feel for accuracy I will most likely go with something like this anyway.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277517,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-18T20:53:51.200000",
          "content": "<p>It wasn't meant to work for comparison with sequences of different lengths. Just top-k classification accuracy like in e.g. image classification. For different lengths I would use Levenshtein distance, makes way more sense.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277519,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-18T20:59:04.623000",
          "content": "<p>Levenshtein distance faces the same problems as the accuracy calculation above. I appreciate your help though! </p>\n<p>I think something like what you suggested is the best I can do right now.</p>\n<hr>\n<p><strong><em>[EDIT] - Did you downvote me because I tried to explain what I was looking for? Seems a bit harsh…</em></strong></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277996,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T13:18:53.170000",
          "content": "<p>Well you can  [see comment below] as described for the accuracy. Although I am very confused why you have potentially two vectors of separate lengths. Do you generate all sequences from scratch during training, instead of parallelism? That would be really slow.<br>\nI mean during inference you (should) convert the outputs back to strings and then use Levenshtein distance. And your tokenizer obviously ignores the padding token. So there is no problem.</p>\n<p><code>Did you downvote me because I tried to explain what I was looking for? Seems a bit harsh…</code><br>\nNo. Was a mistake. I fixed that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278094,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T14:52:07.787000",
          "content": "<p>If you want to calculate Levenshtein distance without transforming to strings first (keep in mind that some tokens may make up 2 characters or more) you could to as follows: <br>\nTensorflow already as an edit distance function, so you don't need to implement that part.<br>\nTherefore you replace all tokens in your prediction X after the end/eos token with the padding token id. If your ground truth Y is shorter than X, you pad it with the padding id, until it has the same length. Then you use the edit distance function.<br>\n<br>\nYou still have to correct it by length difference, -_-</p>\n<p>Example:<br>\nX = predictions; Y = ground truth; 0 = padding id; 1 = end/eos id</p>\n<ol>\n<li>Our vectors:<br>\nX = [ [2, 5, <strong>1</strong>,  7, 3], [2, 4, 3, 9, <strong>1</strong>] ]<br>\nY = [ [2, 5, 6, <strong>1</strong>], [2, 4, 3, <strong>1</strong>]]</li>\n<li>Replace token ids after eos with padding id:<br>\nX = [ [2, 5, 1,  <strong>0</strong>, <strong>0</strong>], [2, 4, 3, 9, <strong>1</strong>] ]</li>\n<li>Pad X &amp; Y to same length:<br>\nY = [ [2, 5, 6,  1, <strong>0</strong>], [2, 4, 3, 1, <strong>0</strong>]]</li>\n<li>Compute distance<br>\nX1/Y1 = 2<br>\nX2/Y2 = 2</li>\n<li>Correct by absolute length difference:<br>\nX1/Y1 = 2 -1 = 1<br>\nX2/Y2 = 2 -1 = 1</li>\n</ol>\n<p>How to replace ids after eos id? Match for eos id, and take index of first occurence in vector, then replace all positions afterward with padding id: X[0, (eos_id +1):] = pad_id. There might be better, inbuild ways to do that.</p>\n<p>I think due to the necessary length difference correction you can just compute edit distance directly, then find the first eos positions, compute true length of X, and then correct the distance calculation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1275410,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2021-04-16T09:26:53.960000",
      "content": "<p>Your model should learn to always predict &lt;pad&gt; after an &lt;end&gt; token fast, this should therefore not cause any problems. I am using the same model architecture as you (CNN encoder and LSTM decoder with attention) and in my experience you should see you model consistently predict &lt;pad&gt; tokens after an &lt;end&gt; token in a few hundred training steps.</p>\n<p>Conceptually this also makes sense, predicting the correct InChI is a complex task, learning to always say &lt;pad&gt; after you see an &lt;end&gt; token is quite simple ;)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1275581,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-04-16T13:13:00.147000",
          "content": "<p><strong>My current loss function masks out the padding token</strong>. This means my model never learns to predict the PAD token after the END token. See below:</p>\n<pre><code>loss_object = tf.keras.losses.SparseCategoricalCrossentropy(\n            from_logits=True, reduction=tf.keras.losses.Reduction.NONE\n        )\n\ndef loss_fn(real, pred):\n    mask = tf.math.not_equal(real, 0) # 0 is my pad token\n    loss_ = loss_object(real, pred)\n    loss_ *= tf.cast(mask, dtype=loss_.dtype)\n    loss_ = tf.nn.compute_average_loss(\n        loss_, global_batch_size=GLOBAL_BATCH_SIZE\n    )\n\n    return loss_\n</code></pre>\n<p>I initially used a loss function similar to yours. However, I'm predicting the full-lengths (max length is ~280 using my tokenization method) of the inchi… not a truncated version (as you are). The average inchi length (using my tokenization method) is around 90 tokens. This means that on average there are 190 PAD tokens in every INCHI string. Because of this massive class imbalance, the model will just start predicting PAD tokens all the time. I guess it would eventually break out of this local minima, however, I thought it made sense to just mask them out so the gradient is smoother.</p>\n<p>Hope this makes sense.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1278151,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-19T15:40:47.640000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1274642": "Hi there. I had a dumb question. I'm using a simple CNN->LSTM pipeline w/ Attention.\n\n---\n\nIn my training/validation loop I'm performing the decoding character by character. After each predicted character I update my accuracy.\n\nThe problem is that I've **defined a loss function that disregards padding**. This means that my model will correctly predict the string ONLY UP UNTIL IT PREDICTS THE END TOKEN. At which point it will just predict the remaining tokens as gibberish. This is completely fine from a prediction standpoint as I would remove everything after the END token anyway.\n\nHowever, when I calculate accuracy, all the gibberish lowers my score (same for LD calculation) as it scores negatively when comparing my gibberish to the PAD token.\n\ni.e. Let's assume I am doing evaluation along the string all at once... even though during training I do it one character at a time.\n\n```python\n\n# '2' is the <END> token and '0' is the <PAD> token\nground_truth = [15, 11, 16, 26, 88, 2, 0, 0, 0, 0]\nprediction   = [15, 11, 16, 26, 89, 2, 44, 50, 2, 15]\n\n# Oversimplified\nraw_accuracy = [1.0, 1.0, 1.0, 1.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0]\n\n# This includes the <END> token\nmean_accuracy   = 0.50\nmean_relevant_accuracy = 83.33 # only up to and including <END> token\n```\n\nIs there some way to calculate the accuracy using masking? I would need to probably wrap the existing metric [(**`SparseCategoricalAccuracy`**)](https://www.tensorflow.org/api_docs/python/tf/keras/metrics/SparseCategoricalAccuracy).\n\nBut the same problem arises w.r.t. calculating Levenshtein distance. My current workaround is to use the ground truth padding tokens as a mask over irrelevant data... this works but if my model predicts sequences that are longer than the GT this error will be masked.\n\n---\n\nDoes anyone have any suggestions? \n\n---\n\n*I'll be sharing a notebook in the next couple of days (once my TPU quota resets). It will show my training with the broken metrics. Hopefully, we can fix it together!*",
    "1278177": "Is this closer to what you are looking for?\n```\npreds = tensor([[5, 6, 2, 2, 3, 7, 4, 5],\n                [6, 7, 7, 5, 7, 5, 2, 6],\n                [6, 5, 6, 3, 4, 3, 6, 6]])\nidxs = (preds == 2).max(dim=1)[1]\n# idxs -> tensor([2, 6, 0])\nidxs[idxs==0] = preds.size(1)\n# idxs -> tensor([2, 6, 8])\npred_mask = (torch.arange(preds.size(1)).unsqueeze(0) <= idxs.unsqueeze(1)) * 1.\n# pred_mask -> tensor([[1., 1., 1., 0., 0., 0., 0., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 0.],\n                       [1., 1., 1., 1., 1., 1., 1., 1.]])\nacc_mask = ((ground_truth!=0) + pred_mask).clamp(max=1.)\n```",
    "1275988": "You can just pad your sequences as given and then count the number of unpadded token-id positions. \nThen you calculate how many of your predictions match the ground-truth. This matching count you divide by the unpadded id count. This will give you the true accuracy. \nYou should mask the padding ids in your loss function, otherwise your model will spend some of its predictive power to learn useless patterns. Try the difference in accuracy after you have implemented the masked accuracy function.\nI have here one for Pytorch, maybe you still understand it. Could also be helpful for other people. I have added comments.\n\nAdapted from https://github.com/catalyst-team/catalyst/blob/master/catalyst/metrics/functional/_accuracy.py\nTherefore the licence notice. I added the masked accuracy functionality:\nCopyright 2018 Sergey Kolesnikov\n\n   Licensed under the Apache License, Version 2.0 (the \"License\");\n   you may not use this file except in compliance with the License.\n   You may obtain a copy of the License at\n\n       http://www.apache.org/licenses/LICENSE-2.0\n\n   Unless required by applicable law or agreed to in writing, software\n   distributed under the License is distributed on an \"AS IS\" BASIS,\n   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n   See the License for the specific language governing permissions and\n   limitations under the License.\n\n`mask_idx = int(-sys.maxsize) if mask_idx is None else mask_idx   # Your padding index id\n\n        max_k = max(topk)\n        # Here you count the number of unpadded token positions:\n        batch_size = (targets != mask_idx).to(dtype=torch.uint8).sum(dim=0).item()\n\n        if len(outputs.shape) == 1 or outputs.shape[1] == 1:\n            # binary accuracy\n            pred = outputs.t()\n        else:\n            # multiclass accuracy\n            _, pred = outputs.topk(max_k, 1, True, True)  # noqa: WPS425\n            pred = pred.t()\n        # Here you compute which token predictions where correct and which not:\n        correct = pred.eq(targets.long().view(1, -1).expand_as(pred))\n        # Calculate mask, which shows all unpadded positions (as True):\n        mask = (targets.long() != mask_idx)\n        mask = mask.view(1, -1)\n\n        output = []\n        for k in topk:\n            # Here you only return those predictions which correspond to a unpadded position: \n            correct_k = torch.masked_select(correct[:k], mask)\n            correct_k = correct_k.float().sum(0, keepdim=True)\n            output.append(correct_k.mul_(1.0 / batch_size))`   # Normalization by unpadded position count\n           # output contains the accuracies for our batch & top-k \n           # Note, that you shouldn't place the function at a position where a gradient might be calculated\n           # OR use with 'torch.no_grad():' and outputs.detach() to be safe.`\n\n(Maybe one day I will get the code blocks to work properly)",
    "1275410": "Your model should learn to always predict <pad\\> after an <end\\> token fast, this should therefore not cause any problems. I am using the same model architecture as you (CNN encoder and LSTM decoder with attention) and in my experience you should see you model consistently predict <pad\\> tokens after an <end\\> token in a few hundred training steps.\n\nConceptually this also makes sense, predicting the correct InChI is a complex task, learning to always say <pad\\> after you see an <end\\> token is quite simple ;)",
    "1278151": ""
  }
}