{
  "id": 242582,
  "title": "Tokenizer face palm",
  "url": "/competitions/bms-molecular-translation/discussion/242582",
  "author_name": "Alexander Soare",
  "post_date": "2021-05-29T18:25:18.715000",
  "votes": 12,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Maybe this will save 0.5% of you (so 4 of you) 30 minutes and lots of head scratching.</p>\n<p>If you ever run into a weird issue where your tokenizer can't read a token in one of your predictions (which feels crazy because the tokenizer produced the predictions in the first place) it's probably because you fed in normalized predictions. <code>/p</code> and <code>/q</code> are two tokens that appear in normalized predictions but never in the training dataset.</p>",
  "messages": [
    {
      "id": 1327915,
      "postDate": "2021-05-29T18:25:18.717Z",
      "content": "<p>Maybe this will save 0.5% of you (so 4 of you) 30 minutes and lots of head scratching.</p>\n<p>If you ever run into a weird issue where your tokenizer can't read a token in one of your predictions (which feels crazy because the tokenizer produced the predictions in the first place) it's probably because you fed in normalized predictions. <code>/p</code> and <code>/q</code> are two tokens that appear in normalized predictions but never in the training dataset.</p>",
      "rawMarkdown": "Maybe this will save 0.5% of you (so 4 of you) 30 minutes and lots of head scratching.\n\nIf you ever run into a weird issue where your tokenizer can't read a token in one of your predictions (which feels crazy because the tokenizer produced the predictions in the first place) it's probably because you fed in normalized predictions. `/p` and `/q` are two tokens that appear in normalized predictions but never in the training dataset.",
      "votes": 12
    },
    {
      "id": 1329660,
      "postDate": "2021-05-31T09:32:40.497Z",
      "content": "<p>Wanted to upvote, but a fifth vote would question your perfect prediction 😉</p>",
      "rawMarkdown": "Wanted to upvote, but a fifth vote would question your perfect prediction 😉",
      "votes": 1,
      "replies": [
        {
          "id": 1330082,
          "postDate": "2021-05-31T15:26:35.493Z",
          "content": "<p>I liked it so now you can, too. :D</p>",
          "rawMarkdown": "I liked it so now you can, too. :D",
          "votes": 1
        },
        {
          "id": 1330222,
          "postDate": "2021-05-31T16:35:54.973Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1330080,
      "postDate": "2021-05-31T15:25:33.687Z",
      "content": "<p>There can also be question marks.</p>",
      "rawMarkdown": "There can also be question marks.",
      "votes": 2,
      "replies": [
        {
          "id": 1330254,
          "postDate": "2021-05-31T16:48:32.620Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> why aren't you submitting any more?</p>",
          "rawMarkdown": "@nofreewill why aren't you submitting any more?"
        },
        {
          "id": 1330266,
          "postDate": "2021-05-31T16:56:33.827Z",
          "content": "<p>I'm stuck with a solution, that tells me I'll have ~0.73 LB.<br>\nNow I think I have a new idea that a lot of teams are doing already. I don't know how I didn't think of it, as there are even mentions of it (kind of it) in discussions. I hope I'll be able to go lower, but now I have to work on it pretty hard, as there is only that much time left. I even bought some, I mean a lot of coffee for the evening… :D</p>",
          "rawMarkdown": "I'm stuck with a solution, that tells me I'll have ~0.73 LB.\nNow I think I have a new idea that a lot of teams are doing already. I don't know how I didn't think of it, as there are even mentions of it (kind of it) in discussions. I hope I'll be able to go lower, but now I have to work on it pretty hard, as there is only that much time left. I even bought some, I mean a lot of coffee for the evening... :D"
        },
        {
          "id": 1330276,
          "postDate": "2021-05-31T17:06:59.790Z",
          "content": "<p>I see, well best of luck. I also had a realisation 2 days ago and had to do a back of the envelope calculation to see if running my GPU 24/7 would make it in time. No coffee for me though, maybe if the GPU had feelings I'd get it coffee.</p>",
          "rawMarkdown": "I see, well best of luck. I also had a realisation 2 days ago and had to do a back of the envelope calculation to see if running my GPU 24/7 would make it in time. No coffee for me though, maybe if the GPU had feelings I'd get it coffee.",
          "votes": 1
        },
        {
          "id": 1330295,
          "postDate": "2021-05-31T17:17:48.373Z",
          "content": "<p>If the problem lies with doing the test-set inference in time, you could just do the percentage of the test-set with your new method you can do in time. And then submit the next-best predictions for the rest of the test-set. Better then nothing at all. <br>\n<em>(Says the guy without any submission at all)</em></p>",
          "rawMarkdown": "If the problem lies with doing the test-set inference in time, you could just do the percentage of the test-set with your new method you can do in time. And then submit the next-best predictions for the rest of the test-set. Better then nothing at all. \n*(Says the guy without any submission at all)*",
          "votes": 2
        },
        {
          "id": 1330319,
          "postDate": "2021-05-31T17:28:25.517Z",
          "content": "<p>Yeah makes sense, that's part of the problem indeed</p>",
          "rawMarkdown": "Yeah makes sense, that's part of the problem indeed"
        },
        {
          "id": 1330381,
          "postDate": "2021-05-31T18:28:37.650Z",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> I agree.<br>\nBefore doing my ensemble + beam search, I predicted all test samples with each of my models and I only make the heavy predictions on the ones that have any disagreement between those.<br>\nI predict only ~390k images this way.</p>",
          "rawMarkdown": "@cepheidq I agree.\nBefore doing my ensemble + beam search, I predicted all test samples with each of my models and I only make the heavy predictions on the ones that have any disagreement between those.\nI predict only ~390k images this way.",
          "votes": 3
        },
        {
          "id": 1330425,
          "postDate": "2021-05-31T19:19:41.007Z",
          "content": "<p>May I ask, how many percent of your validation predictions have zero Levenshtein-distance?<br>\nFor me it's 72.6% </p>\n<p>(But I think I have a bug somewhere, because my valid-loss is 0.0096 and I have a validation LD of 5.03 which is way worse then it should be, given some published results here. I have also a top-1 accuracy of 99.54% )<br>\n<em>-&gt; Edit due to wrong accuracy</em></p>",
          "rawMarkdown": "May I ask, how many percent of your validation predictions have zero Levenshtein-distance?\nFor me it's 72.6% \n\n(But I think I have a bug somewhere, because my valid-loss is 0.0096 and I have a validation LD of 5.03 which is way worse then it should be, given some published results here. I have also a top-1 accuracy of 99.54% )\n*-> Edit due to wrong accuracy*"
        },
        {
          "id": 1330427,
          "postDate": "2021-05-31T19:21:22.733Z",
          "content": "<p>Seems about right to me. I had a valid loss of something around 0.002 with 85% of predictions with 0 LD, and 2 as the average <a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> <br>\nDon't read too much into that loss though. My loss fn was a little different for that one.</p>",
          "rawMarkdown": "Seems about right to me. I had a valid loss of something around 0.002 with 85% of predictions with 0 LD, and 2 as the average @cepheidq \nDon't read too much into that loss though. My loss fn was a little different for that one.",
          "votes": 1
        },
        {
          "id": 1330451,
          "postDate": "2021-05-31T19:55:43.610Z",
          "content": "<p>Hmm, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> has a LD of 1.4 with loss 0.010. I just checked their <a href=\"https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing\" target=\"_blank\">code</a>, they use cross-entropy loss.<br>\nBut maybe that was just with teacher-forcing</p>",
          "rawMarkdown": "Hmm, @hengck23 has a LD of 1.4 with loss 0.010. I just checked their [code](https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing), they use cross-entropy loss.\nBut maybe that was just with teacher-forcing"
        },
        {
          "id": 1330476,
          "postDate": "2021-05-31T20:43:40.080Z",
          "content": "<p>UPDATE (For the 4 people who don't know that yet):<br>\nBe the comparison with hengck23 results as they may - using pretrained weights makes a HUGE difference so far. I didn't use them beforehand, because the commercial usage is questionable. But I have a feeling (almost) nobody cares about that fact anyway here. So yeah, I am training with them and see If I can get a better performance in the time left.</p>\n<p>A bit to late for fancy stuff though - I even think I had the same idea which could improve results by quite a margin, but we will see after the competition. <br>\n(Yes, I know that pretrained models are better, but I wanted be responsible and not do things which shouldn't be allowed)</p>",
          "rawMarkdown": "UPDATE (For the 4 people who don't know that yet):\nBe the comparison with hengck23 results as they may - using pretrained weights makes a HUGE difference so far. I didn't use them beforehand, because the commercial usage is questionable. But I have a feeling (almost) nobody cares about that fact anyway here. So yeah, I am training with them and see If I can get a better performance in the time left.\n\nA bit to late for fancy stuff though - I even think I had the same idea which could improve results by quite a margin, but we will see after the competition. \n(Yes, I know that pretrained models are better, but I wanted be responsible and not do things which shouldn't be allowed)",
          "votes": 2
        },
        {
          "id": 1331155,
          "postDate": "2021-06-01T09:59:07.213Z",
          "content": "<p>UPDATE 2: <br>\nI found at least one of my issues:<br>\nUse more attention blocks if you use transformer architecture. Don't feel clever like me and use the <a href=\"http://export.arxiv.org/abs/2009.04534\" target=\"_blank\">Pay Attention when Required Transformer </a> (PAR). It isn't so viable for context depended generation as it has not much encoder-decoder attention. So it works still well with teacher forcing as it sees previous correct tokens, but not during inference. That is just a guess based on my current run.<br>\nAlthough to late to get good performance now. <br>\n(~5.4 LD with 99.09% top-1 valid accuracy and 0.02 loss vs. ~5 LD with 99.54% top-1 &amp; 0.0096 loss)</p>\n<p><strong>Edit</strong>: My hypothesis so far: The decoder performance seems to be a function of the number of attention blocks (considering everything else is well designed). </p>\n<p>92.1% top-1 accuracy leads to LD 4.64+-0.1 (one epoch more than last post). I have aborted training (after 13h) and started again with now 7 attention blocks instead of 5. This is sufficient to falsify my hypothesis. And it might prevent me from submitting any decent result.</p>",
          "rawMarkdown": "UPDATE 2: \nI found at least one of my issues:\nUse more attention blocks if you use transformer architecture. Don't feel clever like me and use the [Pay Attention when Required Transformer ](http://export.arxiv.org/abs/2009.04534) (PAR). It isn't so viable for context depended generation as it has not much encoder-decoder attention. So it works still well with teacher forcing as it sees previous correct tokens, but not during inference. That is just a guess based on my current run.\nAlthough to late to get good performance now. \n(~5.4 LD with 99.09% top-1 valid accuracy and 0.02 loss vs. ~5 LD with 99.54% top-1 & 0.0096 loss)\n\n**Edit**: My hypothesis so far: The decoder performance seems to be a function of the number of attention blocks (considering everything else is well designed). \n\n92.1% top-1 accuracy leads to LD 4.64+-0.1 (one epoch more than last post). I have aborted training (after 13h) and started again with now 7 attention blocks instead of 5. This is sufficient to falsify my hypothesis. And it might prevent me from submitting any decent result."
        },
        {
          "id": 1331183,
          "postDate": "2021-06-01T10:20:32.880Z",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> I'm newish to transformers so I'm not sure how teaching forcing is a choice here. With a single transformer decoder for the whole string you default to 100% teacher forcing during training isn't that right?  There's no way to get around that unless you explicitly break up the transformer into pieces (sequence-wise)</p>",
          "rawMarkdown": "@cepheidq I'm newish to transformers so I'm not sure how teaching forcing is a choice here. With a single transformer decoder for the whole string you default to 100% teacher forcing during training isn't that right?  There's no way to get around that unless you explicitly break up the transformer into pieces (sequence-wise)"
        },
        {
          "id": 1331226,
          "postDate": "2021-06-01T10:47:05.883Z",
          "content": "<p>Yes, you have a parallel input and therefore teacher-forcing by default.<br>\nThe advantage is the speed and that a transformer can always attend to all unmasked tokens - contrary to a RNN which gets past/future tokens through a temporal chain.</p>\n<p>There isn't really a drawback considering the vast advantages. It is just that in case of the PAR-Transformer and a translation task we need more focus on the context (encoded image). But a PAR-transformer has by default only half of the attention blocks. <br>\nSo I think it just learns to much to rely on past tokens which it can do as well as a normal transformer with 2x the attention block. But during inference that doesn't perform swell any more.<br>\nSo no worries for normal transformer architecture (As they are SOTA for machine-translation &amp; image captioning).</p>\n<p>Regarding sequence-wise: You have that with very long sequences where it is called a memory. There you cut off &amp; store past hidden embeddings of the respective attention layer input. Then you use them still as context for your attention, but you don't calculate any attention for the memory tokens itself.</p>\n<p><strong>TLDR</strong>: Teacher-forcing no problem in itself. Just PAR-transformer sub-optimal for encoder-decoder tasks. Sequence-wise just in case of memory usage for very long sequences due to vanilla O(n²d) attention.</p>\n<p>Should you have more questions that are more in-depth, ask away :) But I have to sleep first before answering anything more complex.</p>",
          "rawMarkdown": "Yes, you have a parallel input and therefore teacher-forcing by default.\nThe advantage is the speed and that a transformer can always attend to all unmasked tokens - contrary to a RNN which gets past/future tokens through a temporal chain.\n\nThere isn't really a drawback considering the vast advantages. It is just that in case of the PAR-Transformer and a translation task we need more focus on the context (encoded image). But a PAR-transformer has by default only half of the attention blocks. \nSo I think it just learns to much to rely on past tokens which it can do as well as a normal transformer with 2x the attention block. But during inference that doesn't perform swell any more.\nSo no worries for normal transformer architecture (As they are SOTA for machine-translation & image captioning).\n\nRegarding sequence-wise: You have that with very long sequences where it is called a memory. There you cut off & store past hidden embeddings of the respective attention layer input. Then you use them still as context for your attention, but you don't calculate any attention for the memory tokens itself.\n\n**TLDR**: Teacher-forcing no problem in itself. Just PAR-transformer sub-optimal for encoder-decoder tasks. Sequence-wise just in case of memory usage for very long sequences due to vanilla O(n²d) attention.\n\nShould you have more questions that are more in-depth, ask away :) But I have to sleep first before answering anything more complex.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1329660,
      "author_name": "Alexander Bader",
      "author_url": "",
      "post_date": "2021-05-31T09:32:40.497000",
      "content": "<p>Wanted to upvote, but a fifth vote would question your perfect prediction 😉</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1330082,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-31T15:26:35.493000",
          "content": "<p>I liked it so now you can, too. :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1330222,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-31T16:35:54.973000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1330080,
      "author_name": "nofreewill42",
      "author_url": "",
      "post_date": "2021-05-31T15:25:33.687000",
      "content": "<p>There can also be question marks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1330254,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-31T16:48:32.620000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> why aren't you submitting any more?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1330266,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-31T16:56:33.827000",
          "content": "<p>I'm stuck with a solution, that tells me I'll have ~0.73 LB.<br>\nNow I think I have a new idea that a lot of teams are doing already. I don't know how I didn't think of it, as there are even mentions of it (kind of it) in discussions. I hope I'll be able to go lower, but now I have to work on it pretty hard, as there is only that much time left. I even bought some, I mean a lot of coffee for the evening… :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1330276,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-31T17:06:59.790000",
          "content": "<p>I see, well best of luck. I also had a realisation 2 days ago and had to do a back of the envelope calculation to see if running my GPU 24/7 would make it in time. No coffee for me though, maybe if the GPU had feelings I'd get it coffee.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1330295,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-05-31T17:17:48.373000",
          "content": "<p>If the problem lies with doing the test-set inference in time, you could just do the percentage of the test-set with your new method you can do in time. And then submit the next-best predictions for the rest of the test-set. Better then nothing at all. <br>\n<em>(Says the guy without any submission at all)</em></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1330319,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-31T17:28:25.517000",
          "content": "<p>Yeah makes sense, that's part of the problem indeed</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1330381,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-31T18:28:37.650000",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> I agree.<br>\nBefore doing my ensemble + beam search, I predicted all test samples with each of my models and I only make the heavy predictions on the ones that have any disagreement between those.<br>\nI predict only ~390k images this way.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1330425,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-05-31T19:19:41.007000",
          "content": "<p>May I ask, how many percent of your validation predictions have zero Levenshtein-distance?<br>\nFor me it's 72.6% </p>\n<p>(But I think I have a bug somewhere, because my valid-loss is 0.0096 and I have a validation LD of 5.03 which is way worse then it should be, given some published results here. I have also a top-1 accuracy of 99.54% )<br>\n<em>-&gt; Edit due to wrong accuracy</em></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1330427,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-31T19:21:22.733000",
          "content": "<p>Seems about right to me. I had a valid loss of something around 0.002 with 85% of predictions with 0 LD, and 2 as the average <a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> <br>\nDon't read too much into that loss though. My loss fn was a little different for that one.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1330451,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-05-31T19:55:43.610000",
          "content": "<p>Hmm, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> has a LD of 1.4 with loss 0.010. I just checked their <a href=\"https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing\" target=\"_blank\">code</a>, they use cross-entropy loss.<br>\nBut maybe that was just with teacher-forcing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1330476,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-05-31T20:43:40.080000",
          "content": "<p>UPDATE (For the 4 people who don't know that yet):<br>\nBe the comparison with hengck23 results as they may - using pretrained weights makes a HUGE difference so far. I didn't use them beforehand, because the commercial usage is questionable. But I have a feeling (almost) nobody cares about that fact anyway here. So yeah, I am training with them and see If I can get a better performance in the time left.</p>\n<p>A bit to late for fancy stuff though - I even think I had the same idea which could improve results by quite a margin, but we will see after the competition. <br>\n(Yes, I know that pretrained models are better, but I wanted be responsible and not do things which shouldn't be allowed)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1331155,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-06-01T09:59:07.213000",
          "content": "<p>UPDATE 2: <br>\nI found at least one of my issues:<br>\nUse more attention blocks if you use transformer architecture. Don't feel clever like me and use the <a href=\"http://export.arxiv.org/abs/2009.04534\" target=\"_blank\">Pay Attention when Required Transformer </a> (PAR). It isn't so viable for context depended generation as it has not much encoder-decoder attention. So it works still well with teacher forcing as it sees previous correct tokens, but not during inference. That is just a guess based on my current run.<br>\nAlthough to late to get good performance now. <br>\n(~5.4 LD with 99.09% top-1 valid accuracy and 0.02 loss vs. ~5 LD with 99.54% top-1 &amp; 0.0096 loss)</p>\n<p><strong>Edit</strong>: My hypothesis so far: The decoder performance seems to be a function of the number of attention blocks (considering everything else is well designed). </p>\n<p>92.1% top-1 accuracy leads to LD 4.64+-0.1 (one epoch more than last post). I have aborted training (after 13h) and started again with now 7 attention blocks instead of 5. This is sufficient to falsify my hypothesis. And it might prevent me from submitting any decent result.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331183,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-06-01T10:20:32.880000",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> I'm newish to transformers so I'm not sure how teaching forcing is a choice here. With a single transformer decoder for the whole string you default to 100% teacher forcing during training isn't that right?  There's no way to get around that unless you explicitly break up the transformer into pieces (sequence-wise)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331226,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-06-01T10:47:05.883000",
          "content": "<p>Yes, you have a parallel input and therefore teacher-forcing by default.<br>\nThe advantage is the speed and that a transformer can always attend to all unmasked tokens - contrary to a RNN which gets past/future tokens through a temporal chain.</p>\n<p>There isn't really a drawback considering the vast advantages. It is just that in case of the PAR-Transformer and a translation task we need more focus on the context (encoded image). But a PAR-transformer has by default only half of the attention blocks. <br>\nSo I think it just learns to much to rely on past tokens which it can do as well as a normal transformer with 2x the attention block. But during inference that doesn't perform swell any more.<br>\nSo no worries for normal transformer architecture (As they are SOTA for machine-translation &amp; image captioning).</p>\n<p>Regarding sequence-wise: You have that with very long sequences where it is called a memory. There you cut off &amp; store past hidden embeddings of the respective attention layer input. Then you use them still as context for your attention, but you don't calculate any attention for the memory tokens itself.</p>\n<p><strong>TLDR</strong>: Teacher-forcing no problem in itself. Just PAR-transformer sub-optimal for encoder-decoder tasks. Sequence-wise just in case of memory usage for very long sequences due to vanilla O(n²d) attention.</p>\n<p>Should you have more questions that are more in-depth, ask away :) But I have to sleep first before answering anything more complex.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1327915": "Maybe this will save 0.5% of you (so 4 of you) 30 minutes and lots of head scratching.\n\nIf you ever run into a weird issue where your tokenizer can't read a token in one of your predictions (which feels crazy because the tokenizer produced the predictions in the first place) it's probably because you fed in normalized predictions. `/p` and `/q` are two tokens that appear in normalized predictions but never in the training dataset.",
    "1329660": "Wanted to upvote, but a fifth vote would question your perfect prediction 😉",
    "1330080": "There can also be question marks."
  }
}