{
  "id": 417433,
  "title": "Why was the test fragment rotated?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417433",
  "author_name": "",
  "post_date": "2023-06-15T17:17:23.851860300Z",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I think by now it's pretty obvious that the hidden test fragment was rotated by 90 degrees. What I'm wondering is why this decision was made?</p>\n<p>I understand that when scanning a scroll you don't necessarily know the orientation of the text, or even if there is a consistent orientation. However, this feels like an additional challenge that could have been left outside of the scope of this Kaggle competition. If you have a good model to detect text for a known orientation, then you can try different rotations in inference to get the best results. Which is pretty much what happened here.</p>\n<p>But why didn't we focus only on finding the best models for a consistent orientation? It's certainly a challenge by itself, and might have produced cleaner results &amp; progress without the additional distraction of rotation. What am I missing here?</p>",
  "messages": [
    {
      "id": "2304101",
      "postDate": "06/15/2023 17:17:23",
      "content": "<p>I think by now it's pretty obvious that the hidden test fragment was rotated by 90 degrees. What I'm wondering is why this decision was made?</p>\n<p>I understand that when scanning a scroll you don't necessarily know the orientation of the text, or even if there is a consistent orientation. However, this feels like an additional challenge that could have been left outside of the scope of this Kaggle competition. If you have a good model to detect text for a known orientation, then you can try different rotations in inference to get the best results. Which is pretty much what happened here.</p>\n<p>But why didn't we focus only on finding the best models for a consistent orientation? It's certainly a challenge by itself, and might have produced cleaner results &amp; progress without the additional distraction of rotation. What am I missing here?</p>",
      "rawMarkdown": "I think by now it's pretty obvious that the hidden test fragment was rotated by 90 degrees. What I'm wondering is why this decision was made?\n\nI understand that when scanning a scroll you don't necessarily know the orientation of the text, or even if there is a consistent orientation. However, this feels like an additional challenge that could have been left outside of the scope of this Kaggle competition. If you have a good model to detect text for a known orientation, then you can try different rotations in inference to get the best results. Which is pretty much what happened here.\n\nBut why didn't we focus only on finding the best models for a consistent orientation? It's certainly a challenge by itself, and might have produced cleaner results & progress without the additional distraction of rotation. What am I missing here?",
      "votes": null
    },
    {
      "id": "2304264",
      "postDate": "06/15/2023 21:19:23",
      "content": "<p>That's a good point! I spent a lot of time before understanding why the rotation data augmentation was useful.</p>",
      "rawMarkdown": "That's a good point! I spent a lot of time before understanding why the rotation data augmentation was useful.",
      "votes": null
    },
    {
      "id": "2304659",
      "postDate": "06/16/2023 06:36:55",
      "content": "<p>There was some discussion about the test fragment here (you may have seen this already) -<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/412513\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/412513</a></p>\n<p>in it from the competition host - <br>\n\"We double checked the data and did not find any unexpected anomalies. Keep in mind that you should not assume any specific alignment or letter scale of the hidden fragment. Submissions should aim to detect ink and be robust to changes in rotation, translation, warping, and other such transformations, as they will occur in the real world. The ultimate use for the work here is prediction on the full scrolls, which will not be as neat as the fragments.\"</p>\n<p>Think there was more than just orientation difference in the private test fragment and even the public 10% vs the rest of it.  Had one submission in the 0.6ish public and 0 in the private and it had rotation in inference.  Did not select it. </p>\n<p>It will be interesting to see if the work here can make progress on the Grand Prize and Other Prize.     </p>",
      "rawMarkdown": "There was some discussion about the test fragment here (you may have seen this already) -\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/412513\n\nin it from the competition host - \n\"We double checked the data and did not find any unexpected anomalies. Keep in mind that you should not assume any specific alignment or letter scale of the hidden fragment. Submissions should aim to detect ink and be robust to changes in rotation, translation, warping, and other such transformations, as they will occur in the real world. The ultimate use for the work here is prediction on the full scrolls, which will not be as neat as the fragments.\"\n\nThink there was more than just orientation difference in the private test fragment and even the public 10% vs the rest of it.  Had one submission in the 0.6ish public and 0 in the private and it had rotation in inference.  Did not select it. \n\nIt will be interesting to see if the work here can make progress on the Grand Prize and Other Prize.",
      "votes": null
    },
    {
      "id": "2304696",
      "postDate": "06/16/2023 07:08:10",
      "content": "<p>actually i am more interested how did the first kaggler discover that we need rotation.</p>\n<p>data abnormality is common in work (and competition)<br>\nwhen we did not get good results, we tend to improve our model and hyper-parameters.</p>\n<p>i hope this experinece will make me think of \"data\" if i encounter poor performance in the future</p>",
      "rawMarkdown": "actually i am more interested how did the first kaggler discover that we need rotation.\n\ndata abnormality is common in work (and competition)\nwhen we did not get good results, we tend to improve our model and hyper-parameters.\n\ni hope this experinece will make me think of \"data\" if i encounter poor performance in the future",
      "votes": null
    },
    {
      "id": "2304764",
      "postDate": "06/16/2023 08:09:01",
      "content": "<blockquote>\n  <p>actually i am more interested how did the first kaggler discover that we need rotation.</p>\n</blockquote>\n<p>Yes, I was also wondering about that!</p>",
      "rawMarkdown": "> actually i am more interested how did the first kaggler discover that we need rotation.\n\nYes, I was also wondering about that!",
      "votes": null
    },
    {
      "id": "2304944",
      "postDate": "06/16/2023 10:39:13",
      "content": "<p>Seems like it was probably an accidental discovery and found when tta had a disproportionately high performance impact</p>",
      "rawMarkdown": "Seems like it was probably an accidental discovery and found when tta had a disproportionately high performance impact",
      "votes": null
    },
    {
      "id": "2305343",
      "postDate": "06/16/2023 16:01:27",
      "content": "<p>This is a great find, thank you! I think I saw the original post, but then missed the replies. </p>",
      "rawMarkdown": "This is a great find, thank you! I think I saw the original post, but then missed the replies.",
      "votes": null
    },
    {
      "id": "2305348",
      "postDate": "06/16/2023 16:06:47",
      "content": "<p>Good point! I think it comes down to implementing robust QA checks to make sure that your inference/production data stays within the parameters of the training data. E.g. anomaly detection or drift detection. In work settings, you usually have more control over your data; as opposed to a competition setting where we're predicting blindly.</p>\n<p>I also think this was discovered through tta experiments, then narrowed down to rotation as the main impact.</p>",
      "rawMarkdown": "Good point! I think it comes down to implementing robust QA checks to make sure that your inference/production data stays within the parameters of the training data. E.g. anomaly detection or drift detection. In work settings, you usually have more control over your data; as opposed to a competition setting where we're predicting blindly.\n\nI also think this was discovered through tta experiments, then narrowed down to rotation as the main impact.",
      "votes": null
    },
    {
      "id": "2305784",
      "postDate": "06/16/2023 22:55:28",
      "content": "<p>For what it's worth, I'm not sure why rotation was an issue.  Here is one clear letter from the ink mask for fragment 1:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2F69edbb3c57348363e1472df1c566a1d3%2Fletter.png?generation=1686955477919765&amp;alt=media\" alt=\"\"><br>\nThis letter is about 700x700 pixels.  From the solutions I have read so far, many were using crops around 256x256, which would not have been large enough to contain multiple letters or see text orientation.  I guess there is some kind of difference whether you rotate the text or not, but it seems like it would have to be more related to subtleties like letters having more vertical strokes than horizontal ones, as opposed to the rows of letters that are visible at a much more zoomed out scale.  Am I thinking correctly about this?</p>\n<p>(FYI, we did rotations and flips in our training augmentations and 4x rotations in our TTA.)</p>",
      "rawMarkdown": "For what it's worth, I'm not sure why rotation was an issue.  Here is one clear letter from the ink mask for fragment 1:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2F69edbb3c57348363e1472df1c566a1d3%2Fletter.png?generation=1686955477919765&alt=media)\nThis letter is about 700x700 pixels.  From the solutions I have read so far, many were using crops around 256x256, which would not have been large enough to contain multiple letters or see text orientation.  I guess there is some kind of difference whether you rotate the text or not, but it seems like it would have to be more related to subtleties like letters having more vertical strokes than horizontal ones, as opposed to the rows of letters that are visible at a much more zoomed out scale.  Am I thinking correctly about this?\n\n(FYI, we did rotations and flips in our training augmentations and 4x rotations in our TTA.)",
      "votes": null
    },
    {
      "id": "2305885",
      "postDate": "06/17/2023 01:18:46",
      "content": "<p>Our team primarily used image sizes of 384 and 512, but your point still stands. If I were to speculate, there might be something here about certain angles in most letters being easier to recognize in their original orientation than when rotated by 90 degrees. Think about an \"A\" or a \"P\" (or their Greek equivalents). If you cut out small parts of those, then the angles should look different after rotation.</p>\n<p>In addition, the organizer's comment also mentions \"warping\"; so maybe there were some other slight distortions applied that made it harder to match patterns after rotation?</p>",
      "rawMarkdown": "Our team primarily used image sizes of 384 and 512, but your point still stands. If I were to speculate, there might be something here about certain angles in most letters being easier to recognize in their original orientation than when rotated by 90 degrees. Think about an \"A\" or a \"P\" (or their Greek equivalents). If you cut out small parts of those, then the angles should look different after rotation.\n\nIn addition, the organizer's comment also mentions \"warping\"; so maybe there were some other slight distortions applied that made it harder to match patterns after rotation?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2304264,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/15/2023 21:19:23",
      "content": "<p>That's a good point! I spent a lot of time before understanding why the rotation data augmentation was useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2304659,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "06/16/2023 06:36:55",
      "content": "<p>There was some discussion about the test fragment here (you may have seen this already) -<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/412513\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/412513</a></p>\n<p>in it from the competition host - <br>\n\"We double checked the data and did not find any unexpected anomalies. Keep in mind that you should not assume any specific alignment or letter scale of the hidden fragment. Submissions should aim to detect ink and be robust to changes in rotation, translation, warping, and other such transformations, as they will occur in the real world. The ultimate use for the work here is prediction on the full scrolls, which will not be as neat as the fragments.\"</p>\n<p>Think there was more than just orientation difference in the private test fragment and even the public 10% vs the rest of it.  Had one submission in the 0.6ish public and 0 in the private and it had rotation in inference.  Did not select it. </p>\n<p>It will be interesting to see if the work here can make progress on the Grand Prize and Other Prize.     </p>",
      "votes": null,
      "replies": [
        {
          "id": 2305343,
          "author_name": "headsortails",
          "author_url": "",
          "post_date": "06/16/2023 16:01:27",
          "content": "<p>This is a great find, thank you! I think I saw the original post, but then missed the replies. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2305784,
          "author_name": "socated",
          "author_url": "",
          "post_date": "06/16/2023 22:55:28",
          "content": "<p>For what it's worth, I'm not sure why rotation was an issue.  Here is one clear letter from the ink mask for fragment 1:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2F69edbb3c57348363e1472df1c566a1d3%2Fletter.png?generation=1686955477919765&amp;alt=media\" alt=\"\"><br>\nThis letter is about 700x700 pixels.  From the solutions I have read so far, many were using crops around 256x256, which would not have been large enough to contain multiple letters or see text orientation.  I guess there is some kind of difference whether you rotate the text or not, but it seems like it would have to be more related to subtleties like letters having more vertical strokes than horizontal ones, as opposed to the rows of letters that are visible at a much more zoomed out scale.  Am I thinking correctly about this?</p>\n<p>(FYI, we did rotations and flips in our training augmentations and 4x rotations in our TTA.)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2305885,
              "author_name": "headsortails",
              "author_url": "",
              "post_date": "06/17/2023 01:18:46",
              "content": "<p>Our team primarily used image sizes of 384 and 512, but your point still stands. If I were to speculate, there might be something here about certain angles in most letters being easier to recognize in their original orientation than when rotated by 90 degrees. Think about an \"A\" or a \"P\" (or their Greek equivalents). If you cut out small parts of those, then the angles should look different after rotation.</p>\n<p>In addition, the organizer's comment also mentions \"warping\"; so maybe there were some other slight distortions applied that made it harder to match patterns after rotation?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2304696,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/16/2023 07:08:10",
      "content": "<p>actually i am more interested how did the first kaggler discover that we need rotation.</p>\n<p>data abnormality is common in work (and competition)<br>\nwhen we did not get good results, we tend to improve our model and hyper-parameters.</p>\n<p>i hope this experinece will make me think of \"data\" if i encounter poor performance in the future</p>",
      "votes": null,
      "replies": [
        {
          "id": 2304764,
          "author_name": "lucasvw",
          "author_url": "",
          "post_date": "06/16/2023 08:09:01",
          "content": "<blockquote>\n  <p>actually i am more interested how did the first kaggler discover that we need rotation.</p>\n</blockquote>\n<p>Yes, I was also wondering about that!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2304944,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "06/16/2023 10:39:13",
          "content": "<p>Seems like it was probably an accidental discovery and found when tta had a disproportionately high performance impact</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2305348,
          "author_name": "headsortails",
          "author_url": "",
          "post_date": "06/16/2023 16:06:47",
          "content": "<p>Good point! I think it comes down to implementing robust QA checks to make sure that your inference/production data stays within the parameters of the training data. E.g. anomaly detection or drift detection. In work settings, you usually have more control over your data; as opposed to a competition setting where we're predicting blindly.</p>\n<p>I also think this was discovered through tta experiments, then narrowed down to rotation as the main impact.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2304101": "I think by now it's pretty obvious that the hidden test fragment was rotated by 90 degrees. What I'm wondering is why this decision was made?\n\nI understand that when scanning a scroll you don't necessarily know the orientation of the text, or even if there is a consistent orientation. However, this feels like an additional challenge that could have been left outside of the scope of this Kaggle competition. If you have a good model to detect text for a known orientation, then you can try different rotations in inference to get the best results. Which is pretty much what happened here.\n\nBut why didn't we focus only on finding the best models for a consistent orientation? It's certainly a challenge by itself, and might have produced cleaner results & progress without the additional distraction of rotation. What am I missing here?",
    "2304264": "That's a good point! I spent a lot of time before understanding why the rotation data augmentation was useful.",
    "2304659": "There was some discussion about the test fragment here (you may have seen this already) -\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/412513\n\nin it from the competition host - \n\"We double checked the data and did not find any unexpected anomalies. Keep in mind that you should not assume any specific alignment or letter scale of the hidden fragment. Submissions should aim to detect ink and be robust to changes in rotation, translation, warping, and other such transformations, as they will occur in the real world. The ultimate use for the work here is prediction on the full scrolls, which will not be as neat as the fragments.\"\n\nThink there was more than just orientation difference in the private test fragment and even the public 10% vs the rest of it.  Had one submission in the 0.6ish public and 0 in the private and it had rotation in inference.  Did not select it. \n\nIt will be interesting to see if the work here can make progress on the Grand Prize and Other Prize.",
    "2304696": "actually i am more interested how did the first kaggler discover that we need rotation.\n\ndata abnormality is common in work (and competition)\nwhen we did not get good results, we tend to improve our model and hyper-parameters.\n\ni hope this experinece will make me think of \"data\" if i encounter poor performance in the future",
    "2304764": "> actually i am more interested how did the first kaggler discover that we need rotation.\n\nYes, I was also wondering about that!",
    "2304944": "Seems like it was probably an accidental discovery and found when tta had a disproportionately high performance impact",
    "2305343": "This is a great find, thank you! I think I saw the original post, but then missed the replies.",
    "2305348": "Good point! I think it comes down to implementing robust QA checks to make sure that your inference/production data stays within the parameters of the training data. E.g. anomaly detection or drift detection. In work settings, you usually have more control over your data; as opposed to a competition setting where we're predicting blindly.\n\nI also think this was discovered through tta experiments, then narrowed down to rotation as the main impact.",
    "2305784": "For what it's worth, I'm not sure why rotation was an issue.  Here is one clear letter from the ink mask for fragment 1:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2F69edbb3c57348363e1472df1c566a1d3%2Fletter.png?generation=1686955477919765&alt=media)\nThis letter is about 700x700 pixels.  From the solutions I have read so far, many were using crops around 256x256, which would not have been large enough to contain multiple letters or see text orientation.  I guess there is some kind of difference whether you rotate the text or not, but it seems like it would have to be more related to subtleties like letters having more vertical strokes than horizontal ones, as opposed to the rows of letters that are visible at a much more zoomed out scale.  Am I thinking correctly about this?\n\n(FYI, we did rotations and flips in our training augmentations and 4x rotations in our TTA.)",
    "2305885": "Our team primarily used image sizes of 384 and 512, but your point still stands. If I were to speculate, there might be something here about certain angles in most letters being easier to recognize in their original orientation than when rotated by 90 degrees. Think about an \"A\" or a \"P\" (or their Greek equivalents). If you cut out small parts of those, then the angles should look different after rotation.\n\nIn addition, the organizer's comment also mentions \"warping\"; so maybe there were some other slight distortions applied that made it harder to match patterns after rotation?"
  },
  "source": "meta"
}