{
  "id": 410747,
  "title": "Weird samples?",
  "url": "/competitions/asl-fingerspelling/discussion/410747",
  "author_name": "",
  "post_date": "2023-05-16T12:47:34.926694700Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.</p>\n<p><a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
  "messages": [
    {
      "id": "2261593",
      "postDate": "05/16/2023 12:47:34",
      "content": "<p>What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.</p>\n<p><a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
      "rawMarkdown": "What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.\n\n<a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "2261813",
      "postDate": "05/16/2023 15:06:55",
      "content": "<p>Please fix your image, it says \"image not found\"</p>",
      "rawMarkdown": "Please fix your image, it says \"image not found\"",
      "votes": null
    },
    {
      "id": "2261829",
      "postDate": "05/16/2023 15:14:33",
      "content": "<p>Sorry, Mark. Here is the new image</p>\n<p><a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
      "rawMarkdown": "Sorry, Mark. Here is the new image\n\n<a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "2262118",
      "postDate": "05/16/2023 18:34:30",
      "content": "<p>Thanks and sharp observation.</p>\n<p>There seem to be plenty of single frame samples labelled with long phrases as shown below.<br>\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&amp;alt=media\" alt=\"\"></p>\n<p>It might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.</p>\n<p>The general take away from your observation I believe is the dataset is quite noisy.</p>\n<p>I hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks and sharp observation.\n\nThere seem to be plenty of single frame samples labelled with long phrases as shown below.\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&alt=media)\n\nIt might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.\n\nThe general take away from your observation I believe is the dataset is quite noisy.\n\nI hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&alt=media)",
      "votes": null
    },
    {
      "id": "2262141",
      "postDate": "05/16/2023 18:50:57",
      "content": "<pre><code>There seem to be plenty of single frame samples labelled with long phrases as shown below.\n</code></pre>\n<p>Yep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. </p>\n<pre><code>The general take away from your observation I believe is the dataset is quite noisy.\n</code></pre>\n<p>Yes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"</p>\n<pre><code>I hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n</code></pre>\n<p>I hope so. </p>",
      "rawMarkdown": "```\nThere seem to be plenty of single frame samples labelled with long phrases as shown below.\n```\n\nYep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. \n\n\n```\nThe general take away from your observation I believe is the dataset is quite noisy.\n```\n\nYes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"\n\n```\nI hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n```\n\nI hope so.",
      "votes": null
    },
    {
      "id": "2295935",
      "postDate": "06/11/2023 12:06:09",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "sohier \nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&alt=media)",
      "votes": null
    },
    {
      "id": "2296130",
      "postDate": "06/11/2023 15:06:06",
      "content": "<p>there are also samples with only 1 frame 👀</p>",
      "rawMarkdown": "there are also samples with only 1 frame 👀",
      "votes": null
    },
    {
      "id": "2369334",
      "postDate": "08/01/2023 17:07:51",
      "content": "<p>Hey! Did you find an answer if the test dataset would have a similar pattern?</p>",
      "rawMarkdown": "Hey! Did you find an answer if the test dataset would have a similar pattern?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2261813,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "05/16/2023 15:06:55",
      "content": "<p>Please fix your image, it says \"image not found\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 2261829,
          "author_name": "vad13irt",
          "author_url": "",
          "post_date": "05/16/2023 15:14:33",
          "content": "<p>Sorry, Mark. Here is the new image</p>\n<p><a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2262118,
              "author_name": "markwijkhuizen",
              "author_url": "",
              "post_date": "05/16/2023 18:34:30",
              "content": "<p>Thanks and sharp observation.</p>\n<p>There seem to be plenty of single frame samples labelled with long phrases as shown below.<br>\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&amp;alt=media\" alt=\"\"></p>\n<p>It might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.</p>\n<p>The general take away from your observation I believe is the dataset is quite noisy.</p>\n<p>I hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2262141,
                  "author_name": "vad13irt",
                  "author_url": "",
                  "post_date": "05/16/2023 18:50:57",
                  "content": "<pre><code>There seem to be plenty of single frame samples labelled with long phrases as shown below.\n</code></pre>\n<p>Yep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. </p>\n<pre><code>The general take away from your observation I believe is the dataset is quite noisy.\n</code></pre>\n<p>Yes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"</p>\n<pre><code>I hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n</code></pre>\n<p>I hope so. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2295935,
                      "author_name": "kurupical",
                      "author_url": "",
                      "post_date": "06/11/2023 12:06:09",
                      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&amp;alt=media\" alt=\"\"></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2296130,
                          "author_name": "yuanzhezhou",
                          "author_url": "",
                          "post_date": "06/11/2023 15:06:06",
                          "content": "<p>there are also samples with only 1 frame 👀</p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 2369334,
                          "author_name": "quentinparrenin",
                          "author_url": "",
                          "post_date": "08/01/2023 17:07:51",
                          "content": "<p>Hey! Did you find an answer if the test dataset would have a similar pattern?</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2261593": "What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.\n\n<a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
    "2261813": "Please fix your image, it says \"image not found\"",
    "2261829": "Sorry, Mark. Here is the new image\n\n<a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
    "2262118": "Thanks and sharp observation.\n\nThere seem to be plenty of single frame samples labelled with long phrases as shown below.\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&alt=media)\n\nIt might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.\n\nThe general take away from your observation I believe is the dataset is quite noisy.\n\nI hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&alt=media)",
    "2262141": "```\nThere seem to be plenty of single frame samples labelled with long phrases as shown below.\n```\n\nYep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. \n\n\n```\nThe general take away from your observation I believe is the dataset is quite noisy.\n```\n\nYes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"\n\n```\nI hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n```\n\nI hope so.",
    "2295935": "sohier \nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&alt=media)",
    "2296130": "there are also samples with only 1 frame 👀",
    "2369334": "Hey! Did you find an answer if the test dataset would have a similar pattern?"
  },
  "source": "meta"
}