{
  "id": 574971,
  "title": "Data description",
  "url": "/competitions/beyond-visible-spectrum-ai-for-agriculture-2025/discussion/574971",
  "author_name": "Yauhen Tratsiak",
  "post_date": "2025-04-25T06:44:34.307000",
  "votes": 5,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi. The data description is not clear. Could you explain data structure in a bit more detailed way? It will be cool to have visualisation example with explanation</p>\n<p>What is wrong with '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot/sample2451.npy'?</p>",
  "messages": [
    {
      "id": 3186773,
      "postDate": "2025-04-25T06:44:34.307Z",
      "content": "<p>Hi. The data description is not clear. Could you explain data structure in a bit more detailed way? It will be cool to have visualisation example with explanation</p>\n<p>What is wrong with '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot/sample2451.npy'?</p>",
      "rawMarkdown": "Hi. The data description is not clear. Could you explain data structure in a bit more detailed way? It will be cool to have visualisation example with explanation\n\nWhat is wrong with '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot/sample2451.npy'?",
      "votes": 5
    },
    {
      "id": 3189422,
      "postDate": "2025-04-29T07:10:27.613Z",
      "content": "<p>I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). Thus, I decided to look on the data again. I removed some bad data (wrong shape or with empty data) and made some plots. All files were grouped by label. y - is sample, x - is channel<br>\nMean is in intensity<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fccc267a9bc0ee560412af2a617bfb347%2Fmean_by_label_by_channel.jpg?generation=1745910083798576&amp;alt=media\" alt=\"Mean is target\"></p>\n<p>STD in in intensity<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fc2d63d334dac8d694bef7a383df233d9%2Fstd_by_label_by_channel.jpg?generation=1745910110411656&amp;alt=media\" alt=\"STD is target\"></p>\n<p>We can see that the data are divided into two regions: before 70 channel (~730 nm ) and after it. Thus, one of the major questions could be possible light absorption by downy mildew to find the most appropriate channels for analysis. In general, the data looks messy and does not demonstrate any visible pattern (like increasing an intensity in specific channels with labels)</p>\n<p>it would be interesting to know how the labels were calculated</p>",
      "rawMarkdown": "I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). Thus, I decided to look on the data again. I removed some bad data (wrong shape or with empty data) and made some plots. All files were grouped by label. y - is sample, x - is channel\nMean is in intensity\n![Mean is target](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fccc267a9bc0ee560412af2a617bfb347%2Fmean_by_label_by_channel.jpg?generation=1745910083798576&alt=media)\n\nSTD in in intensity\n![STD is target](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fc2d63d334dac8d694bef7a383df233d9%2Fstd_by_label_by_channel.jpg?generation=1745910110411656&alt=media)\n\nWe can see that the data are divided into two regions: before 70 channel (~730 nm ) and after it. Thus, one of the major questions could be possible light absorption by downy mildew to find the most appropriate channels for analysis. In general, the data looks messy and does not demonstrate any visible pattern (like increasing an intensity in specific channels with labels)\n\nit would be interesting to know how the labels were calculated",
      "votes": 3,
      "replies": [
        {
          "id": 3189514,
          "postDate": "2025-04-29T10:11:54.977Z",
          "content": "<p>The data seems quite questionable, and I've uploaded my notebook with the analysis. Feel free to take a look.</p>",
          "rawMarkdown": "The data seems quite questionable, and I've uploaded my notebook with the analysis. Feel free to take a look."
        },
        {
          "id": 3190411,
          "postDate": "2025-04-30T16:20:33.510Z",
          "content": "<blockquote>\n  <p>I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). </p>\n</blockquote>\n<p>Pretty sure the data does not permit better than random guessing.</p>",
          "rawMarkdown": "> I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). \n\nPretty sure the data does not permit better than random guessing."
        }
      ]
    },
    {
      "id": 3187805,
      "postDate": "2025-04-26T14:55:25.323Z",
      "content": "<p>I made small research:<br>\nWhile trying to load /ot/ot/sample2451.npy, the following error occurred: ValueError: cannot reshape array of size 1785856 into shape (128,128,125)<br>\nManually opened the file and checked its size:<br>\nFile size: 3,571,840 bytes.<br>\nSince float32 uses 4 bytes per value, this corresponds to only 892,960 elements.<br>\nHowever, the expected array shape (128, 128, 125) requires 2,048,000 elements.</p>",
      "rawMarkdown": "I made small research:\nWhile trying to load /ot/ot/sample2451.npy, the following error occurred: ValueError: cannot reshape array of size 1785856 into shape (128,128,125)\nManually opened the file and checked its size:\nFile size: 3,571,840 bytes.\nSince float32 uses 4 bytes per value, this corresponds to only 892,960 elements.\nHowever, the expected array shape (128, 128, 125) requires 2,048,000 elements.",
      "votes": 1,
      "replies": [
        {
          "id": 3187819,
          "postDate": "2025-04-26T15:22:46.143Z",
          "content": "<p>I did the same. It corresponds 128, 128, 109. The organiser should check the data</p>",
          "rawMarkdown": "I did the same. It corresponds 128, 128, 109. The organiser should check the data",
          "votes": 1
        }
      ]
    },
    {
      "id": 3187704,
      "postDate": "2025-04-26T12:14:25.157Z",
      "content": "<p>The data hasn’t been cleaned yet, so the first 10 and last 14 bands are excluded due to noise. This is because of post-processing issues in the UAV hyperspectral data. All other files are clean, though.</p>",
      "rawMarkdown": "The data hasn’t been cleaned yet, so the first 10 and last 14 bands are excluded due to noise. This is because of post-processing issues in the UAV hyperspectral data. All other files are clean, though.",
      "votes": 1,
      "replies": [
        {
          "id": 3187745,
          "postDate": "2025-04-26T13:14:31.163Z",
          "content": "<p>one sample is corrupt and around 50 differ from usual size of 128x128 which can be padded or dropped but they're also in the test set so can't be dropped. </p>",
          "rawMarkdown": "one sample is corrupt and around 50 differ from usual size of 128x128 which can be padded or dropped but they're also in the test set so can't be dropped. ",
          "votes": 1,
          "replies": [
            {
              "id": 3187770,
              "postDate": "2025-04-26T13:56:11.380Z",
              "content": "<p>Train dataset<br>\nBad size : 345   (128, 57, 125)<br>\nBad size : 399   (128, 57, 125)<br>\nBad size : 442   (128, 57, 125)<br>\nBad size : 573   (128, 57, 125)<br>\nBad size : 618   (128, 57, 125)<br>\nBad size : 737   (128, 57, 125)<br>\nBad size : 742   (128, 57, 125)<br>\nBad size : 778   (128, 57, 125)<br>\nBad size : 781   (128, 57, 125)<br>\nBad size : 1130   (128, 57, 125)<br>\nBad size : 1201   (128, 57, 125)<br>\nBad size : 1249   (128, 57, 125)<br>\nBad size : 1272   (128, 57, 125)<br>\nBad size : 1296   (128, 57, 125)<br>\nBad size : 1355   (128, 57, 125)<br>\nBad size : 1377   (128, 57, 125)<br>\nBad size : 1393   (128, 57, 125)<br>\nBad size : 1399   (128, 57, 125)<br>\nBad size : 1472   (128, 57, 125)<br>\nBad size : 1487   (128, 57, 125)<br>\nBad size : 1544   (128, 57, 125)<br>\nBad size : 1644   (128, 57, 125)<br>\nBad size : 1690   (128, 57, 125)<br>\nBad size : 1784   (128, 57, 125)<br>\nBad size : 1818   (128, 57, 125)<br>\nBad size : 1832   (128, 57, 125)<br>\nBad size : 1845   (128, 57, 125)<br>\nBad size : 1875   (128, 57, 125)<br>\nBad size : 1881   (128, 57, 125)<br>\nBad size : 1900   (128, 57, 125)<br>\nBad size : 2152   (128, 57, 125)</p>\n<p>test dataset<br>\nBad size : 111   (128, 57, 125)<br>\nBad size : 131   (128, 57, 125)<br>\nBad size : 168   (128, 57, 125)<br>\nBad size : 201   (128, 57, 125)<br>\nBad size : 210   (128, 57, 125)<br>\nBad size : 261   (128, 57, 125)<br>\nBad size : 388   (128, 57, 125)<br>\nBad size : 429   (128, 57, 125)<br>\nBad size : 528   (128, 57, 125)</p>",
              "rawMarkdown": "Train dataset\nBad size : 345   (128, 57, 125)\nBad size : 399   (128, 57, 125)\nBad size : 442   (128, 57, 125)\nBad size : 573   (128, 57, 125)\nBad size : 618   (128, 57, 125)\nBad size : 737   (128, 57, 125)\nBad size : 742   (128, 57, 125)\nBad size : 778   (128, 57, 125)\nBad size : 781   (128, 57, 125)\nBad size : 1130   (128, 57, 125)\nBad size : 1201   (128, 57, 125)\nBad size : 1249   (128, 57, 125)\nBad size : 1272   (128, 57, 125)\nBad size : 1296   (128, 57, 125)\nBad size : 1355   (128, 57, 125)\nBad size : 1377   (128, 57, 125)\nBad size : 1393   (128, 57, 125)\nBad size : 1399   (128, 57, 125)\nBad size : 1472   (128, 57, 125)\nBad size : 1487   (128, 57, 125)\nBad size : 1544   (128, 57, 125)\nBad size : 1644   (128, 57, 125)\nBad size : 1690   (128, 57, 125)\nBad size : 1784   (128, 57, 125)\nBad size : 1818   (128, 57, 125)\nBad size : 1832   (128, 57, 125)\nBad size : 1845   (128, 57, 125)\nBad size : 1875   (128, 57, 125)\nBad size : 1881   (128, 57, 125)\nBad size : 1900   (128, 57, 125)\nBad size : 2152   (128, 57, 125)\n\ntest dataset\nBad size : 111   (128, 57, 125)\nBad size : 131   (128, 57, 125)\nBad size : 168   (128, 57, 125)\nBad size : 201   (128, 57, 125)\nBad size : 210   (128, 57, 125)\nBad size : 261   (128, 57, 125)\nBad size : 388   (128, 57, 125)\nBad size : 429   (128, 57, 125)\nBad size : 528   (128, 57, 125)\n",
              "votes": 2
            },
            {
              "id": 3191867,
              "postDate": "2025-05-02T08:38:55.417Z",
              "content": "<p>Yes, I have the same problem</p>",
              "rawMarkdown": "Yes, I have the same problem"
            },
            {
              "id": 3203901,
              "postDate": "2025-05-17T12:59:34.227Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3307362,
      "postDate": "2025-10-26T18:38:25.263Z",
      "content": "<p>what are the 100 classes how they label it ?</p>",
      "rawMarkdown": "what are the 100 classes how they label it ?\n"
    },
    {
      "id": 3194655,
      "postDate": "2025-05-06T06:43:06.853Z",
      "content": "<p>I am facing loading the data itself. I am getting the following error \"ValueError: num_samples should be a positive integer value, but got num_samples=0\". Really looking for some help as i am new to this. </p>",
      "rawMarkdown": "I am facing loading the data itself. I am getting the following error \"ValueError: num_samples should be a positive integer value, but got num_samples=0\". Really looking for some help as i am new to this. ",
      "replies": [
        {
          "id": 3194782,
          "postDate": "2025-05-06T09:34:14.517Z",
          "content": "<p>Don't worry, I can understand that you are new to kaggle .</p>\n<p>The error message you're encountering:<br>\n<code>ValueError: num_samples should be a positive integer value, but got num_samples=0</code><br>\nindicates that your dataset is empty or not properly loaded. This typically occurs when using a DataLoader in PyTorch with a dataset that has zero length.</p>\n<p>Here are some steps to resolve this issue:</p>\n<ol>\n<li>Check Dataset Path and Structure</li>\n<li>Validate Custom Dataset Implementation:-<br>\nIf you've implemented a custom dataset, make sure that the <strong>len</strong> method returns the correct number of samples.</li>\n</ol>\n<pre><code> (torch.utils.data.):\n     ():\n        .data = data\n\n     ():\n         len(.data)  \n\n     ():\n         .data[idx]\n</code></pre>\n<ol>\n<li>When creating a DataLoader, check the parameters passed, especially shuffle and sampler. Conflicting settings can lead to issues.</li>\n</ol>",
          "rawMarkdown": "Don't worry, I can understand that you are new to kaggle .\n\nThe error message you're encountering:\n`ValueError: num_samples should be a positive integer value, but got num_samples=0`\nindicates that your dataset is empty or not properly loaded. This typically occurs when using a DataLoader in PyTorch with a dataset that has zero length.\n\nHere are some steps to resolve this issue:\n1. Check Dataset Path and Structure\n2. Validate Custom Dataset Implementation:-\nIf you've implemented a custom dataset, make sure that the __len__ method returns the correct number of samples.\n```\nclass CustomDataset(torch.utils.data.Dataset):\n    def __init__(self, data):\n        self.data = data\n\n    def __len__(self):\n        return len(self.data)  # Ensure this returns a positive integer\n\n    def __getitem__(self, idx):\n        return self.data[idx]\n```\n3. When creating a DataLoader, check the parameters passed, especially shuffle and sampler. Conflicting settings can lead to issues.",
          "replies": [
            {
              "id": 3196618,
              "postDate": "2025-05-07T08:24:12.430Z",
              "content": "<p>Thanks sarthak, I was able to overcome the issue. </p>",
              "rawMarkdown": "Thanks sarthak, I was able to overcome the issue. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3186876,
      "postDate": "2025-04-25T09:13:15.727Z",
      "content": "<p>Will do it later in Code.</p>",
      "rawMarkdown": "Will do it later in Code.",
      "replies": [
        {
          "id": 3191866,
          "postDate": "2025-05-02T08:38:34.623Z",
          "content": "<p>The test set has 9 malformed examples:</p>\n<p>Expected array shape: (128, 128, 125)<br>\nAnomalous shapes and counts: {(128, 57, 125): 9}<br>\n['sample2520.npy',<br>\n 'sample1406.npy',<br>\n 'sample870.npy',<br>\n 'sample66.npy',<br>\n 'sample401.npy',<br>\n 'sample1138.npy',<br>\n 'sample2210.npy',<br>\n 'sample1540.npy',<br>\n 'sample2721.npy']</p>\n<p>And each malformed example has zero entries.</p>\n<p>Can you please update the data.</p>",
          "rawMarkdown": "The test set has 9 malformed examples:\n\nExpected array shape: (128, 128, 125)\nAnomalous shapes and counts: {(128, 57, 125): 9}\n['sample2520.npy',\n 'sample1406.npy',\n 'sample870.npy',\n 'sample66.npy',\n 'sample401.npy',\n 'sample1138.npy',\n 'sample2210.npy',\n 'sample1540.npy',\n 'sample2721.npy']\n\nAnd each malformed example has zero entries.\n\nCan you please update the data.\n",
          "replies": [
            {
              "id": 3191952,
              "postDate": "2025-05-02T10:00:02.687Z",
              "content": "<p>I believe it would be really helpful if the organizers could upload clean and uncorrupted samples, preferably from a noise-free version of the dataset. consistent evaluation for all participants.</p>",
              "rawMarkdown": "I believe it would be really helpful if the organizers could upload clean and uncorrupted samples, preferably from a noise-free version of the dataset. consistent evaluation for all participants.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3189422,
      "author_name": "Yauhen Tratsiak",
      "author_url": "",
      "post_date": "2025-04-29T07:10:27.613000",
      "content": "<p>I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). Thus, I decided to look on the data again. I removed some bad data (wrong shape or with empty data) and made some plots. All files were grouped by label. y - is sample, x - is channel<br>\nMean is in intensity<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fccc267a9bc0ee560412af2a617bfb347%2Fmean_by_label_by_channel.jpg?generation=1745910083798576&amp;alt=media\" alt=\"Mean is target\"></p>\n<p>STD in in intensity<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fc2d63d334dac8d694bef7a383df233d9%2Fstd_by_label_by_channel.jpg?generation=1745910110411656&amp;alt=media\" alt=\"STD is target\"></p>\n<p>We can see that the data are divided into two regions: before 70 channel (~730 nm ) and after it. Thus, one of the major questions could be possible light absorption by downy mildew to find the most appropriate channels for analysis. In general, the data looks messy and does not demonstrate any visible pattern (like increasing an intensity in specific channels with labels)</p>\n<p>it would be interesting to know how the labels were calculated</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3189514,
          "author_name": "Nikita Manaenkov",
          "author_url": "",
          "post_date": "2025-04-29T10:11:54.977000",
          "content": "<p>The data seems quite questionable, and I've uploaded my notebook with the analysis. Feel free to take a look.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3190411,
          "author_name": "Jan Reinecke",
          "author_url": "",
          "post_date": "2025-04-30T16:20:33.510000",
          "content": "<blockquote>\n  <p>I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). </p>\n</blockquote>\n<p>Pretty sure the data does not permit better than random guessing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3187805,
      "author_name": "Nikita Manaenkov",
      "author_url": "",
      "post_date": "2025-04-26T14:55:25.323000",
      "content": "<p>I made small research:<br>\nWhile trying to load /ot/ot/sample2451.npy, the following error occurred: ValueError: cannot reshape array of size 1785856 into shape (128,128,125)<br>\nManually opened the file and checked its size:<br>\nFile size: 3,571,840 bytes.<br>\nSince float32 uses 4 bytes per value, this corresponds to only 892,960 elements.<br>\nHowever, the expected array shape (128, 128, 125) requires 2,048,000 elements.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3187819,
          "author_name": "Yauhen Tratsiak",
          "author_url": "",
          "post_date": "2025-04-26T15:22:46.143000",
          "content": "<p>I did the same. It corresponds 128, 128, 109. The organiser should check the data</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3187704,
      "author_name": "Nikita Manaenkov",
      "author_url": "",
      "post_date": "2025-04-26T12:14:25.157000",
      "content": "<p>The data hasn’t been cleaned yet, so the first 10 and last 14 bands are excluded due to noise. This is because of post-processing issues in the UAV hyperspectral data. All other files are clean, though.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3187745,
          "author_name": "g john rao",
          "author_url": "",
          "post_date": "2025-04-26T13:14:31.163000",
          "content": "<p>one sample is corrupt and around 50 differ from usual size of 128x128 which can be padded or dropped but they're also in the test set so can't be dropped. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3187770,
              "author_name": "Yauhen Tratsiak",
              "author_url": "",
              "post_date": "2025-04-26T13:56:11.380000",
              "content": "<p>Train dataset<br>\nBad size : 345   (128, 57, 125)<br>\nBad size : 399   (128, 57, 125)<br>\nBad size : 442   (128, 57, 125)<br>\nBad size : 573   (128, 57, 125)<br>\nBad size : 618   (128, 57, 125)<br>\nBad size : 737   (128, 57, 125)<br>\nBad size : 742   (128, 57, 125)<br>\nBad size : 778   (128, 57, 125)<br>\nBad size : 781   (128, 57, 125)<br>\nBad size : 1130   (128, 57, 125)<br>\nBad size : 1201   (128, 57, 125)<br>\nBad size : 1249   (128, 57, 125)<br>\nBad size : 1272   (128, 57, 125)<br>\nBad size : 1296   (128, 57, 125)<br>\nBad size : 1355   (128, 57, 125)<br>\nBad size : 1377   (128, 57, 125)<br>\nBad size : 1393   (128, 57, 125)<br>\nBad size : 1399   (128, 57, 125)<br>\nBad size : 1472   (128, 57, 125)<br>\nBad size : 1487   (128, 57, 125)<br>\nBad size : 1544   (128, 57, 125)<br>\nBad size : 1644   (128, 57, 125)<br>\nBad size : 1690   (128, 57, 125)<br>\nBad size : 1784   (128, 57, 125)<br>\nBad size : 1818   (128, 57, 125)<br>\nBad size : 1832   (128, 57, 125)<br>\nBad size : 1845   (128, 57, 125)<br>\nBad size : 1875   (128, 57, 125)<br>\nBad size : 1881   (128, 57, 125)<br>\nBad size : 1900   (128, 57, 125)<br>\nBad size : 2152   (128, 57, 125)</p>\n<p>test dataset<br>\nBad size : 111   (128, 57, 125)<br>\nBad size : 131   (128, 57, 125)<br>\nBad size : 168   (128, 57, 125)<br>\nBad size : 201   (128, 57, 125)<br>\nBad size : 210   (128, 57, 125)<br>\nBad size : 261   (128, 57, 125)<br>\nBad size : 388   (128, 57, 125)<br>\nBad size : 429   (128, 57, 125)<br>\nBad size : 528   (128, 57, 125)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3191867,
              "author_name": "Utku Arslan",
              "author_url": "",
              "post_date": "2025-05-02T08:38:55.417000",
              "content": "<p>Yes, I have the same problem</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3203901,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-17T12:59:34.227000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3307362,
      "author_name": "IXRBHIII",
      "author_url": "",
      "post_date": "2025-10-26T18:38:25.263000",
      "content": "<p>what are the 100 classes how they label it ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3194655,
      "author_name": "Abhishek Sawkar",
      "author_url": "",
      "post_date": "2025-05-06T06:43:06.853000",
      "content": "<p>I am facing loading the data itself. I am getting the following error \"ValueError: num_samples should be a positive integer value, but got num_samples=0\". Really looking for some help as i am new to this. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3194782,
          "author_name": "Sarthak",
          "author_url": "",
          "post_date": "2025-05-06T09:34:14.517000",
          "content": "<p>Don't worry, I can understand that you are new to kaggle .</p>\n<p>The error message you're encountering:<br>\n<code>ValueError: num_samples should be a positive integer value, but got num_samples=0</code><br>\nindicates that your dataset is empty or not properly loaded. This typically occurs when using a DataLoader in PyTorch with a dataset that has zero length.</p>\n<p>Here are some steps to resolve this issue:</p>\n<ol>\n<li>Check Dataset Path and Structure</li>\n<li>Validate Custom Dataset Implementation:-<br>\nIf you've implemented a custom dataset, make sure that the <strong>len</strong> method returns the correct number of samples.</li>\n</ol>\n<pre><code> (torch.utils.data.):\n     ():\n        .data = data\n\n     ():\n         len(.data)  \n\n     ():\n         .data[idx]\n</code></pre>\n<ol>\n<li>When creating a DataLoader, check the parameters passed, especially shuffle and sampler. Conflicting settings can lead to issues.</li>\n</ol>",
          "votes": 0,
          "replies": [
            {
              "id": 3196618,
              "author_name": "Abhishek Sawkar",
              "author_url": "",
              "post_date": "2025-05-07T08:24:12.430000",
              "content": "<p>Thanks sarthak, I was able to overcome the issue. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3186876,
      "author_name": "robeson",
      "author_url": "",
      "post_date": "2025-04-25T09:13:15.727000",
      "content": "<p>Will do it later in Code.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3191866,
          "author_name": "Utku Arslan",
          "author_url": "",
          "post_date": "2025-05-02T08:38:34.623000",
          "content": "<p>The test set has 9 malformed examples:</p>\n<p>Expected array shape: (128, 128, 125)<br>\nAnomalous shapes and counts: {(128, 57, 125): 9}<br>\n['sample2520.npy',<br>\n 'sample1406.npy',<br>\n 'sample870.npy',<br>\n 'sample66.npy',<br>\n 'sample401.npy',<br>\n 'sample1138.npy',<br>\n 'sample2210.npy',<br>\n 'sample1540.npy',<br>\n 'sample2721.npy']</p>\n<p>And each malformed example has zero entries.</p>\n<p>Can you please update the data.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3191952,
              "author_name": "Sarthak",
              "author_url": "",
              "post_date": "2025-05-02T10:00:02.687000",
              "content": "<p>I believe it would be really helpful if the organizers could upload clean and uncorrupted samples, preferably from a noise-free version of the dataset. consistent evaluation for all participants.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3186773": "Hi. The data description is not clear. Could you explain data structure in a bit more detailed way? It will be cool to have visualisation example with explanation\n\nWhat is wrong with '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot/sample2451.npy'?",
    "3189422": "I see the best score is around 850, which most probably corresponds to random predictions (I think we all have a problem with shitty predictions). Thus, I decided to look on the data again. I removed some bad data (wrong shape or with empty data) and made some plots. All files were grouped by label. y - is sample, x - is channel\nMean is in intensity\n![Mean is target](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fccc267a9bc0ee560412af2a617bfb347%2Fmean_by_label_by_channel.jpg?generation=1745910083798576&alt=media)\n\nSTD in in intensity\n![STD is target](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6635386%2Fc2d63d334dac8d694bef7a383df233d9%2Fstd_by_label_by_channel.jpg?generation=1745910110411656&alt=media)\n\nWe can see that the data are divided into two regions: before 70 channel (~730 nm ) and after it. Thus, one of the major questions could be possible light absorption by downy mildew to find the most appropriate channels for analysis. In general, the data looks messy and does not demonstrate any visible pattern (like increasing an intensity in specific channels with labels)\n\nit would be interesting to know how the labels were calculated",
    "3187805": "I made small research:\nWhile trying to load /ot/ot/sample2451.npy, the following error occurred: ValueError: cannot reshape array of size 1785856 into shape (128,128,125)\nManually opened the file and checked its size:\nFile size: 3,571,840 bytes.\nSince float32 uses 4 bytes per value, this corresponds to only 892,960 elements.\nHowever, the expected array shape (128, 128, 125) requires 2,048,000 elements.",
    "3187704": "The data hasn’t been cleaned yet, so the first 10 and last 14 bands are excluded due to noise. This is because of post-processing issues in the UAV hyperspectral data. All other files are clean, though.",
    "3307362": "what are the 100 classes how they label it ?\n",
    "3194655": "I am facing loading the data itself. I am getting the following error \"ValueError: num_samples should be a positive integer value, but got num_samples=0\". Really looking for some help as i am new to this. ",
    "3186876": "Will do it later in Code."
  }
}