{
  "id": 407563,
  "title": "Question about the Hidden Dataset ",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/407563",
  "author_name": "",
  "post_date": "2023-05-07T06:07:24.826988500Z",
  "votes": -1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>During my submit, I found that the hidden dataset may have very large tif files (I think it is greater than four times test/b).</p>\n<p>So I tried to use <a href=\"optimized-data-reading-on-example-submission\" target=\"_blank\">https://www.kaggle.com/code/riferji/optimized-data-reading-on-example-submission</a> to split the data into four parts to load into my model.</p>\n<p>But I've run into some really tough problems. I tested my code's ability to handle large images with train/2 data (because it is also very large), and it works. But when I submit my code, I keep getting <code>Notebook Threw Exception</code>, and I have no way to reproduce it. 😥</p>\n<p>Can anyone provide some solution ideas for this kind of bug that has no way to be reproduced? Thank you very much!❤️</p>",
  "messages": [
    {
      "id": "2248714",
      "postDate": "05/07/2023 06:07:24",
      "content": "<p>During my submit, I found that the hidden dataset may have very large tif files (I think it is greater than four times test/b).</p>\n<p>So I tried to use <a href=\"optimized-data-reading-on-example-submission\" target=\"_blank\">https://www.kaggle.com/code/riferji/optimized-data-reading-on-example-submission</a> to split the data into four parts to load into my model.</p>\n<p>But I've run into some really tough problems. I tested my code's ability to handle large images with train/2 data (because it is also very large), and it works. But when I submit my code, I keep getting <code>Notebook Threw Exception</code>, and I have no way to reproduce it. 😥</p>\n<p>Can anyone provide some solution ideas for this kind of bug that has no way to be reproduced? Thank you very much!❤️</p>",
      "rawMarkdown": "During my submit, I found that the hidden dataset may have very large tif files (I think it is greater than four times test/b).\n\nSo I tried to use [https://www.kaggle.com/code/riferji/optimized-data-reading-on-example-submission](optimized-data-reading-on-example-submission) to split the data into four parts to load into my model.\n\nBut I've run into some really tough problems. I tested my code's ability to handle large images with train/2 data (because it is also very large), and it works. But when I submit my code, I keep getting `Notebook Threw Exception`, and I have no way to reproduce it. 😥\n\nCan anyone provide some solution ideas for this kind of bug that has no way to be reproduced? Thank you very much!❤️",
      "votes": null
    },
    {
      "id": "2248726",
      "postDate": "05/07/2023 06:19:48",
      "content": "<p>I'm wondering if it's possible that the file for the hidden dataset is a bit different from the file for the public dataset (for example, the tif for the public dataset is (5454, 6330) when read, while the hidden dataset is (1, 5454, 6330) when read)</p>",
      "rawMarkdown": "I'm wondering if it's possible that the file for the hidden dataset is a bit different from the file for the public dataset (for example, the tif for the public dataset is (5454, 6330) when read, while the hidden dataset is (1, 5454, 6330) when read)",
      "votes": null
    },
    {
      "id": "2249261",
      "postDate": "05/07/2023 15:24:26",
      "content": "<p>it could be a lot of thing, i just started here so i had a lot problems includind rejected submission, it was mainly due to:</p>\n<ul>\n<li>internet enabled (it need internet disabled:<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/401945\" target=\"_blank\"> so you cant install library directly in the submission notebook with internet</a>) so you cant install with pip or whatever with internet<br>\n-bad rle/csv file (something like an index in the dataframe i think)<br>\n-random  key press to convert a code cell in markdown cell (the worst)<br>\nmaybe it could help you to find the issue…</li>\n</ul>",
      "rawMarkdown": "it could be a lot of thing, i just started here so i had a lot problems includind rejected submission, it was mainly due to:\n- internet enabled (it need internet disabled:[ so you cant install library directly in the submission notebook with internet](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/401945)) so you cant install with pip or whatever with internet\n -bad rle/csv file (something like an index in the dataframe i think)\n-random  key press to convert a code cell in markdown cell (the worst)\nmaybe it could help you to find the issue...",
      "votes": null
    },
    {
      "id": "2249657",
      "postDate": "05/08/2023 01:38:44",
      "content": "<p>Thanks so much, I will try it😁</p>",
      "rawMarkdown": "Thanks so much, I will try it😁",
      "votes": null
    },
    {
      "id": "2255585",
      "postDate": "05/11/2023 20:33:11",
      "content": "<p>I'm still struggling on getting a valid submission, and I noticed a lot of others are as well. Although the hidden data does seem to be larger than the available test fragments, I don't get the \"Notebook Exceeded Allowed Compute\" Error, so I think you might be okay on that front. Here are a couple more things you can try:</p>\n<ul>\n<li>Check the size of the model output. Make sure it is the same size as the original</li>\n<li>Make sure that the file you want actually exists</li>\n<li>See if the test submission file has data</li>\n<li>As a last resort, you can split up the notebook to find the section that causes the rejected submission.</li>\n</ul>\n<p>Hope this helps!</p>",
      "rawMarkdown": "I'm still struggling on getting a valid submission, and I noticed a lot of others are as well. Although the hidden data does seem to be larger than the available test fragments, I don't get the \"Notebook Exceeded Allowed Compute\" Error, so I think you might be okay on that front. Here are a couple more things you can try:\n- Check the size of the model output. Make sure it is the same size as the original\n- Make sure that the file you want actually exists\n- See if the test submission file has data\n- As a last resort, you can split up the notebook to find the section that causes the rejected submission.\n\nHope this helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2248726,
      "author_name": "jimmyisme1",
      "author_url": "",
      "post_date": "05/07/2023 06:19:48",
      "content": "<p>I'm wondering if it's possible that the file for the hidden dataset is a bit different from the file for the public dataset (for example, the tif for the public dataset is (5454, 6330) when read, while the hidden dataset is (1, 5454, 6330) when read)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2249261,
          "author_name": "iraqbot",
          "author_url": "",
          "post_date": "05/07/2023 15:24:26",
          "content": "<p>it could be a lot of thing, i just started here so i had a lot problems includind rejected submission, it was mainly due to:</p>\n<ul>\n<li>internet enabled (it need internet disabled:<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/401945\" target=\"_blank\"> so you cant install library directly in the submission notebook with internet</a>) so you cant install with pip or whatever with internet<br>\n-bad rle/csv file (something like an index in the dataframe i think)<br>\n-random  key press to convert a code cell in markdown cell (the worst)<br>\nmaybe it could help you to find the issue…</li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2249657,
              "author_name": "jimmyisme1",
              "author_url": "",
              "post_date": "05/08/2023 01:38:44",
              "content": "<p>Thanks so much, I will try it😁</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2255585,
          "author_name": "kmarkey",
          "author_url": "",
          "post_date": "05/11/2023 20:33:11",
          "content": "<p>I'm still struggling on getting a valid submission, and I noticed a lot of others are as well. Although the hidden data does seem to be larger than the available test fragments, I don't get the \"Notebook Exceeded Allowed Compute\" Error, so I think you might be okay on that front. Here are a couple more things you can try:</p>\n<ul>\n<li>Check the size of the model output. Make sure it is the same size as the original</li>\n<li>Make sure that the file you want actually exists</li>\n<li>See if the test submission file has data</li>\n<li>As a last resort, you can split up the notebook to find the section that causes the rejected submission.</li>\n</ul>\n<p>Hope this helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2248714": "During my submit, I found that the hidden dataset may have very large tif files (I think it is greater than four times test/b).\n\nSo I tried to use [https://www.kaggle.com/code/riferji/optimized-data-reading-on-example-submission](optimized-data-reading-on-example-submission) to split the data into four parts to load into my model.\n\nBut I've run into some really tough problems. I tested my code's ability to handle large images with train/2 data (because it is also very large), and it works. But when I submit my code, I keep getting `Notebook Threw Exception`, and I have no way to reproduce it. 😥\n\nCan anyone provide some solution ideas for this kind of bug that has no way to be reproduced? Thank you very much!❤️",
    "2248726": "I'm wondering if it's possible that the file for the hidden dataset is a bit different from the file for the public dataset (for example, the tif for the public dataset is (5454, 6330) when read, while the hidden dataset is (1, 5454, 6330) when read)",
    "2249261": "it could be a lot of thing, i just started here so i had a lot problems includind rejected submission, it was mainly due to:\n- internet enabled (it need internet disabled:[ so you cant install library directly in the submission notebook with internet](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/401945)) so you cant install with pip or whatever with internet\n -bad rle/csv file (something like an index in the dataframe i think)\n-random  key press to convert a code cell in markdown cell (the worst)\nmaybe it could help you to find the issue...",
    "2249657": "Thanks so much, I will try it😁",
    "2255585": "I'm still struggling on getting a valid submission, and I noticed a lot of others are as well. Although the hidden data does seem to be larger than the available test fragments, I don't get the \"Notebook Exceeded Allowed Compute\" Error, so I think you might be okay on that front. Here are a couple more things you can try:\n- Check the size of the model output. Make sure it is the same size as the original\n- Make sure that the file you want actually exists\n- See if the test submission file has data\n- As a last resort, you can split up the notebook to find the section that causes the rejected submission.\n\nHope this helps!"
  },
  "source": "meta"
}