{
  "id": 465824,
  "title": "test data filename - numbers? continous?",
  "url": "/competitions/blood-vessel-segmentation/discussion/465824",
  "author_name": "",
  "post_date": "2024-01-05T19:20:14.784711300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I am having trouble having my prediction to succeed, even though it appears to work all well with train data.</p>\n<p>I am really struggling to understand the reason why but I think it has something to do with the images filenames in the test data.</p>\n<p>The training data is organised as /images/&lt;4digitnumber&gt;.tif<br>\nand so is the small test sample of kidney_5 and kidney_6</p>\n<p>One of my notebooks, I run inference per image, so the csv id output is _&lt;4digitnumber&gt;, extracted from the file path, and next column the respective RLE.</p>\n<p>I have created a 3D version. I am assuming that the &lt;4digitnumber&gt; are continous slice numbers, so that I can create a 3D array when putting slices together. Then the 3D inference is executed, and then RLE-encode slice-by-slice, with numbering starting from 0000 and assuming they are continous numbering 0000, 0001, 0002. It appears that this is resulting in failure and throwing an error saying \"Submission Scoring Error. Your notebook generated a submission file with incorrect format. …\".</p>\n<p>I also tried the procedure of collecting the data to a 3D volume, running inference in 2D (slice-by-slice) and encode each slice. RLE information is then added to the csv file with id from the z-coordinate. Even though this method is identical to the 2D inference described above, the only difference is the way the id in the CSV file was created, sucessful when the filename of the slice is used vs no-success when the z slice coordinate is used.</p>\n<p>If the numbering is different, it may be that the slices are not continuous but jump (slices missing?), or simply there are other characters instead of just digits in the filename.</p>\n<p>Of course the reason why it is failing may be something else. In that case… I'm stuck.</p>\n<p>Do the images filename numbering in the test data follow the same pattern as in the train data (0000 to lastz, step 1) ?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "2588842",
      "postDate": "01/05/2024 19:20:14",
      "content": "<p>Hi,</p>\n<p>I am having trouble having my prediction to succeed, even though it appears to work all well with train data.</p>\n<p>I am really struggling to understand the reason why but I think it has something to do with the images filenames in the test data.</p>\n<p>The training data is organised as /images/&lt;4digitnumber&gt;.tif<br>\nand so is the small test sample of kidney_5 and kidney_6</p>\n<p>One of my notebooks, I run inference per image, so the csv id output is _&lt;4digitnumber&gt;, extracted from the file path, and next column the respective RLE.</p>\n<p>I have created a 3D version. I am assuming that the &lt;4digitnumber&gt; are continous slice numbers, so that I can create a 3D array when putting slices together. Then the 3D inference is executed, and then RLE-encode slice-by-slice, with numbering starting from 0000 and assuming they are continous numbering 0000, 0001, 0002. It appears that this is resulting in failure and throwing an error saying \"Submission Scoring Error. Your notebook generated a submission file with incorrect format. …\".</p>\n<p>I also tried the procedure of collecting the data to a 3D volume, running inference in 2D (slice-by-slice) and encode each slice. RLE information is then added to the csv file with id from the z-coordinate. Even though this method is identical to the 2D inference described above, the only difference is the way the id in the CSV file was created, sucessful when the filename of the slice is used vs no-success when the z slice coordinate is used.</p>\n<p>If the numbering is different, it may be that the slices are not continuous but jump (slices missing?), or simply there are other characters instead of just digits in the filename.</p>\n<p>Of course the reason why it is failing may be something else. In that case… I'm stuck.</p>\n<p>Do the images filename numbering in the test data follow the same pattern as in the train data (0000 to lastz, step 1) ?</p>\n<p>Thank you</p>",
      "rawMarkdown": "Hi,\n\nI am having trouble having my prediction to succeed, even though it appears to work all well with train data.\n\nI am really struggling to understand the reason why but I think it has something to do with the images filenames in the test data.\n\nThe training data is organised as <sample>/images/<4digitnumber>.tif\nand so is the small test sample of kidney_5 and kidney_6\n\nOne of my notebooks, I run inference per image, so the csv id output is <sample>_<4digitnumber>, extracted from the file path, and next column the respective RLE.\n\nI have created a 3D version. I am assuming that the <4digitnumber> are continous slice numbers, so that I can create a 3D array when putting slices together. Then the 3D inference is executed, and then RLE-encode slice-by-slice, with numbering starting from 0000 and assuming they are continous numbering 0000, 0001, 0002. It appears that this is resulting in failure and throwing an error saying \"Submission Scoring Error. Your notebook generated a submission file with incorrect format. ...\".\n\nI also tried the procedure of collecting the data to a 3D volume, running inference in 2D (slice-by-slice) and encode each slice. RLE information is then added to the csv file with id from the z-coordinate. Even though this method is identical to the 2D inference described above, the only difference is the way the id in the CSV file was created, sucessful when the filename of the slice is used vs no-success when the z slice coordinate is used.\n\nIf the numbering is different, it may be that the slices are not continuous but jump (slices missing?), or simply there are other characters instead of just digits in the filename.\n\nOf course the reason why it is failing may be something else. In that case... I'm stuck.\n\nDo the images filename numbering in the test data follow the same pattern as in the train data (0000 to lastz, step 1) ?\n\nThank you",
      "votes": null
    },
    {
      "id": "2588855",
      "postDate": "01/05/2024 19:46:06",
      "content": "<p>They are continuos but not necessary starting from 0. Anyway, read the keys from whatever is in test folders, sort them, and use them whatever they are.</p>",
      "rawMarkdown": "They are continuos but not necessary starting from 0. Anyway, read the keys from whatever is in test folders, sort them, and use them whatever they are.",
      "votes": null
    },
    {
      "id": "2588859",
      "postDate": "01/05/2024 19:53:58",
      "content": "<p>you saved my day, thank you</p>",
      "rawMarkdown": "you saved my day, thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2588855,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "01/05/2024 19:46:06",
      "content": "<p>They are continuos but not necessary starting from 0. Anyway, read the keys from whatever is in test folders, sort them, and use them whatever they are.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2588859,
          "author_name": "perdigao1",
          "author_url": "",
          "post_date": "01/05/2024 19:53:58",
          "content": "<p>you saved my day, thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2588842": "Hi,\n\nI am having trouble having my prediction to succeed, even though it appears to work all well with train data.\n\nI am really struggling to understand the reason why but I think it has something to do with the images filenames in the test data.\n\nThe training data is organised as <sample>/images/<4digitnumber>.tif\nand so is the small test sample of kidney_5 and kidney_6\n\nOne of my notebooks, I run inference per image, so the csv id output is <sample>_<4digitnumber>, extracted from the file path, and next column the respective RLE.\n\nI have created a 3D version. I am assuming that the <4digitnumber> are continous slice numbers, so that I can create a 3D array when putting slices together. Then the 3D inference is executed, and then RLE-encode slice-by-slice, with numbering starting from 0000 and assuming they are continous numbering 0000, 0001, 0002. It appears that this is resulting in failure and throwing an error saying \"Submission Scoring Error. Your notebook generated a submission file with incorrect format. ...\".\n\nI also tried the procedure of collecting the data to a 3D volume, running inference in 2D (slice-by-slice) and encode each slice. RLE information is then added to the csv file with id from the z-coordinate. Even though this method is identical to the 2D inference described above, the only difference is the way the id in the CSV file was created, sucessful when the filename of the slice is used vs no-success when the z slice coordinate is used.\n\nIf the numbering is different, it may be that the slices are not continuous but jump (slices missing?), or simply there are other characters instead of just digits in the filename.\n\nOf course the reason why it is failing may be something else. In that case... I'm stuck.\n\nDo the images filename numbering in the test data follow the same pattern as in the train data (0000 to lastz, step 1) ?\n\nThank you",
    "2588855": "They are continuos but not necessary starting from 0. Anyway, read the keys from whatever is in test folders, sort them, and use them whatever they are.",
    "2588859": "you saved my day, thank you"
  },
  "source": "meta"
}