{
  "id": 473024,
  "title": "Submission scoring error",
  "url": "/competitions/blood-vessel-segmentation/discussion/473024",
  "author_name": "",
  "post_date": "2024-02-03T06:22:03.640056300Z",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>My submission completed after a few minutes, but the scoring process took longer and encountered an error. Has anyone else faced a similar issue?</p>",
  "messages": [
    {
      "id": "2633577",
      "postDate": "02/03/2024 06:22:03",
      "content": "<p>My submission completed after a few minutes, but the scoring process took longer and encountered an error. Has anyone else faced a similar issue?</p>",
      "rawMarkdown": "My submission completed after a few minutes, but the scoring process took longer and encountered an error. Has anyone else faced a similar issue?",
      "votes": null
    },
    {
      "id": "2633851",
      "postDate": "02/03/2024 10:21:38",
      "content": "<p>Hi. The one that completed in few minuts was executed at example test folder with 6 example files. The one that took longer was executed on real test folder. Without more information and If the example didn't fail I'd say that there is something wrong with id. If that's the case don't try generate any id, read them from the files in test and use them sorted as you read them.</p>",
      "rawMarkdown": "Hi. The one that completed in few minuts was executed at example test folder with 6 example files. The one that took longer was executed on real test folder. Without more information and If the example didn't fail I'd say that there is something wrong with id. If that's the case don't try generate any id, read them from the files in test and use them sorted as you read them.",
      "votes": null
    },
    {
      "id": "2634718",
      "postDate": "02/04/2024 00:50:21",
      "content": "<p>I've carefully reviewed the file IDs and resolved the identified issue. I am runing the code and submit again to chck if it is fixed. However, I suspect there might be some problem elsewhere. </p>\n<p>Your suggestion was helpful—thank you!</p>",
      "rawMarkdown": "I've carefully reviewed the file IDs and resolved the identified issue. I am runing the code and submit again to chck if it is fixed. However, I suspect there might be some problem elsewhere. \n\nYour suggestion was helpful—thank you!",
      "votes": null
    },
    {
      "id": "2635878",
      "postDate": "02/04/2024 16:53:04",
      "content": "<p>First of all, thank you very much! Through your comments and tips here in the discussion, I finally managed to submit after 2 months of a lot of effort.</p>\n<p>Secondly, my score is currently at zero, and there are 2 days left. People mentioned that this may be due to the threshold of binarization, saying that the threshold should be high, close to 255. However, I am still scoring zero. Do you have any tips for me to check in my algorithm to really score in this competition?</p>\n<p>I'm kind of desperate because I only have 5 submissions per day, and I've already used today's, so there are only 5 left until the end.</p>",
      "rawMarkdown": "First of all, thank you very much! Through your comments and tips here in the discussion, I finally managed to submit after 2 months of a lot of effort.\n\nSecondly, my score is currently at zero, and there are 2 days left. People mentioned that this may be due to the threshold of binarization, saying that the threshold should be high, close to 255. However, I am still scoring zero. Do you have any tips for me to check in my algorithm to really score in this competition?\n\nI'm kind of desperate because I only have 5 submissions per day, and I've already used today's, so there are only 5 left until the end.",
      "votes": null
    },
    {
      "id": "2636071",
      "postDate": "02/04/2024 19:37:14",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/471055\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/471055</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/471055",
      "votes": null
    },
    {
      "id": "2636082",
      "postDate": "02/04/2024 19:41:46",
      "content": "<p>You're welcome! I'll check this out!!</p>",
      "rawMarkdown": "You're welcome! I'll check this out!!",
      "votes": null
    },
    {
      "id": "2636184",
      "postDate": "02/04/2024 21:40:02",
      "content": "<p>The zero score could be due to the different functions used in training and inference.</p>\n<p>Thanks for the suggestion. I thoroughly reviewed my training and inference notebooks, and I've resolved the issue. </p>",
      "rawMarkdown": "The zero score could be due to the different functions used in training and inference.\n\nThanks for the suggestion. I thoroughly reviewed my training and inference notebooks, and I've resolved the issue.",
      "votes": null
    },
    {
      "id": "2636352",
      "postDate": "02/05/2024 03:16:04",
      "content": "<p>I found the memory usage was a potential cause of slowdown. By clearing memory after each image inference, I was able to improve execution speed and avoid submission timeouts.</p>",
      "rawMarkdown": "I found the memory usage was a potential cause of slowdown. By clearing memory after each image inference, I was able to improve execution speed and avoid submission timeouts.",
      "votes": null
    },
    {
      "id": "2637659",
      "postDate": "02/05/2024 20:02:41",
      "content": "<p>I found the cause of the score 0.</p>\n<p>The model I trained works in the training notebook, generating predicted masks for unknown images.</p>\n<p>However, when I use the same model in the inference notebook, the model produces empty masks, even with the dataset it was trained on.</p>\n<p>Now I don't know if the issue is with the model loading or if it's related to the preprocessing.</p>",
      "rawMarkdown": "I found the cause of the score 0.\n\nThe model I trained works in the training notebook, generating predicted masks for unknown images.\n\nHowever, when I use the same model in the inference notebook, the model produces empty masks, even with the dataset it was trained on.\n\nNow I don't know if the issue is with the model loading or if it's related to the preprocessing.",
      "votes": null
    },
    {
      "id": "2637684",
      "postDate": "02/05/2024 20:12:55",
      "content": "<p>Check the normalization part. If the normalization values are different between train and test then I think this behavior might be expected. I had something similar and when I just used Z-score of train data for both training and inference the problem has been resolved.</p>",
      "rawMarkdown": "Check the normalization part. If the normalization values are different between train and test then I think this behavior might be expected. I had something similar and when I just used Z-score of train data for both training and inference the problem has been resolved.",
      "votes": null
    },
    {
      "id": "2637744",
      "postDate": "02/05/2024 20:48:40",
      "content": "<p>Thanks!</p>\n<p>I'll be honest with you, I don't know what that z-score normalization would be. I'm working with 2D, specifically with kidney_1_dense. I still have a lot to learn to reach that level.</p>",
      "rawMarkdown": "Thanks!\n\nI'll be honest with you, I don't know what that z-score normalization would be. I'm working with 2D, specifically with kidney_1_dense. I still have a lot to learn to reach that level.",
      "votes": null
    },
    {
      "id": "2638858",
      "postDate": "02/06/2024 14:24:13",
      "content": "<p>Friend, I just realized that it wasn't the model that was causing empty masks, it was the binarization that was leaving all the masks empty. I switched to the Otsu method, but it's making the masks dirty. As a result, the score went from 0 to 0.044. I'm going to try the binarization method that Bhavya Dhingra mentioned in his notebook \"Clean Code 📚| Weighted Ensemble [Inference]\". Wish me luck! And thank you very much!</p>",
      "rawMarkdown": "Friend, I just realized that it wasn't the model that was causing empty masks, it was the binarization that was leaving all the masks empty. I switched to the Otsu method, but it's making the masks dirty. As a result, the score went from 0 to 0.044. I'm going to try the binarization method that Bhavya Dhingra mentioned in his notebook \"Clean Code 📚| Weighted Ensemble [Inference]\". Wish me luck! And thank you very much!",
      "votes": null
    },
    {
      "id": "2638887",
      "postDate": "02/06/2024 14:36:38",
      "content": "<p>Good luck!<br>\nAnyway I am suggesting (if there is enough time) to revise the normalization again.<br>\nAssume a dataset with only one feature. This feature ranges from 0 to 100 in the training data, and say you used min_max_scaler to make it from 0 to 1. Then in test data you have the same feature ranges from 20 to 80, and you used min_max_scaler (not the one fitted on training data) to make the range between 0 and 1.<br>\nWhen the model get fitted into the training data, the 0s, for example, of training data are NOT the same as the 0s of the testing data. So, it is expected to produce quite bad results (unless you are lucky enough to have training data and testing data both between 0 and 100). That's why it is a good practice to scale test data with same scaler you used in training data. <br>\nIn the public nbs they fitted scalers for each dataset independently and that worked somehow lol, but this behaviour is risky and can lead to quite bad results in test data.</p>",
      "rawMarkdown": "Good luck!\nAnyway I am suggesting (if there is enough time) to revise the normalization again.\nAssume a dataset with only one feature. This feature ranges from 0 to 100 in the training data, and say you used min_max_scaler to make it from 0 to 1. Then in test data you have the same feature ranges from 20 to 80, and you used min_max_scaler (not the one fitted on training data) to make the range between 0 and 1.\nWhen the model get fitted into the training data, the 0s, for example, of training data are NOT the same as the 0s of the testing data. So, it is expected to produce quite bad results (unless you are lucky enough to have training data and testing data both between 0 and 100). That's why it is a good practice to scale test data with same scaler you used in training data. \nIn the public nbs they fitted scalers for each dataset independently and that worked somehow lol, but this behaviour is risky and can lead to quite bad results in test data.",
      "votes": null
    },
    {
      "id": "2640343",
      "postDate": "02/06/2024 20:50:11",
      "content": "<p>What I did is to compare training and inference notebooks for function inconsistencies, especially scaling and processing functions. </p>\n<p>Always used training notebook functions to ensure consistency across both notebooks.</p>",
      "rawMarkdown": "What I did is to compare training and inference notebooks for function inconsistencies, especially scaling and processing functions. \n\nAlways used training notebook functions to ensure consistency across both notebooks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2633851,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/03/2024 10:21:38",
      "content": "<p>Hi. The one that completed in few minuts was executed at example test folder with 6 example files. The one that took longer was executed on real test folder. Without more information and If the example didn't fail I'd say that there is something wrong with id. If that's the case don't try generate any id, read them from the files in test and use them sorted as you read them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2634718,
          "author_name": "minhsienweng",
          "author_url": "",
          "post_date": "02/04/2024 00:50:21",
          "content": "<p>I've carefully reviewed the file IDs and resolved the identified issue. I am runing the code and submit again to chck if it is fixed. However, I suspect there might be some problem elsewhere. </p>\n<p>Your suggestion was helpful—thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2635878,
          "author_name": "fernandoaesteves",
          "author_url": "",
          "post_date": "02/04/2024 16:53:04",
          "content": "<p>First of all, thank you very much! Through your comments and tips here in the discussion, I finally managed to submit after 2 months of a lot of effort.</p>\n<p>Secondly, my score is currently at zero, and there are 2 days left. People mentioned that this may be due to the threshold of binarization, saying that the threshold should be high, close to 255. However, I am still scoring zero. Do you have any tips for me to check in my algorithm to really score in this competition?</p>\n<p>I'm kind of desperate because I only have 5 submissions per day, and I've already used today's, so there are only 5 left until the end.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2636071,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "02/04/2024 19:37:14",
              "content": "<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/471055\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/471055</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2636082,
                  "author_name": "fernandoaesteves",
                  "author_url": "",
                  "post_date": "02/04/2024 19:41:46",
                  "content": "<p>You're welcome! I'll check this out!!</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2636184,
                  "author_name": "minhsienweng",
                  "author_url": "",
                  "post_date": "02/04/2024 21:40:02",
                  "content": "<p>The zero score could be due to the different functions used in training and inference.</p>\n<p>Thanks for the suggestion. I thoroughly reviewed my training and inference notebooks, and I've resolved the issue. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2637659,
                      "author_name": "fernandoaesteves",
                      "author_url": "",
                      "post_date": "02/05/2024 20:02:41",
                      "content": "<p>I found the cause of the score 0.</p>\n<p>The model I trained works in the training notebook, generating predicted masks for unknown images.</p>\n<p>However, when I use the same model in the inference notebook, the model produces empty masks, even with the dataset it was trained on.</p>\n<p>Now I don't know if the issue is with the model loading or if it's related to the preprocessing.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2637684,
                          "author_name": "mohammad2012191",
                          "author_url": "",
                          "post_date": "02/05/2024 20:12:55",
                          "content": "<p>Check the normalization part. If the normalization values are different between train and test then I think this behavior might be expected. I had something similar and when I just used Z-score of train data for both training and inference the problem has been resolved.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2637744,
                              "author_name": "fernandoaesteves",
                              "author_url": "",
                              "post_date": "02/05/2024 20:48:40",
                              "content": "<p>Thanks!</p>\n<p>I'll be honest with you, I don't know what that z-score normalization would be. I'm working with 2D, specifically with kidney_1_dense. I still have a lot to learn to reach that level.</p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 2638858,
                              "author_name": "fernandoaesteves",
                              "author_url": "",
                              "post_date": "02/06/2024 14:24:13",
                              "content": "<p>Friend, I just realized that it wasn't the model that was causing empty masks, it was the binarization that was leaving all the masks empty. I switched to the Otsu method, but it's making the masks dirty. As a result, the score went from 0 to 0.044. I'm going to try the binarization method that Bhavya Dhingra mentioned in his notebook \"Clean Code 📚| Weighted Ensemble [Inference]\". Wish me luck! And thank you very much!</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2638887,
                                  "author_name": "mohammad2012191",
                                  "author_url": "",
                                  "post_date": "02/06/2024 14:36:38",
                                  "content": "<p>Good luck!<br>\nAnyway I am suggesting (if there is enough time) to revise the normalization again.<br>\nAssume a dataset with only one feature. This feature ranges from 0 to 100 in the training data, and say you used min_max_scaler to make it from 0 to 1. Then in test data you have the same feature ranges from 20 to 80, and you used min_max_scaler (not the one fitted on training data) to make the range between 0 and 1.<br>\nWhen the model get fitted into the training data, the 0s, for example, of training data are NOT the same as the 0s of the testing data. So, it is expected to produce quite bad results (unless you are lucky enough to have training data and testing data both between 0 and 100). That's why it is a good practice to scale test data with same scaler you used in training data. <br>\nIn the public nbs they fitted scalers for each dataset independently and that worked somehow lol, but this behaviour is risky and can lead to quite bad results in test data.</p>",
                                  "votes": null,
                                  "replies": []
                                },
                                {
                                  "id": 2640343,
                                  "author_name": "minhsienweng",
                                  "author_url": "",
                                  "post_date": "02/06/2024 20:50:11",
                                  "content": "<p>What I did is to compare training and inference notebooks for function inconsistencies, especially scaling and processing functions. </p>\n<p>Always used training notebook functions to ensure consistency across both notebooks.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2636352,
      "author_name": "minhsienweng",
      "author_url": "",
      "post_date": "02/05/2024 03:16:04",
      "content": "<p>I found the memory usage was a potential cause of slowdown. By clearing memory after each image inference, I was able to improve execution speed and avoid submission timeouts.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2633577": "My submission completed after a few minutes, but the scoring process took longer and encountered an error. Has anyone else faced a similar issue?",
    "2633851": "Hi. The one that completed in few minuts was executed at example test folder with 6 example files. The one that took longer was executed on real test folder. Without more information and If the example didn't fail I'd say that there is something wrong with id. If that's the case don't try generate any id, read them from the files in test and use them sorted as you read them.",
    "2634718": "I've carefully reviewed the file IDs and resolved the identified issue. I am runing the code and submit again to chck if it is fixed. However, I suspect there might be some problem elsewhere. \n\nYour suggestion was helpful—thank you!",
    "2635878": "First of all, thank you very much! Through your comments and tips here in the discussion, I finally managed to submit after 2 months of a lot of effort.\n\nSecondly, my score is currently at zero, and there are 2 days left. People mentioned that this may be due to the threshold of binarization, saying that the threshold should be high, close to 255. However, I am still scoring zero. Do you have any tips for me to check in my algorithm to really score in this competition?\n\nI'm kind of desperate because I only have 5 submissions per day, and I've already used today's, so there are only 5 left until the end.",
    "2636071": "https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/471055",
    "2636082": "You're welcome! I'll check this out!!",
    "2636184": "The zero score could be due to the different functions used in training and inference.\n\nThanks for the suggestion. I thoroughly reviewed my training and inference notebooks, and I've resolved the issue.",
    "2636352": "I found the memory usage was a potential cause of slowdown. By clearing memory after each image inference, I was able to improve execution speed and avoid submission timeouts.",
    "2637659": "I found the cause of the score 0.\n\nThe model I trained works in the training notebook, generating predicted masks for unknown images.\n\nHowever, when I use the same model in the inference notebook, the model produces empty masks, even with the dataset it was trained on.\n\nNow I don't know if the issue is with the model loading or if it's related to the preprocessing.",
    "2637684": "Check the normalization part. If the normalization values are different between train and test then I think this behavior might be expected. I had something similar and when I just used Z-score of train data for both training and inference the problem has been resolved.",
    "2637744": "Thanks!\n\nI'll be honest with you, I don't know what that z-score normalization would be. I'm working with 2D, specifically with kidney_1_dense. I still have a lot to learn to reach that level.",
    "2638858": "Friend, I just realized that it wasn't the model that was causing empty masks, it was the binarization that was leaving all the masks empty. I switched to the Otsu method, but it's making the masks dirty. As a result, the score went from 0 to 0.044. I'm going to try the binarization method that Bhavya Dhingra mentioned in his notebook \"Clean Code 📚| Weighted Ensemble [Inference]\". Wish me luck! And thank you very much!",
    "2638887": "Good luck!\nAnyway I am suggesting (if there is enough time) to revise the normalization again.\nAssume a dataset with only one feature. This feature ranges from 0 to 100 in the training data, and say you used min_max_scaler to make it from 0 to 1. Then in test data you have the same feature ranges from 20 to 80, and you used min_max_scaler (not the one fitted on training data) to make the range between 0 and 1.\nWhen the model get fitted into the training data, the 0s, for example, of training data are NOT the same as the 0s of the testing data. So, it is expected to produce quite bad results (unless you are lucky enough to have training data and testing data both between 0 and 100). That's why it is a good practice to scale test data with same scaler you used in training data. \nIn the public nbs they fitted scalers for each dataset independently and that worked somehow lol, but this behaviour is risky and can lead to quite bad results in test data.",
    "2640343": "What I did is to compare training and inference notebooks for function inconsistencies, especially scaling and processing functions. \n\nAlways used training notebook functions to ensure consistency across both notebooks."
  },
  "source": "meta"
}