{
  "id": 475053,
  "title": "20th to 1000th+ Place =)",
  "url": "/competitions/blood-vessel-segmentation/discussion/475053",
  "author_name": "Theo Viel",
  "post_date": "2024-02-07T00:39:12.026000",
  "votes": 28,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Despite scoring okay on public and running successfully, my pipeline gave an astonishing 0.01 private score :)</p>\n<p>Code is public if someone finds an obvious flaw:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/theoviel/sennet-inference-theo\" target=\"_blank\">https://www.kaggle.com/code/theoviel/sennet-inference-theo</a></p>\n</blockquote>\n<p>Even the first versions don't score well on private and they're straight forward. I think the issue is when I create the submission file.</p>\n<p>My best guess so far is that the data structure changed in private (naming is not 0 -&gt; n_frames ?) but that makes no sense. </p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I'd appreciate some help on this</p>\n<p><strong>Update:</strong> My submission would've ranked 20th (private 0.616). Thanks to everyone who helped solving the issue. <br>\nAt least I did not miss a gold medal. </p>",
  "messages": [
    {
      "id": 2640486,
      "postDate": "2024-02-07T00:39:12.027Z",
      "content": "<p>Despite scoring okay on public and running successfully, my pipeline gave an astonishing 0.01 private score :)</p>\n<p>Code is public if someone finds an obvious flaw:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/theoviel/sennet-inference-theo\" target=\"_blank\">https://www.kaggle.com/code/theoviel/sennet-inference-theo</a></p>\n</blockquote>\n<p>Even the first versions don't score well on private and they're straight forward. I think the issue is when I create the submission file.</p>\n<p>My best guess so far is that the data structure changed in private (naming is not 0 -&gt; n_frames ?) but that makes no sense. </p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I'd appreciate some help on this</p>\n<p><strong>Update:</strong> My submission would've ranked 20th (private 0.616). Thanks to everyone who helped solving the issue. <br>\nAt least I did not miss a gold medal. </p>",
      "rawMarkdown": "Despite scoring okay on public and running successfully, my pipeline gave an astonishing 0.01 private score :)\n\nCode is public if someone finds an obvious flaw:\n> https://www.kaggle.com/code/theoviel/sennet-inference-theo\n\nEven the first versions don't score well on private and they're straight forward. I think the issue is when I create the submission file.\n\nMy best guess so far is that the data structure changed in private (naming is not 0 -> n_frames ?) but that makes no sense. \n\n@ryanholbrook I'd appreciate some help on this\n\n**Update:** My submission would've ranked 20th (private 0.616). Thanks to everyone who helped solving the issue. \nAt least I did not miss a gold medal. ",
      "votes": 28
    },
    {
      "id": 2640529,
      "postDate": "2024-02-07T01:14:26.680Z",
      "content": "<p>Pretty sure the issue lies here</p>\n<pre><code> i, p  (preds):\n        sub.append(\n            pd.DataFrame({\n                : [],\n                : [rle_encode(p)],\n            })\n        )\n</code></pre>\n<p>Since a kidney doesn't have to start at 0 slice index (since it's expensive to get dense annotations, just a small continuous part could be annotated and used for evaluation), that might be the issue. In other words, the safer approach would be to extract the actual image names from the dataset.</p>",
      "rawMarkdown": "Pretty sure the issue lies here\n\n```python\nfor i, p in enumerate(preds):\n        sub.append(\n            pd.DataFrame({\n                'id': [f\"{kidney}_{i:04d}\"],\n                'rle': [rle_encode(p)],\n            })\n        )\n\n```\nSince a kidney doesn't have to start at 0 slice index (since it's expensive to get dense annotations, just a small continuous part could be annotated and used for evaluation), that might be the issue. In other words, the safer approach would be to extract the actual image names from the dataset.",
      "votes": 6,
      "replies": [
        {
          "id": 2640576,
          "postDate": "2024-02-07T02:04:07.553Z",
          "content": "<p>You are right. Fixed the issue and now I am getting scores. I'm resubmitting my final subs =)</p>",
          "rawMarkdown": "You are right. Fixed the issue and now I am getting scores. I'm resubmitting my final subs =)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2642122,
      "postDate": "2024-02-07T23:37:35.153Z",
      "content": "<p>I made the same mistake. My true score was 0.597…</p>",
      "rawMarkdown": "I made the same mistake. My true score was 0.597...",
      "votes": 1
    },
    {
      "id": 2642857,
      "postDate": "2024-02-08T13:20:22.657Z",
      "content": "<p>helpfulllll</p>",
      "rawMarkdown": "helpfulllll"
    },
    {
      "id": 2641116,
      "postDate": "2024-02-07T09:45:17.817Z",
      "content": "<p>One of my subs got 0.82 in public and 0.007 in private. The difference between this sub and the others (which got reasonable scores in private) is that I just used bigger inference image size😅.</p>",
      "rawMarkdown": "One of my subs got 0.82 in public and 0.007 in private. The difference between this sub and the others (which got reasonable scores in private) is that I just used bigger inference image size😅."
    },
    {
      "id": 2640555,
      "postDate": "2024-02-07T01:29:28.727Z",
      "content": "<p>I believe I had the same issue, definitely feeling for you right now.</p>",
      "rawMarkdown": "I believe I had the same issue, definitely feeling for you right now."
    },
    {
      "id": 2640521,
      "postDate": "2024-02-07T01:07:36.477Z",
      "content": "<p>Our PB result was an ensemble score of five models, one of which scored 0.003.</p>",
      "rawMarkdown": "Our PB result was an ensemble score of five models, one of which scored 0.003."
    },
    {
      "id": 2640489,
      "postDate": "2024-02-07T00:44:49.073Z",
      "content": "<blockquote>\n  <p>My best guess so far is that the data structure changed in private (naming is not 0 -&gt; n_frames ?) but that makes no sense.</p>\n</blockquote>\n<p>They mentionned that the private LB was a subset of the kidney_6 like kidney_3_dense was. If you look at kidney_3_dense ids, it starts from 496. Sad to see nevertheless, if it is what you pointed out, I hope kaggle can make your subs count.</p>",
      "rawMarkdown": "> My best guess so far is that the data structure changed in private (naming is not 0 -> n_frames ?) but that makes no sense.\n\nThey mentionned that the private LB was a subset of the kidney_6 like kidney_3_dense was. If you look at kidney_3_dense ids, it starts from 496. Sad to see nevertheless, if it is what you pointed out, I hope kaggle can make your subs count.",
      "replies": [
        {
          "id": 2640522,
          "postDate": "2024-02-07T01:07:48.213Z",
          "content": "<p>Really ? That would make sense but data description says \"Continuous 3D part of a whole human kidney\" for both public and private </p>",
          "rawMarkdown": "Really ? That would make sense but data description says \"Continuous 3D part of a whole human kidney\" for both public and private ",
          "replies": [
            {
              "id": 2640533,
              "postDate": "2024-02-07T01:15:10.963Z",
              "content": "<p>Exactly <br>\nA part of a kidney, not a whole kidney </p>",
              "rawMarkdown": "Exactly \nA part of a kidney, not a whole kidney ",
              "votes": 2
            },
            {
              "id": 2640567,
              "postDate": "2024-02-07T01:53:15.620Z",
              "content": "<p>Ooooh I see. Guess I read too fast and interpreted that they annotated the entire kidney …<br>\nWill resubmit my models and see how it goes. </p>\n<p>Still very misleading to have the frames start at a random number. The provided sample data starts at 0 even though they are not the first frames of the kidney.</p>",
              "rawMarkdown": "Ooooh I see. Guess I read too fast and interpreted that they annotated the entire kidney ...\nWill resubmit my models and see how it goes. \n\nStill very misleading to have the frames start at a random number. The provided sample data starts at 0 even though they are not the first frames of the kidney.",
              "votes": 2
            },
            {
              "id": 2641598,
              "postDate": "2024-02-07T15:12:00.727Z",
              "content": "<blockquote>\n  <p>Still very misleading to have the frames start at a random number. The provided sample data starts at 0 even though they are not the first frames of the kidney.</p>\n</blockquote>\n<p>you have a good point about the sample submission having keys that actually wouldn't work in a submission xd </p>\n<blockquote>\n  <p>Will resubmit my models and see how it goes.</p>\n</blockquote>\n<p>Out of curiosity, what did they score ? </p>",
              "rawMarkdown": "> Still very misleading to have the frames start at a random number. The provided sample data starts at 0 even though they are not the first frames of the kidney.\n\nyou have a good point about the sample submission having keys that actually wouldn't work in a submission xd \n\n> Will resubmit my models and see how it goes.\n\nOut of curiosity, what did they score ? "
            },
            {
              "id": 2641624,
              "postDate": "2024-02-07T15:21:40.163Z",
              "content": "<p>My best selected sub scored 0.616 which is 20th or 21st</p>",
              "rawMarkdown": "My best selected sub scored 0.616 which is 20th or 21st",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2640493,
      "postDate": "2024-02-07T00:49:41.047Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2640490,
      "postDate": "2024-02-07T00:45:11.307Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2640529,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2024-02-07T01:14:26.680000",
      "content": "<p>Pretty sure the issue lies here</p>\n<pre><code> i, p  (preds):\n        sub.append(\n            pd.DataFrame({\n                : [],\n                : [rle_encode(p)],\n            })\n        )\n</code></pre>\n<p>Since a kidney doesn't have to start at 0 slice index (since it's expensive to get dense annotations, just a small continuous part could be annotated and used for evaluation), that might be the issue. In other words, the safer approach would be to extract the actual image names from the dataset.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2640576,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2024-02-07T02:04:07.553000",
          "content": "<p>You are right. Fixed the issue and now I am getting scores. I'm resubmitting my final subs =)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2642122,
      "author_name": "D.Imanishi",
      "author_url": "",
      "post_date": "2024-02-07T23:37:35.153000",
      "content": "<p>I made the same mistake. My true score was 0.597…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2642857,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-08T13:20:22.657000",
      "content": "<p>helpfulllll</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2641116,
      "author_name": "Mohamed Eltayeb",
      "author_url": "",
      "post_date": "2024-02-07T09:45:17.817000",
      "content": "<p>One of my subs got 0.82 in public and 0.007 in private. The difference between this sub and the others (which got reasonable scores in private) is that I just used bigger inference image size😅.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2640555,
      "author_name": "N1234567",
      "author_url": "",
      "post_date": "2024-02-07T01:29:28.727000",
      "content": "<p>I believe I had the same issue, definitely feeling for you right now.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2640521,
      "author_name": "ynhuhu",
      "author_url": "",
      "post_date": "2024-02-07T01:07:36.477000",
      "content": "<p>Our PB result was an ensemble score of five models, one of which scored 0.003.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2640489,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2024-02-07T00:44:49.073000",
      "content": "<blockquote>\n  <p>My best guess so far is that the data structure changed in private (naming is not 0 -&gt; n_frames ?) but that makes no sense.</p>\n</blockquote>\n<p>They mentionned that the private LB was a subset of the kidney_6 like kidney_3_dense was. If you look at kidney_3_dense ids, it starts from 496. Sad to see nevertheless, if it is what you pointed out, I hope kaggle can make your subs count.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2640522,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2024-02-07T01:07:48.213000",
          "content": "<p>Really ? That would make sense but data description says \"Continuous 3D part of a whole human kidney\" for both public and private </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2640533,
              "author_name": "Igor Krashenyi",
              "author_url": "",
              "post_date": "2024-02-07T01:15:10.963000",
              "content": "<p>Exactly <br>\nA part of a kidney, not a whole kidney </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2640567,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2024-02-07T01:53:15.620000",
              "content": "<p>Ooooh I see. Guess I read too fast and interpreted that they annotated the entire kidney …<br>\nWill resubmit my models and see how it goes. </p>\n<p>Still very misleading to have the frames start at a random number. The provided sample data starts at 0 even though they are not the first frames of the kidney.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2641598,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2024-02-07T15:12:00.727000",
              "content": "<blockquote>\n  <p>Still very misleading to have the frames start at a random number. The provided sample data starts at 0 even though they are not the first frames of the kidney.</p>\n</blockquote>\n<p>you have a good point about the sample submission having keys that actually wouldn't work in a submission xd </p>\n<blockquote>\n  <p>Will resubmit my models and see how it goes.</p>\n</blockquote>\n<p>Out of curiosity, what did they score ? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2641624,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2024-02-07T15:21:40.163000",
              "content": "<p>My best selected sub scored 0.616 which is 20th or 21st</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2640493,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-07T00:49:41.047000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2640490,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-07T00:45:11.307000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2640486": "Despite scoring okay on public and running successfully, my pipeline gave an astonishing 0.01 private score :)\n\nCode is public if someone finds an obvious flaw:\n> https://www.kaggle.com/code/theoviel/sennet-inference-theo\n\nEven the first versions don't score well on private and they're straight forward. I think the issue is when I create the submission file.\n\nMy best guess so far is that the data structure changed in private (naming is not 0 -> n_frames ?) but that makes no sense. \n\n@ryanholbrook I'd appreciate some help on this\n\n**Update:** My submission would've ranked 20th (private 0.616). Thanks to everyone who helped solving the issue. \nAt least I did not miss a gold medal. ",
    "2640529": "Pretty sure the issue lies here\n\n```python\nfor i, p in enumerate(preds):\n        sub.append(\n            pd.DataFrame({\n                'id': [f\"{kidney}_{i:04d}\"],\n                'rle': [rle_encode(p)],\n            })\n        )\n\n```\nSince a kidney doesn't have to start at 0 slice index (since it's expensive to get dense annotations, just a small continuous part could be annotated and used for evaluation), that might be the issue. In other words, the safer approach would be to extract the actual image names from the dataset.",
    "2642122": "I made the same mistake. My true score was 0.597...",
    "2642857": "helpfulllll",
    "2641116": "One of my subs got 0.82 in public and 0.007 in private. The difference between this sub and the others (which got reasonable scores in private) is that I just used bigger inference image size😅.",
    "2640555": "I believe I had the same issue, definitely feeling for you right now.",
    "2640521": "Our PB result was an ensemble score of five models, one of which scored 0.003.",
    "2640489": "> My best guess so far is that the data structure changed in private (naming is not 0 -> n_frames ?) but that makes no sense.\n\nThey mentionned that the private LB was a subset of the kidney_6 like kidney_3_dense was. If you look at kidney_3_dense ids, it starts from 496. Sad to see nevertheless, if it is what you pointed out, I hope kaggle can make your subs count.",
    "2640493": "",
    "2640490": ""
  }
}