{
  "id": 228654,
  "title": "Probing the Private Test Image Size",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/228654",
  "author_name": "",
  "post_date": "2021-03-25T17:27:19.617638800Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<blockquote>\n  <p>EDIT : It appears <code>HuBMAP-20-dataset_information.csv</code> is not updated for the private dataset, this code doesn't work :(</p>\n</blockquote>\n<p>As a lot of you have noticed, there is one image in the private test set that is bigger than those in the training &amp; public sets.</p>\n<p>This results in countless submission errors, and debugging your code gets really painful. Having an idea of the size of the image could definitely help.</p>\n<p>The following code should allow to check if the maximum image size is higher than <code>ref_size</code> :</p>\n<pre><code>import pandas as pd\n\ndf_info = pd.read_csv(\"../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv\")\ndf_info['nb_pixels'] = df_info['width_pixels'] * df_info['height_pixels']\nmax_ = df_info['nb_pixels'].max()\n\n# Max on train set : 2025172800\n# we want to check if the biggest test image is bigger than :\nref_size= 3000000000  \n\nif ref_size &gt; max_:\n    df_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\n    df_submit.to_csv('submission.csv')\n</code></pre>\n<p>The biggest image in the private having more than <code>ref_size</code> pixels will result in a scoring error, whereas otherwise the notebook should work fine. Please let me know if something is wrong with my approach though.</p>\n<p><code>ref_size</code> can be adjusted iteratively to have a better idea of the image size :) </p>\n<p>Unfortunately I'm out of subs for today so I can't try the idea yet. If anybody feels like trying it today please share your results, i'll share mine as soon as possible otherwise.</p>",
  "messages": [
    {
      "id": "1252435",
      "postDate": "03/25/2021 17:27:19",
      "content": "<blockquote>\n  <p>EDIT : It appears <code>HuBMAP-20-dataset_information.csv</code> is not updated for the private dataset, this code doesn't work :(</p>\n</blockquote>\n<p>As a lot of you have noticed, there is one image in the private test set that is bigger than those in the training &amp; public sets.</p>\n<p>This results in countless submission errors, and debugging your code gets really painful. Having an idea of the size of the image could definitely help.</p>\n<p>The following code should allow to check if the maximum image size is higher than <code>ref_size</code> :</p>\n<pre><code>import pandas as pd\n\ndf_info = pd.read_csv(\"../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv\")\ndf_info['nb_pixels'] = df_info['width_pixels'] * df_info['height_pixels']\nmax_ = df_info['nb_pixels'].max()\n\n# Max on train set : 2025172800\n# we want to check if the biggest test image is bigger than :\nref_size= 3000000000  \n\nif ref_size &gt; max_:\n    df_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\n    df_submit.to_csv('submission.csv')\n</code></pre>\n<p>The biggest image in the private having more than <code>ref_size</code> pixels will result in a scoring error, whereas otherwise the notebook should work fine. Please let me know if something is wrong with my approach though.</p>\n<p><code>ref_size</code> can be adjusted iteratively to have a better idea of the image size :) </p>\n<p>Unfortunately I'm out of subs for today so I can't try the idea yet. If anybody feels like trying it today please share your results, i'll share mine as soon as possible otherwise.</p>",
      "rawMarkdown": "> EDIT : It appears `HuBMAP-20-dataset_information.csv` is not updated for the private dataset, this code doesn't work :(\n\nAs a lot of you have noticed, there is one image in the private test set that is bigger than those in the training & public sets.\n\nThis results in countless submission errors, and debugging your code gets really painful. Having an idea of the size of the image could definitely help.\n\nThe following code should allow to check if the maximum image size is higher than `ref_size` :\n\n```\nimport pandas as pd\n\ndf_info = pd.read_csv(\"../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv\")\ndf_info['nb_pixels'] = df_info['width_pixels'] * df_info['height_pixels']\nmax_ = df_info['nb_pixels'].max()\n\n# Max on train set : 2025172800\n# we want to check if the biggest test image is bigger than :\nref_size= 3000000000  \n\nif ref_size > max_:\n    df_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\n    df_submit.to_csv('submission.csv')\n```\n\nThe biggest image in the private having more than `ref_size` pixels will result in a scoring error, whereas otherwise the notebook should work fine. Please let me know if something is wrong with my approach though.\n\n`ref_size` can be adjusted iteratively to have a better idea of the image size :) \n\nUnfortunately I'm out of subs for today so I can't try the idea yet. If anybody feels like trying it today please share your results, i'll share mine as soon as possible otherwise.",
      "votes": null
    },
    {
      "id": "1252590",
      "postDate": "03/25/2021 20:08:32",
      "content": "<p>I've already got the full collection of possible errors on Kaggle. This time no error!<br>\nBut are you sure <code>../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv</code> contains width_pixels/height_pixels for private data? Maybe opening each image and do the same would work.</p>",
      "rawMarkdown": "I've already got the full collection of possible errors on Kaggle. This time no error!\nBut are you sure `../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv` contains width_pixels/height_pixels for private data? Maybe opening each image and do the same would work.",
      "votes": null
    },
    {
      "id": "1252595",
      "postDate": "03/25/2021 20:16:59",
      "content": "<p>I assumed this file would be updated with the private LB but yes it might not be the case. I'll have to check that!</p>",
      "rawMarkdown": "I assumed this file would be updated with the private LB but yes it might not be the case. I'll have to check that!",
      "votes": null
    },
    {
      "id": "1252599",
      "postDate": "03/25/2021 20:23:26",
      "content": "<p>That would explain why I keep running into submission errors though.</p>",
      "rawMarkdown": "That would explain why I keep running into submission errors though.",
      "votes": null
    },
    {
      "id": "1252630",
      "postDate": "03/25/2021 21:27:40",
      "content": "<p>We would need a new category in Kaggle: \"Failure expert/master/grandmaster\" 😄<br>\nI'm also getting an error I'm not able to tackle since days.</p>",
      "rawMarkdown": "We would need a new category in Kaggle: \"Failure expert/master/grandmaster\" 😄\nI'm also getting an error I'm not able to tackle since days.",
      "votes": null
    },
    {
      "id": "1252643",
      "postDate": "03/25/2021 21:48:18",
      "content": "<p><img src=\"https://nsa40.casimages.com/img/2021/03/25/210325105706636012.png\" alt=\"\"></p>\n<p>Where can I apply for failure GM ?</p>",
      "rawMarkdown": "![](https://nsa40.casimages.com/img/2021/03/25/210325105706636012.png)\n\nWhere can I apply for failure GM ?",
      "votes": null
    },
    {
      "id": "1252656",
      "postDate": "03/25/2021 22:18:26",
      "content": "<p>You cannot be failure GM with only single error multiple times. You need the full collection at least.</p>",
      "rawMarkdown": "You cannot be failure GM with only single error multiple times. You need the full collection at least.",
      "votes": null
    },
    {
      "id": "1252658",
      "postDate": "03/25/2021 22:28:31",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506</a></p>\n<p>🤗 sorry I forgot about this…</p>",
      "rawMarkdown": "theoviel https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506\n\n🤗 sorry I forgot about this...",
      "votes": null
    },
    {
      "id": "1268789",
      "postDate": "04/09/2021 18:52:12",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <br>\nI tried with this code but all result into sub csv not found</p>\n<pre><code>ref=39960* 50680\n        if h*w &gt; int(1.1*ref):\n            test_df.to_csv('submission.csv')\n        elif h*w &gt; int(1.2*ref):\n            None\n        elif h*w &gt; int(1.3)*ref:\n            None\n\n        elif h*w &gt;int(1.5* ref):\n             #test_df.to_csv('submission.csv')\n            None\n        else:\n            print('None')\n</code></pre>",
      "rawMarkdown": "theoviel \nI tried with this code but all result into sub csv not found\n\n```\nref=39960* 50680\n        if h*w > int(1.1*ref):\n            test_df.to_csv('submission.csv')\n        elif h*w > int(1.2*ref):\n            None\n        elif h*w > int(1.3)*ref:\n            None\n            \n        elif h*w >int(1.5* ref):\n             #test_df.to_csv('submission.csv')\n            None\n        else:\n            print('None')\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1252590,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "03/25/2021 20:08:32",
      "content": "<p>I've already got the full collection of possible errors on Kaggle. This time no error!<br>\nBut are you sure <code>../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv</code> contains width_pixels/height_pixels for private data? Maybe opening each image and do the same would work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1252595,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/25/2021 20:16:59",
          "content": "<p>I assumed this file would be updated with the private LB but yes it might not be the case. I'll have to check that!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252599,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/25/2021 20:23:26",
          "content": "<p>That would explain why I keep running into submission errors though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252630,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/25/2021 21:27:40",
          "content": "<p>We would need a new category in Kaggle: \"Failure expert/master/grandmaster\" 😄<br>\nI'm also getting an error I'm not able to tackle since days.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252643,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/25/2021 21:48:18",
          "content": "<p><img src=\"https://nsa40.casimages.com/img/2021/03/25/210325105706636012.png\" alt=\"\"></p>\n<p>Where can I apply for failure GM ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252656,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/25/2021 22:18:26",
          "content": "<p>You cannot be failure GM with only single error multiple times. You need the full collection at least.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252658,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "03/25/2021 22:28:31",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506</a></p>\n<p>🤗 sorry I forgot about this…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1268789,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "04/09/2021 18:52:12",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <br>\nI tried with this code but all result into sub csv not found</p>\n<pre><code>ref=39960* 50680\n        if h*w &gt; int(1.1*ref):\n            test_df.to_csv('submission.csv')\n        elif h*w &gt; int(1.2*ref):\n            None\n        elif h*w &gt; int(1.3)*ref:\n            None\n\n        elif h*w &gt;int(1.5* ref):\n             #test_df.to_csv('submission.csv')\n            None\n        else:\n            print('None')\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1252435": "> EDIT : It appears `HuBMAP-20-dataset_information.csv` is not updated for the private dataset, this code doesn't work :(\n\nAs a lot of you have noticed, there is one image in the private test set that is bigger than those in the training & public sets.\n\nThis results in countless submission errors, and debugging your code gets really painful. Having an idea of the size of the image could definitely help.\n\nThe following code should allow to check if the maximum image size is higher than `ref_size` :\n\n```\nimport pandas as pd\n\ndf_info = pd.read_csv(\"../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv\")\ndf_info['nb_pixels'] = df_info['width_pixels'] * df_info['height_pixels']\nmax_ = df_info['nb_pixels'].max()\n\n# Max on train set : 2025172800\n# we want to check if the biggest test image is bigger than :\nref_size= 3000000000  \n\nif ref_size > max_:\n    df_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\n    df_submit.to_csv('submission.csv')\n```\n\nThe biggest image in the private having more than `ref_size` pixels will result in a scoring error, whereas otherwise the notebook should work fine. Please let me know if something is wrong with my approach though.\n\n`ref_size` can be adjusted iteratively to have a better idea of the image size :) \n\nUnfortunately I'm out of subs for today so I can't try the idea yet. If anybody feels like trying it today please share your results, i'll share mine as soon as possible otherwise.",
    "1252590": "I've already got the full collection of possible errors on Kaggle. This time no error!\nBut are you sure `../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv` contains width_pixels/height_pixels for private data? Maybe opening each image and do the same would work.",
    "1252595": "I assumed this file would be updated with the private LB but yes it might not be the case. I'll have to check that!",
    "1252599": "That would explain why I keep running into submission errors though.",
    "1252630": "We would need a new category in Kaggle: \"Failure expert/master/grandmaster\" 😄\nI'm also getting an error I'm not able to tackle since days.",
    "1252643": "![](https://nsa40.casimages.com/img/2021/03/25/210325105706636012.png)\n\nWhere can I apply for failure GM ?",
    "1252656": "You cannot be failure GM with only single error multiple times. You need the full collection at least.",
    "1252658": "theoviel https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506\n\n🤗 sorry I forgot about this...",
    "1268789": "theoviel \nI tried with this code but all result into sub csv not found\n\n```\nref=39960* 50680\n        if h*w > int(1.1*ref):\n            test_df.to_csv('submission.csv')\n        elif h*w > int(1.2*ref):\n            None\n        elif h*w > int(1.3)*ref:\n            None\n            \n        elif h*w >int(1.5* ref):\n             #test_df.to_csv('submission.csv')\n            None\n        else:\n            print('None')\n```"
  },
  "source": "meta"
}