{
  "id": 202005,
  "title": "iafoss nb with metadata and higher CV / HARDCODED due to submission error",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/202005",
  "author_name": "Issagali Konysbayev [dsmlkz]",
  "post_date": "2020-12-07T20:05:06.324000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>**TL;DR: **</p>\n<ol>\n<li>The metadata, though small, might help improve the model.</li>\n<li>Beware - the private test set metadata could cause 'Submission Scoring Error' (still waiting for response to this <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403\" target=\"_blank\">question</a> in 'Discussions'.</li>\n</ol>\n<p><strong>Update:</strong> kaggle informed <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403\" target=\"_blank\">here</a> that  'HuBMAP-20-dataset_information.csv' is not available for private test set. So I removed the notebooks and models dataset from the public domain as they are not useful anymore.</p>\n<hr>\n<p>In <a href=\"https://www.kaggle.com/isakev/hubmap-hardcode-iafoss-nb-with-metadata\" target=\"_blank\">the notebook</a>, the model architecture from the <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter\" target=\"_blank\">notebook</a> was changed to accept the metadata from 'HuBMAP-20-dataset_information.csv' file and trained exactly as in that notebook.</p>\n<p>The validation score was by 0.0010  higher than in the original <a href=\"https://www.kaggle.com/iafoos\" target=\"_blank\">@iafoos</a> notebook, but PL score was lower. The reason might be that the models used in the <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub\" target=\"_blank\">submission notebook</a> were not the ones trained in the original public notebook or due to differences in test and training sets.</p>\n<p>The metadata used are only those for the public data. When trying to use the private test set metadata from the file 'HuBMAP-20-dataset_information.csv' I stuck repeatedly with 'Submission Error' either due to a bug in my code or some inconsistency of private test set metadata - this (version 1) is my first successful submission out of 16. So, I used the mean and the mode of the public metadata when predicting private test set.</p>\n<p>The CV score improvement due to the addition of metadata might be insignificant since there are only few samples corresponding to raw images and the better CV score could be due to noise.</p>",
  "messages": [
    {
      "id": 1105357,
      "postDate": "2020-12-07T20:05:06.323Z",
      "content": "<p>**TL;DR: **</p>\n<ol>\n<li>The metadata, though small, might help improve the model.</li>\n<li>Beware - the private test set metadata could cause 'Submission Scoring Error' (still waiting for response to this <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403\" target=\"_blank\">question</a> in 'Discussions'.</li>\n</ol>\n<p><strong>Update:</strong> kaggle informed <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403\" target=\"_blank\">here</a> that  'HuBMAP-20-dataset_information.csv' is not available for private test set. So I removed the notebooks and models dataset from the public domain as they are not useful anymore.</p>\n<hr>\n<p>In <a href=\"https://www.kaggle.com/isakev/hubmap-hardcode-iafoss-nb-with-metadata\" target=\"_blank\">the notebook</a>, the model architecture from the <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter\" target=\"_blank\">notebook</a> was changed to accept the metadata from 'HuBMAP-20-dataset_information.csv' file and trained exactly as in that notebook.</p>\n<p>The validation score was by 0.0010  higher than in the original <a href=\"https://www.kaggle.com/iafoos\" target=\"_blank\">@iafoos</a> notebook, but PL score was lower. The reason might be that the models used in the <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub\" target=\"_blank\">submission notebook</a> were not the ones trained in the original public notebook or due to differences in test and training sets.</p>\n<p>The metadata used are only those for the public data. When trying to use the private test set metadata from the file 'HuBMAP-20-dataset_information.csv' I stuck repeatedly with 'Submission Error' either due to a bug in my code or some inconsistency of private test set metadata - this (version 1) is my first successful submission out of 16. So, I used the mean and the mode of the public metadata when predicting private test set.</p>\n<p>The CV score improvement due to the addition of metadata might be insignificant since there are only few samples corresponding to raw images and the better CV score could be due to noise.</p>",
      "rawMarkdown": "**TL;DR: **\n1. The metadata, though small, might help improve the model.\n2. Beware - the private test set metadata could cause 'Submission Scoring Error' (still waiting for response to this [question](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403) in 'Discussions'.\n\n**Update:** kaggle informed [here]( https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403) that  'HuBMAP-20-dataset_information.csv' is not available for private test set. So I removed the notebooks and models dataset from the public domain as they are not useful anymore.\n\n-------------------\n\nIn [the notebook](https://www.kaggle.com/isakev/hubmap-hardcode-iafoss-nb-with-metadata), the model architecture from the @iafoss [notebook](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter) was changed to accept the metadata from 'HuBMAP-20-dataset_information.csv' file and trained exactly as in that notebook.\n\nThe validation score was by 0.0010  higher than in the original @iafoos notebook, but PL score was lower. The reason might be that the models used in the @iafoss [submission notebook](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub) were not the ones trained in the original public notebook or due to differences in test and training sets.\n\nThe metadata used are only those for the public data. When trying to use the private test set metadata from the file 'HuBMAP-20-dataset_information.csv' I stuck repeatedly with 'Submission Error' either due to a bug in my code or some inconsistency of private test set metadata - this (version 1) is my first successful submission out of 16. So, I used the mean and the mode of the public metadata when predicting private test set.\n\nThe CV score improvement due to the addition of metadata might be insignificant since there are only few samples corresponding to raw images and the better CV score could be due to noise.",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1105357": "**TL;DR: **\n1. The metadata, though small, might help improve the model.\n2. Beware - the private test set metadata could cause 'Submission Scoring Error' (still waiting for response to this [question](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403) in 'Discussions'.\n\n**Update:** kaggle informed [here]( https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201506#1104403) that  'HuBMAP-20-dataset_information.csv' is not available for private test set. So I removed the notebooks and models dataset from the public domain as they are not useful anymore.\n\n-------------------\n\nIn [the notebook](https://www.kaggle.com/isakev/hubmap-hardcode-iafoss-nb-with-metadata), the model architecture from the @iafoss [notebook](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter) was changed to accept the metadata from 'HuBMAP-20-dataset_information.csv' file and trained exactly as in that notebook.\n\nThe validation score was by 0.0010  higher than in the original @iafoos notebook, but PL score was lower. The reason might be that the models used in the @iafoss [submission notebook](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub) were not the ones trained in the original public notebook or due to differences in test and training sets.\n\nThe metadata used are only those for the public data. When trying to use the private test set metadata from the file 'HuBMAP-20-dataset_information.csv' I stuck repeatedly with 'Submission Error' either due to a bug in my code or some inconsistency of private test set metadata - this (version 1) is my first successful submission out of 16. So, I used the mean and the mode of the public metadata when predicting private test set.\n\nThe CV score improvement due to the addition of metadata might be insignificant since there are only few samples corresponding to raw images and the better CV score could be due to noise."
  }
}