{
  "id": 203354,
  "title": "Global How-To for these types of competitions",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/203354",
  "author_name": "",
  "post_date": "2020-12-14T23:08:44.967879700Z",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Kagglers,</p>\n<p>NONE OF THIS IS OFFICIAL KAGGLE POSITION. Just my thoughts.</p>\n<p>There have been a number of these image/code competitions recently. Let's see if we can answer the most common questions in one place.</p>\n<p>Please add appropriate links to similar posts in other recent contests.</p>\n<p>Starters:<br>\nFor questions not answered in this competition discussion area, look at recent similar competitions:<br>\n  RSNA Pulmonary Embolism <br>\n  Melanoma - similar contest<br>\n  Cassava Leaf Disease - 5 different targets. Possibly the most similar</p>\n<p>One difference - this competition gives you 18,000 annotations, while the above competitions had only binary labels. The current HuBMAP - Hacking the Kidney has annotations, and might be similar to this competitions.</p>\n<p>Some Notes about submission method:</p>\n<p>You may TRAIN using any resources you have, including Kaggle, Colab or your own resources. No limitations on TRAIN resources.</p>\n<p>To submit a result, you must do the inference in a committed Kaggle notebook.</p>\n<p>Usually, you load your model weights into a notebook.</p>\n<p>The committed notebook cannot use:<br>\n  Internet (because you could leak the hidden test data)<br>\n  TPU (because you need Internet for TPU)</p>\n<p>If your model isn't built into Kaggle and you need it to do your inference, you need to copy the github/wheel files into a dataset. Then you can load it in the committed notebook. See competitions listed above for examples.</p>\n<p>The above competitions give various overviews of the commit process. The important highlights are:</p>\n<p>The public test dataset is 3582 images.<br>\nThe hidden dataset is approximately 14,000 images. Probably the public test (3582) plus approximately 10,500 more.<br>\nMake sure your committed notebook can run in the 9 hours allotted and within available memory. As a test, you can run it against about half the train dataset.</p>\n<p>Committing a notebook for submission is a two step process;<br>\n  Save and Run the notebook - it must complete and produce a submission file. But this is just a \"dry run\". The submission file IS NOT the file you must submit. You can take shortcuts and not process the entire public test data at this step to save time.</p>\n<p>My Submissions - find the result of the Save/Run you just did. Submit that Notebook/result.</p>\n<p>Kaggle will then re-run your notebook against the hidden dataset (which includes the public test data and additional hidden test data). It will take four times longer than processing the public test dataset.</p>\n<p>I believe based on prior contests that the public leaderboard is based on the public test data. The real private leaderboard score is based on the other 75%. So, during testing, you can probably only inference on the public test data and you will still get a public leaderboard score. However, make sure your code is set up to run the entire test dataset in committed mode, especially given time/memory constraints.</p>\n<p>There is no way to see the hidden test data. Therefore there is no way to:</p>\n<p>prebuild a full test submission file and submit it.</p>\n<p>pre-process the full test data. You will need to do this in the committed notebook. You probably don't have the disk space to preprocess it all at once, so do it in smaller batches.</p>\n<p>Ensemble multiple submission.csv files unless you can do all the inference/ensemble in one notebook in the 9 hours allowed.</p>\n<p>sample_submission.csv, test images, test_tfrecords are all replaced in the committed notebook.</p>\n<p>If you have an invalid submission file, it is very difficult to debug. That is because any information could contribute to a leak of the hidden test data:</p>\n<p>For your submission file - <br>\n  must have all rows in the hidden sample_submission.csv file.</p>\n<p>when you write your file, use \"index=False\". Make sure your labels are correct.</p>\n<p>you must capture any errors. For any code that could fail, use try/except blocks.</p>\n<p>be very careful with resources. Enable your GPU if you need it. Watch memory and disk space. Concatenating PANDAS arrays takes up a lot of memory. Use other techniques.</p>\n<p>I think this competition will be much harder than it looks at first. As a radiologist, I look forward to significant advancement in this area.</p>\n<p>If you got this far, one final comment. While this is a friendly competition, it is still a competition. Anybody who scores well should have to put in some real work. Please make sure anything you post contributes to learning and doesn't just let Kagglers leaderboard-climb with a high-scoring notebook. Especially in the last few weeks, but really for the entire competition. Most of us have taken \"time off\" from competitions because of a high scoring notebook scattering the leaderboard.</p>\n<p>Best of luck to all.</p>\n<p>-Rich</p>\n<p>“Intelligence is the ability to learn from your mistakes. Wisdom is the ability to learn from the mistakes of others.” – Anonymous.</p>",
  "messages": [
    {
      "id": "1112815",
      "postDate": "12/14/2020 23:08:44",
      "content": "<p>Kagglers,</p>\n<p>NONE OF THIS IS OFFICIAL KAGGLE POSITION. Just my thoughts.</p>\n<p>There have been a number of these image/code competitions recently. Let's see if we can answer the most common questions in one place.</p>\n<p>Please add appropriate links to similar posts in other recent contests.</p>\n<p>Starters:<br>\nFor questions not answered in this competition discussion area, look at recent similar competitions:<br>\n  RSNA Pulmonary Embolism <br>\n  Melanoma - similar contest<br>\n  Cassava Leaf Disease - 5 different targets. Possibly the most similar</p>\n<p>One difference - this competition gives you 18,000 annotations, while the above competitions had only binary labels. The current HuBMAP - Hacking the Kidney has annotations, and might be similar to this competitions.</p>\n<p>Some Notes about submission method:</p>\n<p>You may TRAIN using any resources you have, including Kaggle, Colab or your own resources. No limitations on TRAIN resources.</p>\n<p>To submit a result, you must do the inference in a committed Kaggle notebook.</p>\n<p>Usually, you load your model weights into a notebook.</p>\n<p>The committed notebook cannot use:<br>\n  Internet (because you could leak the hidden test data)<br>\n  TPU (because you need Internet for TPU)</p>\n<p>If your model isn't built into Kaggle and you need it to do your inference, you need to copy the github/wheel files into a dataset. Then you can load it in the committed notebook. See competitions listed above for examples.</p>\n<p>The above competitions give various overviews of the commit process. The important highlights are:</p>\n<p>The public test dataset is 3582 images.<br>\nThe hidden dataset is approximately 14,000 images. Probably the public test (3582) plus approximately 10,500 more.<br>\nMake sure your committed notebook can run in the 9 hours allotted and within available memory. As a test, you can run it against about half the train dataset.</p>\n<p>Committing a notebook for submission is a two step process;<br>\n  Save and Run the notebook - it must complete and produce a submission file. But this is just a \"dry run\". The submission file IS NOT the file you must submit. You can take shortcuts and not process the entire public test data at this step to save time.</p>\n<p>My Submissions - find the result of the Save/Run you just did. Submit that Notebook/result.</p>\n<p>Kaggle will then re-run your notebook against the hidden dataset (which includes the public test data and additional hidden test data). It will take four times longer than processing the public test dataset.</p>\n<p>I believe based on prior contests that the public leaderboard is based on the public test data. The real private leaderboard score is based on the other 75%. So, during testing, you can probably only inference on the public test data and you will still get a public leaderboard score. However, make sure your code is set up to run the entire test dataset in committed mode, especially given time/memory constraints.</p>\n<p>There is no way to see the hidden test data. Therefore there is no way to:</p>\n<p>prebuild a full test submission file and submit it.</p>\n<p>pre-process the full test data. You will need to do this in the committed notebook. You probably don't have the disk space to preprocess it all at once, so do it in smaller batches.</p>\n<p>Ensemble multiple submission.csv files unless you can do all the inference/ensemble in one notebook in the 9 hours allowed.</p>\n<p>sample_submission.csv, test images, test_tfrecords are all replaced in the committed notebook.</p>\n<p>If you have an invalid submission file, it is very difficult to debug. That is because any information could contribute to a leak of the hidden test data:</p>\n<p>For your submission file - <br>\n  must have all rows in the hidden sample_submission.csv file.</p>\n<p>when you write your file, use \"index=False\". Make sure your labels are correct.</p>\n<p>you must capture any errors. For any code that could fail, use try/except blocks.</p>\n<p>be very careful with resources. Enable your GPU if you need it. Watch memory and disk space. Concatenating PANDAS arrays takes up a lot of memory. Use other techniques.</p>\n<p>I think this competition will be much harder than it looks at first. As a radiologist, I look forward to significant advancement in this area.</p>\n<p>If you got this far, one final comment. While this is a friendly competition, it is still a competition. Anybody who scores well should have to put in some real work. Please make sure anything you post contributes to learning and doesn't just let Kagglers leaderboard-climb with a high-scoring notebook. Especially in the last few weeks, but really for the entire competition. Most of us have taken \"time off\" from competitions because of a high scoring notebook scattering the leaderboard.</p>\n<p>Best of luck to all.</p>\n<p>-Rich</p>\n<p>“Intelligence is the ability to learn from your mistakes. Wisdom is the ability to learn from the mistakes of others.” – Anonymous.</p>",
      "rawMarkdown": "Kagglers,\n\nNONE OF THIS IS OFFICIAL KAGGLE POSITION. Just my thoughts.\n\nThere have been a number of these image/code competitions recently. Let's see if we can answer the most common questions in one place.\n\nPlease add appropriate links to similar posts in other recent contests.\n\nStarters:\nFor questions not answered in this competition discussion area, look at recent similar competitions:\n  RSNA Pulmonary Embolism \n  Melanoma - similar contest\n  Cassava Leaf Disease - 5 different targets. Possibly the most similar\n\nOne difference - this competition gives you 18,000 annotations, while the above competitions had only binary labels. The current HuBMAP - Hacking the Kidney has annotations, and might be similar to this competitions.\n\nSome Notes about submission method:\n\nYou may TRAIN using any resources you have, including Kaggle, Colab or your own resources. No limitations on TRAIN resources.\n\nTo submit a result, you must do the inference in a committed Kaggle notebook.\n\nUsually, you load your model weights into a notebook.\n\nThe committed notebook cannot use:\n  Internet (because you could leak the hidden test data)\n  TPU (because you need Internet for TPU)\n\nIf your model isn't built into Kaggle and you need it to do your inference, you need to copy the github/wheel files into a dataset. Then you can load it in the committed notebook. See competitions listed above for examples.\n\nThe above competitions give various overviews of the commit process. The important highlights are:\n\nThe public test dataset is 3582 images.\nThe hidden dataset is approximately 14,000 images. Probably the public test (3582) plus approximately 10,500 more.\nMake sure your committed notebook can run in the 9 hours allotted and within available memory. As a test, you can run it against about half the train dataset.\n\nCommitting a notebook for submission is a two step process;\n  Save and Run the notebook - it must complete and produce a submission file. But this is just a \"dry run\". The submission file IS NOT the file you must submit. You can take shortcuts and not process the entire public test data at this step to save time.\n\n My Submissions - find the result of the Save/Run you just did. Submit that Notebook/result.\n\nKaggle will then re-run your notebook against the hidden dataset (which includes the public test data and additional hidden test data). It will take four times longer than processing the public test dataset.\n\nI believe based on prior contests that the public leaderboard is based on the public test data. The real private leaderboard score is based on the other 75%. So, during testing, you can probably only inference on the public test data and you will still get a public leaderboard score. However, make sure your code is set up to run the entire test dataset in committed mode, especially given time/memory constraints.\n\nThere is no way to see the hidden test data. Therefore there is no way to:\n\n  prebuild a full test submission file and submit it.\n\n  pre-process the full test data. You will need to do this in the committed notebook. You probably don't have the disk space to preprocess it all at once, so do it in smaller batches.\n\n  Ensemble multiple submission.csv files unless you can do all the inference/ensemble in one notebook in the 9 hours allowed.\n\n  sample_submission.csv, test images, test_tfrecords are all replaced in the committed notebook.\n\nIf you have an invalid submission file, it is very difficult to debug. That is because any information could contribute to a leak of the hidden test data:\n\nFor your submission file - \n  must have all rows in the hidden sample_submission.csv file.\n\n  when you write your file, use \"index=False\". Make sure your labels are correct.\n\n  you must capture any errors. For any code that could fail, use try/except blocks.\n\n  be very careful with resources. Enable your GPU if you need it. Watch memory and disk space. Concatenating PANDAS arrays takes up a lot of memory. Use other techniques.\n\nI think this competition will be much harder than it looks at first. As a radiologist, I look forward to significant advancement in this area.\n\nIf you got this far, one final comment. While this is a friendly competition, it is still a competition. Anybody who scores well should have to put in some real work. Please make sure anything you post contributes to learning and doesn't just let Kagglers leaderboard-climb with a high-scoring notebook. Especially in the last few weeks, but really for the entire competition. Most of us have taken \"time off\" from competitions because of a high scoring notebook scattering the leaderboard.\n\nBest of luck to all.\n\n-Rich\n\n“Intelligence is the ability to learn from your mistakes. Wisdom is the ability to learn from the mistakes of others.” – Anonymous.",
      "votes": null
    },
    {
      "id": "1112854",
      "postDate": "12/15/2020 00:15:26",
      "content": "<p>My last competition was a code competition, but with tabular data. It was too slow to make submissions, i hope this doesn't be true:</p>\n<blockquote>\n  <p>It will take four times longer than processing the public test dataset</p>\n</blockquote>",
      "rawMarkdown": "My last competition was a code competition, but with tabular data. It was too slow to make submissions, i hope this doesn't be true:\n\n>  It will take four times longer than processing the public test dataset",
      "votes": null
    },
    {
      "id": "1112857",
      "postDate": "12/15/2020 00:25:12",
      "content": "<p>If you ran out of time without having images to process, look closely at your references and indexes in pandas dataframes. Simplify your code until you isolate what is too slow. Sometimes the bottleneck is where you least expect it.</p>",
      "rawMarkdown": "If you ran out of time without having images to process, look closely at your references and indexes in pandas dataframes. Simplify your code until you isolate what is too slow. Sometimes the bottleneck is where you least expect it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1112854,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "12/15/2020 00:15:26",
      "content": "<p>My last competition was a code competition, but with tabular data. It was too slow to make submissions, i hope this doesn't be true:</p>\n<blockquote>\n  <p>It will take four times longer than processing the public test dataset</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1112857,
          "author_name": "richardepstein",
          "author_url": "",
          "post_date": "12/15/2020 00:25:12",
          "content": "<p>If you ran out of time without having images to process, look closely at your references and indexes in pandas dataframes. Simplify your code until you isolate what is too slow. Sometimes the bottleneck is where you least expect it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1112815": "Kagglers,\n\nNONE OF THIS IS OFFICIAL KAGGLE POSITION. Just my thoughts.\n\nThere have been a number of these image/code competitions recently. Let's see if we can answer the most common questions in one place.\n\nPlease add appropriate links to similar posts in other recent contests.\n\nStarters:\nFor questions not answered in this competition discussion area, look at recent similar competitions:\n  RSNA Pulmonary Embolism \n  Melanoma - similar contest\n  Cassava Leaf Disease - 5 different targets. Possibly the most similar\n\nOne difference - this competition gives you 18,000 annotations, while the above competitions had only binary labels. The current HuBMAP - Hacking the Kidney has annotations, and might be similar to this competitions.\n\nSome Notes about submission method:\n\nYou may TRAIN using any resources you have, including Kaggle, Colab or your own resources. No limitations on TRAIN resources.\n\nTo submit a result, you must do the inference in a committed Kaggle notebook.\n\nUsually, you load your model weights into a notebook.\n\nThe committed notebook cannot use:\n  Internet (because you could leak the hidden test data)\n  TPU (because you need Internet for TPU)\n\nIf your model isn't built into Kaggle and you need it to do your inference, you need to copy the github/wheel files into a dataset. Then you can load it in the committed notebook. See competitions listed above for examples.\n\nThe above competitions give various overviews of the commit process. The important highlights are:\n\nThe public test dataset is 3582 images.\nThe hidden dataset is approximately 14,000 images. Probably the public test (3582) plus approximately 10,500 more.\nMake sure your committed notebook can run in the 9 hours allotted and within available memory. As a test, you can run it against about half the train dataset.\n\nCommitting a notebook for submission is a two step process;\n  Save and Run the notebook - it must complete and produce a submission file. But this is just a \"dry run\". The submission file IS NOT the file you must submit. You can take shortcuts and not process the entire public test data at this step to save time.\n\n My Submissions - find the result of the Save/Run you just did. Submit that Notebook/result.\n\nKaggle will then re-run your notebook against the hidden dataset (which includes the public test data and additional hidden test data). It will take four times longer than processing the public test dataset.\n\nI believe based on prior contests that the public leaderboard is based on the public test data. The real private leaderboard score is based on the other 75%. So, during testing, you can probably only inference on the public test data and you will still get a public leaderboard score. However, make sure your code is set up to run the entire test dataset in committed mode, especially given time/memory constraints.\n\nThere is no way to see the hidden test data. Therefore there is no way to:\n\n  prebuild a full test submission file and submit it.\n\n  pre-process the full test data. You will need to do this in the committed notebook. You probably don't have the disk space to preprocess it all at once, so do it in smaller batches.\n\n  Ensemble multiple submission.csv files unless you can do all the inference/ensemble in one notebook in the 9 hours allowed.\n\n  sample_submission.csv, test images, test_tfrecords are all replaced in the committed notebook.\n\nIf you have an invalid submission file, it is very difficult to debug. That is because any information could contribute to a leak of the hidden test data:\n\nFor your submission file - \n  must have all rows in the hidden sample_submission.csv file.\n\n  when you write your file, use \"index=False\". Make sure your labels are correct.\n\n  you must capture any errors. For any code that could fail, use try/except blocks.\n\n  be very careful with resources. Enable your GPU if you need it. Watch memory and disk space. Concatenating PANDAS arrays takes up a lot of memory. Use other techniques.\n\nI think this competition will be much harder than it looks at first. As a radiologist, I look forward to significant advancement in this area.\n\nIf you got this far, one final comment. While this is a friendly competition, it is still a competition. Anybody who scores well should have to put in some real work. Please make sure anything you post contributes to learning and doesn't just let Kagglers leaderboard-climb with a high-scoring notebook. Especially in the last few weeks, but really for the entire competition. Most of us have taken \"time off\" from competitions because of a high scoring notebook scattering the leaderboard.\n\nBest of luck to all.\n\n-Rich\n\n“Intelligence is the ability to learn from your mistakes. Wisdom is the ability to learn from the mistakes of others.” – Anonymous.",
    "1112854": "My last competition was a code competition, but with tabular data. It was too slow to make submissions, i hope this doesn't be true:\n\n>  It will take four times longer than processing the public test dataset",
    "1112857": "If you ran out of time without having images to process, look closely at your references and indexes in pandas dataframes. Simplify your code until you isolate what is too slow. Sometimes the bottleneck is where you least expect it."
  },
  "source": "meta"
}