{
  "id": 154297,
  "title": "New to Machine Learning or Kaggle?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154297",
  "author_name": "Julia Elliott",
  "post_date": "2020-05-28T00:18:31.577000",
  "votes": 17,
  "comment_count": 62,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\">how to enter a competition using Kaggle Notebooks</a>.</p>\n\n<p>Ready to dive into this competition? Review the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\">Overview Description</a> and start to work with the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/data\">Data</a>!</p>",
  "messages": [
    {
      "id": 864318,
      "postDate": "2020-05-28T00:18:31.577Z",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\">how to enter a competition using Kaggle Notebooks</a>.</p>\n\n<p>Ready to dive into this competition? Review the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\">Overview Description</a> and start to work with the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/data\">Data</a>!</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).\n\nReady to dive into this competition? Review the [Overview Description](https://www.kaggle.com/c/siim-isic-melanoma-classification/overview) and start to work with the [Data](https://www.kaggle.com/c/siim-isic-melanoma-classification/data)!",
      "votes": 15
    },
    {
      "id": 876951,
      "postDate": "2020-06-07T07:39:51.757Z",
      "content": "<p>I am new to kaggle but have done many projects in DL.</p>",
      "rawMarkdown": "I am new to kaggle but have done many projects in DL.",
      "votes": 6
    },
    {
      "id": 964778,
      "postDate": "2020-08-10T06:53:02.043Z",
      "content": "<p>Not new but trying to do kaggle at least one hour each day. Need Motivations. Thanks.</p>",
      "rawMarkdown": "Not new but trying to do kaggle at least one hour each day. Need Motivations. Thanks.",
      "votes": 1
    },
    {
      "id": 952564,
      "postDate": "2020-07-31T03:50:14.233Z",
      "content": "<p>This is my first cv competition. Got to learn a lot of stuff.</p>",
      "rawMarkdown": "This is my first cv competition. Got to learn a lot of stuff.",
      "votes": 1,
      "replies": [
        {
          "id": 953369,
          "postDate": "2020-07-31T18:20:37.460Z",
          "content": "<p>Same here. There is seriously a lot more that I need to learn.</p>",
          "rawMarkdown": "Same here. There is seriously a lot more that I need to learn."
        }
      ]
    },
    {
      "id": 945810,
      "postDate": "2020-07-26T06:33:33.530Z",
      "content": "<p>HI. I'm a beginner of competitions using images  and deep learning. I want to know if there are any public notebooks that kindly teaches the basics of modules and how it works?</p>",
      "rawMarkdown": "HI. I'm a beginner of competitions using images  and deep learning. I want to know if there are any public notebooks that kindly teaches the basics of modules and how it works?",
      "votes": 1,
      "replies": [
        {
          "id": 946152,
          "postDate": "2020-07-26T11:49:05Z",
          "content": "<p>For PyTorch I think this both are great for begginers:\n- <a href=\"https://www.kaggle.com/ankitjha/practical-deep-learning-using-pytorch\">https://www.kaggle.com/ankitjha/practical-deep-learning-using-pytorch</a>\n- <a href=\"https://www.kaggle.com/mratsim/starting-kit-for-pytorch-deep-learning\">https://www.kaggle.com/mratsim/starting-kit-for-pytorch-deep-learning</a></p>",
          "rawMarkdown": "For PyTorch I think this both are great for begginers:\n- https://www.kaggle.com/ankitjha/practical-deep-learning-using-pytorch\n- https://www.kaggle.com/mratsim/starting-kit-for-pytorch-deep-learning\n",
          "votes": 2
        },
        {
          "id": 946229,
          "postDate": "2020-07-26T12:56:50.303Z",
          "content": "<p>Thanks a lot for your information!</p>",
          "rawMarkdown": "Thanks a lot for your information!",
          "votes": 1
        }
      ]
    },
    {
      "id": 935375,
      "postDate": "2020-07-19T10:33:49.040Z",
      "content": "<p>Hi                        </p>",
      "rawMarkdown": "Hi                        \n",
      "votes": 1
    },
    {
      "id": 930499,
      "postDate": "2020-07-15T14:05:04.173Z",
      "content": "<p>hi <a href=\"/juliaelliott\">@juliaelliott</a> \nI have a question regarding public LB.\nI submitted several submissions with a 0.957 score in public LB, but got a different rank.\nI went from 27th to 22nd to 17th rank. Is there a hidden 4th post decimal value we don't see or is the private lb score used to rank the values with the same public LB score?</p>\n\n<p>best Roman</p>",
      "rawMarkdown": "hi @juliaelliott \nI have a question regarding public LB.\nI submitted several submissions with a 0.957 score in public LB, but got a different rank.\nI went from 27th to 22nd to 17th rank. Is there a hidden 4th post decimal value we don't see or is the private lb score used to rank the values with the same public LB score?\n\nbest Roman",
      "votes": 1,
      "replies": [
        {
          "id": 930658,
          "postDate": "2020-07-15T16:17:38.113Z",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> The public LB ranking will always be based on the public test subset, and based on the full precision of your score. We often only reveal a smaller number (3 or 4) of decimal places at the outset to protect against leaderboard probing, and then consider increasing precision as warranted. Therefore, yes, the ranking of multiple scores which appear to have the same score out to 3 decimal places will be based on hidden decimal values. We've gone ahead and updated the precision to 4 decimal places to help differentiate the playing field a bit.</p>",
          "rawMarkdown": "@romanweilguny The public LB ranking will always be based on the public test subset, and based on the full precision of your score. We often only reveal a smaller number (3 or 4) of decimal places at the outset to protect against leaderboard probing, and then consider increasing precision as warranted. Therefore, yes, the ranking of multiple scores which appear to have the same score out to 3 decimal places will be based on hidden decimal values. We've gone ahead and updated the precision to 4 decimal places to help differentiate the playing field a bit.",
          "votes": 2
        },
        {
          "id": 930708,
          "postDate": "2020-07-15T16:58:15.923Z",
          "content": "<p>thx - that was quick :)</p>",
          "rawMarkdown": "thx - that was quick :)"
        },
        {
          "id": 938595,
          "postDate": "2020-07-21T16:12:18.633Z",
          "content": "<p>thanks for  clarifying. i also wondered what the process was</p>",
          "rawMarkdown": "thanks for  clarifying. i also wondered what the process was"
        },
        {
          "id": 953367,
          "postDate": "2020-07-31T18:19:31.137Z",
          "content": "<p>Thanks for the answer. Even I was wondering the same thing. :)</p>",
          "rawMarkdown": "Thanks for the answer. Even I was wondering the same thing. :)"
        }
      ]
    },
    {
      "id": 869395,
      "postDate": "2020-06-01T01:51:48.583Z",
      "content": "<p>How can I download the data on a server? Is there a link I can wget from</p>",
      "rawMarkdown": "How can I download the data on a server? Is there a link I can wget from",
      "votes": 1,
      "replies": [
        {
          "id": 877944,
          "postDate": "2020-06-08T06:22:12.780Z",
          "content": "<p>you can't use wget , you have to use kaggle api to download the dataset through code.</p>",
          "rawMarkdown": "you can't use wget , you have to use kaggle api to download the dataset through code.",
          "votes": 5
        }
      ]
    },
    {
      "id": 864789,
      "postDate": "2020-05-28T07:29:06.710Z",
      "content": "<p>Hello,\nI'm new to kaggle! \nRegards,\nLittisha</p>",
      "rawMarkdown": "Hello,\nI'm new to kaggle! \nRegards,\nLittisha",
      "votes": 1
    },
    {
      "id": 935866,
      "postDate": "2020-07-19T18:04:47.347Z",
      "content": "<p>New to kaggle. I am confused with why the GPU is not getting used when I am training my model. </p>\n\n<p>I have the GPU session enabled, but the GPU is not working as evident by the GPU bar in the top right of the notebook.</p>\n\n<p>How do I enable GPU if I already am in a GPU session? Thanks in advance.</p>",
      "rawMarkdown": "New to kaggle. I am confused with why the GPU is not getting used when I am training my model. \n\nI have the GPU session enabled, but the GPU is not working as evident by the GPU bar in the top right of the notebook.\n\nHow do I enable GPU if I already am in a GPU session? Thanks in advance.",
      "votes": 2,
      "replies": [
        {
          "id": 938894,
          "postDate": "2020-07-21T20:47:28.660Z",
          "content": "<p>I think you are missing the device part in code where you select de GPU. If you using Pytorch it should be something like:\ndevice = torch.device(\"cuda\")\nmodel.to(device)</p>",
          "rawMarkdown": "I think you are missing the device part in code where you select de GPU. If you using Pytorch it should be something like:\ndevice = torch.device(\"cuda\")\nmodel.to(device)",
          "votes": 1
        },
        {
          "id": 982889,
          "postDate": "2020-08-23T18:54:53.417Z",
          "content": "<p>Hello, I am not an expert myself but started to use GPU and TPU for competition. here are some notebooks I have made to present how I managed to do it:<br>\n<a href=\"https://www.kaggle.com/foucardm/tutorial-top-15-mnist-digit-recognizer-on-tpu\" target=\"_blank\">GPU and TPU for TensorFlow</a> and this one with <a href=\"https://www.kaggle.com/foucardm/end-to-end-mnist-introduction-to-pytoch\" target=\"_blank\">PyTorch</a></p>",
          "rawMarkdown": "Hello, I am not an expert myself but started to use GPU and TPU for competition. here are some notebooks I have made to present how I managed to do it:\n[GPU and TPU for TensorFlow](https://www.kaggle.com/foucardm/tutorial-top-15-mnist-digit-recognizer-on-tpu) and this one with [PyTorch](https://www.kaggle.com/foucardm/end-to-end-mnist-introduction-to-pytoch)"
        }
      ]
    },
    {
      "id": 869718,
      "postDate": "2020-06-01T08:31:13.057Z",
      "content": "<p>Hi, I am a recent joinee too. I have started a new series on Kaggle named <strong>\"Beginners' Mistakes\"</strong> over <a href=\"https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned\">here</a>. Do check it out  I hope it helps you. \nI would appreciate your views, perspectives, and support to continue it :)</p>",
      "rawMarkdown": "Hi, I am a recent joinee too. I have started a new series on Kaggle named **\"Beginners' Mistakes\"** over [here](https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned). Do check it out  I hope it helps you. \nI would appreciate your views, perspectives, and support to continue it :)",
      "votes": 2
    },
    {
      "id": 866300,
      "postDate": "2020-05-29T09:25:08.193Z",
      "content": "<p>Hi im new to kaggle.</p>",
      "rawMarkdown": "Hi im new to kaggle."
    },
    {
      "id": 983164,
      "postDate": "2020-08-24T04:59:53.580Z",
      "content": "<p>Hello here !</p>\n<p>I'm new too. I began by learning with the micro-courses.</p>",
      "rawMarkdown": "Hello here !\n\nI'm new too. I began by learning with the micro-courses."
    },
    {
      "id": 982653,
      "postDate": "2020-08-23T14:54:42.807Z",
      "content": "<p>Hi everyone! I'm relatively new to kaggle and I want to know how to find the images present in the HAM10000 dataset in this dataset for an upcoming project. Can someone please explain to me how they would go about doing that? I would really appreciate it.</p>",
      "rawMarkdown": "Hi everyone! I'm relatively new to kaggle and I want to know how to find the images present in the HAM10000 dataset in this dataset for an upcoming project. Can someone please explain to me how they would go about doing that? I would really appreciate it."
    },
    {
      "id": 968052,
      "postDate": "2020-08-12T17:02:37.063Z",
      "content": "<p>Hello, newbie here. Just started a few python programming courses after using Matlab for engineering during several years</p>",
      "rawMarkdown": "Hello, newbie here. Just started a few python programming courses after using Matlab for engineering during several years"
    },
    {
      "id": 952648,
      "postDate": "2020-07-31T05:40:26.183Z",
      "content": "<p>Hi </p>",
      "rawMarkdown": "Hi "
    },
    {
      "id": 935449,
      "postDate": "2020-07-19T11:54:15.983Z",
      "content": "<p>Hi everyone, quite a newbie student here on my first kaggle project. <br>\nI have a question regaring the evaluation of a submission:<br>\nI've seen different submission file with different ranges of values, some go up to 0.8 and some have a maximal value of 0.2. How is the threshold then set for the classification, I mean if my maximal value of the submission file is 0.6 and someone else's is 0.2? Is the entire column normalized and then sliced in the middle?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi everyone, quite a newbie student here on my first kaggle project. \nI have a question regaring the evaluation of a submission:\nI've seen different submission file with different ranges of values, some go up to 0.8 and some have a maximal value of 0.2. How is the threshold then set for the classification, I mean if my maximal value of the submission file is 0.6 and someone else's is 0.2? Is the entire column normalized and then sliced in the middle?\n\nThanks"
    },
    {
      "id": 932702,
      "postDate": "2020-07-17T08:09:38.983Z",
      "content": "<p>Hi</p>\n\n<p>I just ran into a problem with a self-created dataset I wanted to use in my kernel.</p>\n\n<p>I get an error message: Unexpected response from the service. Response: {'errors': ['Google Cloud SDK must be authorized before copying private datasets.'], 'error': {'code': 9}, 'wasSuccessful': False}.</p>\n\n<p>when accessing the dataset, which has been imported successfully.</p>\n\n<p>GCS_PATH_PSEUDO =  KaggleDatasets().get_gcs_path('pseudolabeled')</p>\n\n<p>Any idea?</p>",
      "rawMarkdown": "Hi\n\nI just ran into a problem with a self-created dataset I wanted to use in my kernel.\n\nI get an error message: Unexpected response from the service. Response: {'errors': ['Google Cloud SDK must be authorized before copying private datasets.'], 'error': {'code': 9}, 'wasSuccessful': False}.\n\nwhen accessing the dataset, which has been imported successfully.\n\nGCS_PATH_PSEUDO =  KaggleDatasets().get_gcs_path('pseudolabeled')\n\nAny idea?",
      "replies": [
        {
          "id": 933447,
          "postDate": "2020-07-17T17:50:26.387Z",
          "content": "<p>Hi Roman, were you able to try the instructions here <a href=\"https://www.kaggle.com/product-feedback/163416\">https://www.kaggle.com/product-feedback/163416</a> ?</p>",
          "rawMarkdown": "Hi Roman, were you able to try the instructions here https://www.kaggle.com/product-feedback/163416 ?"
        },
        {
          "id": 933485,
          "postDate": "2020-07-17T18:11:42.940Z",
          "content": "<p>Hi </p>\n\n<p>thanks - I have tried this in the meanwhile and it works.\nWhat puzzles me a bit - this worked for me without the new code before - but anyway... thx</p>",
          "rawMarkdown": "Hi \n\nthanks - I have tried this in the meanwhile and it works.\nWhat puzzles me a bit - this worked for me without the new code before - but anyway... thx"
        }
      ]
    },
    {
      "id": 916815,
      "postDate": "2020-07-06T04:07:35.630Z",
      "content": "<p>newbie here </p>",
      "rawMarkdown": "newbie here "
    },
    {
      "id": 916802,
      "postDate": "2020-07-06T03:46:46.693Z",
      "content": "<p>Hello</p>",
      "rawMarkdown": "Hello"
    },
    {
      "id": 916227,
      "postDate": "2020-07-05T13:39:06.657Z",
      "content": "<p>Hello everyone. I'm new to Kaggle. I'd like to know if I can use any programming language to solve this, since the submission is only the predictions. </p>",
      "rawMarkdown": "Hello everyone. I'm new to Kaggle. I'd like to know if I can use any programming language to solve this, since the submission is only the predictions. ",
      "replies": [
        {
          "id": 916267,
          "postDate": "2020-07-05T14:18:38.490Z",
          "content": "<p>I think the only limitation would be if you used a programming language that did not have a free version. </p>",
          "rawMarkdown": "I think the only limitation would be if you used a programming language that did not have a free version. ",
          "votes": 1
        },
        {
          "id": 916324,
          "postDate": "2020-07-05T14:53:18.380Z",
          "content": "<p>So this means it's forbidden to use a proprietary software (not open source) to generate the predictions?</p>",
          "rawMarkdown": "So this means it's forbidden to use a proprietary software (not open source) to generate the predictions?"
        },
        {
          "id": 916447,
          "postDate": "2020-07-05T16:50:05.310Z",
          "content": "<p>I found this  in the rules. So I guess \"generally commercially available software...without undue expense\", allows some commercial software products. If you are using commercial software, you might want to verify with the Kaggle team that the specific software you want to use is allowed.</p>\n\n<p>(a) deliver to the Competition Sponsor the final model's software code as used to generate the winning Submission and associated documentation. The delivered software code should follow these documentation guidelines, must be capable of generating the winning Submission, and contain a description of resources required to build and/or run the executable code successfully. If the final model’s software code includes generally commercially available software that is not owned by you, but that can be procured by the Competition Sponsor without undue expense, then instead of delivering the code for that software to the Competition Sponsor, you must identify that software, method for procuring it, and any parameters or other information necessary to replicate the winning Submission;</p>",
          "rawMarkdown": "I found this  in the rules. So I guess \"generally commercially available software...without undue expense\", allows some commercial software products. If you are using commercial software, you might want to verify with the Kaggle team that the specific software you want to use is allowed.\n\n(a) deliver to the Competition Sponsor the final model's software code as used to generate the winning Submission and associated documentation. The delivered software code should follow these documentation guidelines, must be capable of generating the winning Submission, and contain a description of resources required to build and/or run the executable code successfully. If the final model’s software code includes generally commercially available software that is not owned by you, but that can be procured by the Competition Sponsor without undue expense, then instead of delivering the code for that software to the Competition Sponsor, you must identify that software, method for procuring it, and any parameters or other information necessary to replicate the winning Submission;",
          "votes": 1
        }
      ]
    },
    {
      "id": 914407,
      "postDate": "2020-07-03T20:56:29.957Z",
      "content": "<p>hi, how can i load the datasets in google colab ?</p>",
      "rawMarkdown": "hi, how can i load the datasets in google colab ?"
    },
    {
      "id": 910692,
      "postDate": "2020-07-01T10:02:00.943Z",
      "content": "<p>Hello to everyone. I am new to Kaggle and this almost the first time. It seems to me that the test.csv file will be used to create predictions that will be submitted to the competition. The test file does not have 'target' column, which is of course the column to be predicted. My question is: How do you evaluate your models? Do you have to split again train.csv into train and test sets?</p>",
      "rawMarkdown": "Hello to everyone. I am new to Kaggle and this almost the first time. It seems to me that the test.csv file will be used to create predictions that will be submitted to the competition. The test file does not have 'target' column, which is of course the column to be predicted. My question is: How do you evaluate your models? Do you have to split again train.csv into train and test sets?",
      "replies": [
        {
          "id": 910708,
          "postDate": "2020-07-01T10:14:58.713Z",
          "content": "<p>I'm new too, but I think I can help you, please give me details of your problem and i'll try to help you!</p>",
          "rawMarkdown": "I'm new too, but I think I can help you, please give me details of your problem and i'll try to help you!"
        },
        {
          "id": 910968,
          "postDate": "2020-07-01T13:24:21.637Z",
          "content": "<p>You are on the right track. The train data is typically split into \"train\" and \"validation\" subsets. You train on the \"train\" data, and validate with the \"validation\" subset. The use of \"train\" in two different meanings can be confusing. Some datasets will come with pre-defined \"train\", \"validate\" and \"test\" datasets, so everybody can be consistent on what is the \"validate\" data. However, splitting out the validate data is part of the art of machine learning.</p>\n\n<p>It is common to split out 20% of the data for the validation subset. Often you will make five splits, each with a different 20% of the data as the validation and the other 80% of the data for training. Then you average or otherwise combine your 5 sets of results to get a submission. So if you put your data into 5 buckets, you would have:</p>\n\n<p>fold 1 - T T T T V\nfold 2 - T T T V T\nfold 3 - T T V T T\nfold 4 - T V T T T \nfold 5 - V T T T T</p>\n\n<p>Where you train on the T's in each fold and validate on the V's.</p>\n\n<p>For this contest you might want to make sure a given patient is not split between the train and validate datasets, since that could be a source of leakage.</p>\n\n<p>Separately, if you use External data, you might want the External data only in the Train dataset and not in the Validate dataset. Otherwise, your validate dataset results will be artificially improved. So, I split out my validation data before I add in the External data.</p>\n\n<p>Best of luck</p>\n\n<p>-Rich</p>",
          "rawMarkdown": "You are on the right track. The train data is typically split into \"train\" and \"validation\" subsets. You train on the \"train\" data, and validate with the \"validation\" subset. The use of \"train\" in two different meanings can be confusing. Some datasets will come with pre-defined \"train\", \"validate\" and \"test\" datasets, so everybody can be consistent on what is the \"validate\" data. However, splitting out the validate data is part of the art of machine learning.\n\nIt is common to split out 20% of the data for the validation subset. Often you will make five splits, each with a different 20% of the data as the validation and the other 80% of the data for training. Then you average or otherwise combine your 5 sets of results to get a submission. So if you put your data into 5 buckets, you would have:\n\nfold 1 - T T T T V\nfold 2 - T T T V T\nfold 3 - T T V T T\nfold 4 - T V T T T \nfold 5 - V T T T T\n\nWhere you train on the T's in each fold and validate on the V's.\n\nFor this contest you might want to make sure a given patient is not split between the train and validate datasets, since that could be a source of leakage.\n\nSeparately, if you use External data, you might want the External data only in the Train dataset and not in the Validate dataset. Otherwise, your validate dataset results will be artificially improved. So, I split out my validation data before I add in the External data.\n\nBest of luck\n\n-Rich\n\n ",
          "votes": 1
        }
      ]
    },
    {
      "id": 899050,
      "postDate": "2020-06-23T23:31:24.713Z",
      "content": "<p>Hi, can I do \"Fake\" submission on kaggle to see the distribution of melanoma and benign classes to get a rough idea ? The classes are highly imbalanced in terms of data. I was wondering If I can do a fake submission will all zeros to see  the rough percentage estimation of benign images. Is it allowed ?</p>",
      "rawMarkdown": "Hi, can I do \"Fake\" submission on kaggle to see the distribution of melanoma and benign classes to get a rough idea ? The classes are highly imbalanced in terms of data. I was wondering If I can do a fake submission will all zeros to see  the rough percentage estimation of benign images. Is it allowed ?",
      "replies": [
        {
          "id": 906902,
          "postDate": "2020-06-29T15:51:17.927Z",
          "content": "<p><a href=\"/khizarhussain\">@khizarhussain</a> There is nothing that prevents you from making an all zeros submission. You may not hand-label the test set.</p>",
          "rawMarkdown": "@khizarhussain There is nothing that prevents you from making an all zeros submission. You may not hand-label the test set.",
          "votes": 1
        },
        {
          "id": 907138,
          "postDate": "2020-06-29T17:48:39.033Z",
          "content": "<p>Okay thank you</p>",
          "rawMarkdown": "Okay thank you"
        }
      ]
    },
    {
      "id": 896250,
      "postDate": "2020-06-22T02:44:48.170Z",
      "content": "<p>Hi Everyone,</p>\n\n<p>I have a question, In Leaderboard it says Public leaderboard calculates the score based on 30% of test data. What does that mean? Is it like if we have submitted for 100 test images, the score will be calculated for 30 test images and the remaining 70 images will be scored in the private leaderboard?</p>",
      "rawMarkdown": "Hi Everyone,\n\nI have a question, In Leaderboard it says Public leaderboard calculates the score based on 30% of test data. What does that mean? Is it like if we have submitted for 100 test images, the score will be calculated for 30 test images and the remaining 70 images will be scored in the private leaderboard?",
      "replies": [
        {
          "id": 896459,
          "postDate": "2020-06-22T07:36:14.900Z",
          "content": "<p>Quoting <a href=\"/juliaelliott\">@juliaelliott</a> - </p>\n\n<p>\"\"\nThis is a prediction submission competition, which means the test set that we're sharing on the Data page constitutes the entire test set. For your submission to Kaggle, you simply need to upload a CSV in the format of the sample submission on the Data page (containing the test set's observations and your corresponding predictions). The models you build locally, on Kaggle, or using your own methods are reserved for sharing if you are in winners' standing. We then have a hidden split of the test set's public and private id's, which provide you a public score, and reserve the private score for when the competition ends. \"\" </p>",
          "rawMarkdown": "Quoting @juliaelliott - \n\n\"\"\nThis is a prediction submission competition, which means the test set that we're sharing on the Data page constitutes the entire test set. For your submission to Kaggle, you simply need to upload a CSV in the format of the sample submission on the Data page (containing the test set's observations and your corresponding predictions). The models you build locally, on Kaggle, or using your own methods are reserved for sharing if you are in winners' standing. We then have a hidden split of the test set's public and private id's, which provide you a public score, and reserve the private score for when the competition ends. \"\" "
        },
        {
          "id": 897026,
          "postDate": "2020-06-22T15:22:45.223Z",
          "content": "<p>Thank you <a href=\"/akashram\">@akashram</a> </p>",
          "rawMarkdown": "Thank you @akashram "
        }
      ]
    },
    {
      "id": 884092,
      "postDate": "2020-06-13T07:29:44.330Z",
      "content": "<p>I'm going through fast.ai right now. First competition ever. I figure I'll learn a lot faster keeping things project-based. </p>",
      "rawMarkdown": "I'm going through fast.ai right now. First competition ever. I figure I'll learn a lot faster keeping things project-based. "
    },
    {
      "id": 881460,
      "postDate": "2020-06-11T04:04:29.910Z",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Can I use the datasets here for my research paper publication?</p>",
      "rawMarkdown": "@juliaelliott Can I use the datasets here for my research paper publication?",
      "replies": [
        {
          "id": 883593,
          "postDate": "2020-06-12T18:58:19.973Z",
          "content": "<p>Per the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">competition rules</a>, the dataset is licensed as <a href=\"https://creativecommons.org/licenses/by-nc/4.0/legalcode.txt\">CC BY-NC 4.0</a> for non-commercial use. So please be sure you include the correct attribution.</p>",
          "rawMarkdown": "Per the [competition rules](https://www.kaggle.com/c/siim-isic-melanoma-classification/rules), the dataset is licensed as [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/legalcode.txt) for non-commercial use. So please be sure you include the correct attribution.",
          "votes": 1
        },
        {
          "id": 883989,
          "postDate": "2020-06-13T06:42:14.293Z",
          "content": "<p>Can you tell me the correct attribution?</p>",
          "rawMarkdown": "Can you tell me the correct attribution?"
        },
        {
          "id": 885078,
          "postDate": "2020-06-13T21:20:37.033Z",
          "content": "<p>She means that you have to give credit to the author of the database. For example: \"All credit for this dataset goes to Mr. John Doe. I haven't done anything in creating it, so give credit to him and him only, okay?!?\" You get the picture.\nGo to this link: (<a href=\"https://en.wikipedia.org/wiki/Creative_Commons_license\">https://en.wikipedia.org/wiki/Creative_Commons_license</a>)\nScroll down until you see the table with icons. You'll understand.</p>",
          "rawMarkdown": "She means that you have to give credit to the author of the database. For example: \"All credit for this dataset goes to Mr. John Doe. I haven't done anything in creating it, so give credit to him and him only, okay?!?\" You get the picture.\nGo to this link: (https://en.wikipedia.org/wiki/Creative_Commons_license)\nScroll down until you see the table with icons. You'll understand.",
          "votes": 1
        },
        {
          "id": 885361,
          "postDate": "2020-06-14T06:07:55.077Z",
          "content": "<p><a href=\"/danoozy44\">@danoozy44</a> Thanks for clearing that out</p>",
          "rawMarkdown": "@danoozy44 Thanks for clearing that out"
        }
      ]
    },
    {
      "id": 879619,
      "postDate": "2020-06-09T15:32:06.537Z",
      "content": "<p>Hi Julia!\nHow is the final LB value calculated?\nIf one ensembles submissions, how do you use this information on the private test data?</p>\n\n<p>best Roman</p>",
      "rawMarkdown": "Hi Julia!\nHow is the final LB value calculated?\nIf one ensembles submissions, how do you use this information on the private test data?\n\nbest Roman",
      "replies": [
        {
          "id": 879671,
          "postDate": "2020-06-09T16:15:40.217Z",
          "content": "<p>Final LB is calculated from a private dataset which isn't shared to anyone , this is done to prevent people for overfitting due to competitiveness but you get a fair idea from the public LB which the test CSV is already in the dataset.</p>",
          "rawMarkdown": "Final LB is calculated from a private dataset which isn't shared to anyone , this is done to prevent people for overfitting due to competitiveness but you get a fair idea from the public LB which the test CSV is already in the dataset.",
          "votes": 3
        },
        {
          "id": 879705,
          "postDate": "2020-06-09T16:48:17.223Z",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> This is a prediction submission competition, which means the test set that we're sharing on the Data page constitutes the entire test set. For your submission to Kaggle, you simply need to upload a CSV in the format of the sample submission on the Data page (containing the test set's observations and your corresponding predictions). The models you build locally, on Kaggle, or using your own methods are reserved for sharing if you are in winners' standing. We then have a hidden split of the test set's public and private id's, which provide you a public score, and reserve the private score for when the competition ends.</p>",
          "rawMarkdown": "@romanweilguny This is a prediction submission competition, which means the test set that we're sharing on the Data page constitutes the entire test set. For your submission to Kaggle, you simply need to upload a CSV in the format of the sample submission on the Data page (containing the test set's observations and your corresponding predictions). The models you build locally, on Kaggle, or using your own methods are reserved for sharing if you are in winners' standing. We then have a hidden split of the test set's public and private id's, which provide you a public score, and reserve the private score for when the competition ends.",
          "votes": 1
        },
        {
          "id": 879805,
          "postDate": "2020-06-09T18:24:08.280Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a>  I see :)\nthx</p>",
          "rawMarkdown": "@juliaelliott  I see :)\nthx"
        }
      ]
    },
    {
      "id": 877975,
      "postDate": "2020-06-08T06:58:34.873Z",
      "content": "<p>Hi,</p>\n\n<p>For SIIM-ISIC Melanoma Classification competition , It has 180 GB of data.</p>\n\n<p>Do I need to download it on my local before I start working on it ?</p>",
      "rawMarkdown": "Hi,\n\nFor SIIM-ISIC Melanoma Classification competition , It has 180 GB of data.\n\nDo I need to download it on my local before I start working on it ?",
      "replies": [
        {
          "id": 878004,
          "postDate": "2020-06-08T07:23:02.003Z",
          "content": "<p>Why not use kaggle kernels instead ? It would save time as you don't need to download and it has all the libraries</p>",
          "rawMarkdown": "Why not use kaggle kernels instead ? It would save time as you don't need to download and it has all the libraries",
          "votes": 5
        }
      ]
    },
    {
      "id": 877971,
      "postDate": "2020-06-08T06:54:40.433Z",
      "content": "<p>Hello Everyone,</p>\n\n<p>I am new to kaggle,</p>\n\n<p>Thanks\nRahul</p>",
      "rawMarkdown": "Hello Everyone,\n\nI am new to kaggle,\n\n\nThanks\nRahul",
      "replies": [
        {
          "id": 964780,
          "postDate": "2020-08-10T06:55:09.630Z",
          "content": "<p>Welcome, Rahul. Tell me if you need any help. </p>\n\n<p>You can check peoples kernels to learn or goto kaggle.com/learn</p>",
          "rawMarkdown": "Welcome, Rahul. Tell me if you need any help. \n\nYou can check peoples kernels to learn or goto kaggle.com/learn"
        }
      ]
    },
    {
      "id": 903651,
      "postDate": "2020-06-27T02:59:55.117Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 903729,
          "postDate": "2020-06-27T04:20:02.030Z",
          "content": "<p>For this competition you submit your predictions. There are no time/processing limits on your models.</p>",
          "rawMarkdown": "For this competition you submit your predictions. There are no time/processing limits on your models."
        },
        {
          "id": 906906,
          "postDate": "2020-06-29T15:52:35.793Z",
          "content": "<p><a href=\"/tniac3\">@tniac3</a> quadcore’s response is correct.</p>",
          "rawMarkdown": "@tniac3 quadcore’s response is correct."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 876951,
      "author_name": "yash chaudhary",
      "author_url": "",
      "post_date": "2020-06-07T07:39:51.757000",
      "content": "<p>I am new to kaggle but have done many projects in DL.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 964778,
      "author_name": "Atri Saxena",
      "author_url": "",
      "post_date": "2020-08-10T06:53:02.043000",
      "content": "<p>Not new but trying to do kaggle at least one hour each day. Need Motivations. Thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 952564,
      "author_name": "Jaseem C K",
      "author_url": "",
      "post_date": "2020-07-31T03:50:14.233000",
      "content": "<p>This is my first cv competition. Got to learn a lot of stuff.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 953369,
          "author_name": "Ravinder Saluja",
          "author_url": "",
          "post_date": "2020-07-31T18:20:37.460000",
          "content": "<p>Same here. There is seriously a lot more that I need to learn.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945810,
      "author_name": "twintree",
      "author_url": "",
      "post_date": "2020-07-26T06:33:33.530000",
      "content": "<p>HI. I'm a beginner of competitions using images  and deep learning. I want to know if there are any public notebooks that kindly teaches the basics of modules and how it works?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 946152,
          "author_name": "Wwwwwwww",
          "author_url": "",
          "post_date": "2020-07-26T11:49:05",
          "content": "<p>For PyTorch I think this both are great for begginers:\n- <a href=\"https://www.kaggle.com/ankitjha/practical-deep-learning-using-pytorch\">https://www.kaggle.com/ankitjha/practical-deep-learning-using-pytorch</a>\n- <a href=\"https://www.kaggle.com/mratsim/starting-kit-for-pytorch-deep-learning\">https://www.kaggle.com/mratsim/starting-kit-for-pytorch-deep-learning</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 946229,
          "author_name": "twintree",
          "author_url": "",
          "post_date": "2020-07-26T12:56:50.303000",
          "content": "<p>Thanks a lot for your information!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 935375,
      "author_name": "Saurav",
      "author_url": "",
      "post_date": "2020-07-19T10:33:49.040000",
      "content": "<p>Hi                        </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 930499,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-15T14:05:04.173000",
      "content": "<p>hi <a href=\"/juliaelliott\">@juliaelliott</a> \nI have a question regarding public LB.\nI submitted several submissions with a 0.957 score in public LB, but got a different rank.\nI went from 27th to 22nd to 17th rank. Is there a hidden 4th post decimal value we don't see or is the private lb score used to rank the values with the same public LB score?</p>\n\n<p>best Roman</p>",
      "votes": 1,
      "replies": [
        {
          "id": 930658,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-07-15T16:17:38.113000",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> The public LB ranking will always be based on the public test subset, and based on the full precision of your score. We often only reveal a smaller number (3 or 4) of decimal places at the outset to protect against leaderboard probing, and then consider increasing precision as warranted. Therefore, yes, the ranking of multiple scores which appear to have the same score out to 3 decimal places will be based on hidden decimal values. We've gone ahead and updated the precision to 4 decimal places to help differentiate the playing field a bit.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 930708,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-15T16:58:15.923000",
          "content": "<p>thx - that was quick :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938595,
          "author_name": "Kamal Das",
          "author_url": "",
          "post_date": "2020-07-21T16:12:18.633000",
          "content": "<p>thanks for  clarifying. i also wondered what the process was</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 953367,
          "author_name": "Ravinder Saluja",
          "author_url": "",
          "post_date": "2020-07-31T18:19:31.137000",
          "content": "<p>Thanks for the answer. Even I was wondering the same thing. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 869395,
      "author_name": "Hamza Ghani",
      "author_url": "",
      "post_date": "2020-06-01T01:51:48.583000",
      "content": "<p>How can I download the data on a server? Is there a link I can wget from</p>",
      "votes": 1,
      "replies": [
        {
          "id": 877944,
          "author_name": "yash chaudhary",
          "author_url": "",
          "post_date": "2020-06-08T06:22:12.780000",
          "content": "<p>you can't use wget , you have to use kaggle api to download the dataset through code.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 864789,
      "author_name": "Littisha Lawrance",
      "author_url": "",
      "post_date": "2020-05-28T07:29:06.710000",
      "content": "<p>Hello,\nI'm new to kaggle! \nRegards,\nLittisha</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 935866,
      "author_name": "Adam Mehdi",
      "author_url": "",
      "post_date": "2020-07-19T18:04:47.347000",
      "content": "<p>New to kaggle. I am confused with why the GPU is not getting used when I am training my model. </p>\n\n<p>I have the GPU session enabled, but the GPU is not working as evident by the GPU bar in the top right of the notebook.</p>\n\n<p>How do I enable GPU if I already am in a GPU session? Thanks in advance.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 938894,
          "author_name": "Wwwwwwww",
          "author_url": "",
          "post_date": "2020-07-21T20:47:28.660000",
          "content": "<p>I think you are missing the device part in code where you select de GPU. If you using Pytorch it should be something like:\ndevice = torch.device(\"cuda\")\nmodel.to(device)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 982889,
          "author_name": "FoucardM",
          "author_url": "",
          "post_date": "2020-08-23T18:54:53.417000",
          "content": "<p>Hello, I am not an expert myself but started to use GPU and TPU for competition. here are some notebooks I have made to present how I managed to do it:<br>\n<a href=\"https://www.kaggle.com/foucardm/tutorial-top-15-mnist-digit-recognizer-on-tpu\" target=\"_blank\">GPU and TPU for TensorFlow</a> and this one with <a href=\"https://www.kaggle.com/foucardm/end-to-end-mnist-introduction-to-pytoch\" target=\"_blank\">PyTorch</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 869718,
      "author_name": "Navin",
      "author_url": "",
      "post_date": "2020-06-01T08:31:13.057000",
      "content": "<p>Hi, I am a recent joinee too. I have started a new series on Kaggle named <strong>\"Beginners' Mistakes\"</strong> over <a href=\"https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned\">here</a>. Do check it out  I hope it helps you. \nI would appreciate your views, perspectives, and support to continue it :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 866300,
      "author_name": "gaurav gupta",
      "author_url": "",
      "post_date": "2020-05-29T09:25:08.193000",
      "content": "<p>Hi im new to kaggle.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 983164,
      "author_name": "Cecile Guillot",
      "author_url": "",
      "post_date": "2020-08-24T04:59:53.580000",
      "content": "<p>Hello here !</p>\n<p>I'm new too. I began by learning with the micro-courses.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 982653,
      "author_name": "Abhi Govind",
      "author_url": "",
      "post_date": "2020-08-23T14:54:42.807000",
      "content": "<p>Hi everyone! I'm relatively new to kaggle and I want to know how to find the images present in the HAM10000 dataset in this dataset for an upcoming project. Can someone please explain to me how they would go about doing that? I would really appreciate it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 968052,
      "author_name": "Matias Micheletto",
      "author_url": "",
      "post_date": "2020-08-12T17:02:37.063000",
      "content": "<p>Hello, newbie here. Just started a few python programming courses after using Matlab for engineering during several years</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952648,
      "author_name": "Kaustubh Chaudhari",
      "author_url": "",
      "post_date": "2020-07-31T05:40:26.183000",
      "content": "<p>Hi </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 935449,
      "author_name": "ES.TAU2020",
      "author_url": "",
      "post_date": "2020-07-19T11:54:15.983000",
      "content": "<p>Hi everyone, quite a newbie student here on my first kaggle project. <br>\nI have a question regaring the evaluation of a submission:<br>\nI've seen different submission file with different ranges of values, some go up to 0.8 and some have a maximal value of 0.2. How is the threshold then set for the classification, I mean if my maximal value of the submission file is 0.6 and someone else's is 0.2? Is the entire column normalized and then sliced in the middle?</p>\n<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 932702,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-17T08:09:38.983000",
      "content": "<p>Hi</p>\n\n<p>I just ran into a problem with a self-created dataset I wanted to use in my kernel.</p>\n\n<p>I get an error message: Unexpected response from the service. Response: {'errors': ['Google Cloud SDK must be authorized before copying private datasets.'], 'error': {'code': 9}, 'wasSuccessful': False}.</p>\n\n<p>when accessing the dataset, which has been imported successfully.</p>\n\n<p>GCS_PATH_PSEUDO =  KaggleDatasets().get_gcs_path('pseudolabeled')</p>\n\n<p>Any idea?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 933447,
          "author_name": "devrishi",
          "author_url": "",
          "post_date": "2020-07-17T17:50:26.387000",
          "content": "<p>Hi Roman, were you able to try the instructions here <a href=\"https://www.kaggle.com/product-feedback/163416\">https://www.kaggle.com/product-feedback/163416</a> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 933485,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-17T18:11:42.940000",
          "content": "<p>Hi </p>\n\n<p>thanks - I have tried this in the meanwhile and it works.\nWhat puzzles me a bit - this worked for me without the new code before - but anyway... thx</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 916815,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-06T04:07:35.630000",
      "content": "<p>newbie here </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 916802,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-06T03:46:46.693000",
      "content": "<p>Hello</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 916227,
      "author_name": "Samantha Cruz",
      "author_url": "",
      "post_date": "2020-07-05T13:39:06.657000",
      "content": "<p>Hello everyone. I'm new to Kaggle. I'd like to know if I can use any programming language to solve this, since the submission is only the predictions. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 916267,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-05T14:18:38.490000",
          "content": "<p>I think the only limitation would be if you used a programming language that did not have a free version. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 916324,
          "author_name": "Samantha Cruz",
          "author_url": "",
          "post_date": "2020-07-05T14:53:18.380000",
          "content": "<p>So this means it's forbidden to use a proprietary software (not open source) to generate the predictions?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916447,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-05T16:50:05.310000",
          "content": "<p>I found this  in the rules. So I guess \"generally commercially available software...without undue expense\", allows some commercial software products. If you are using commercial software, you might want to verify with the Kaggle team that the specific software you want to use is allowed.</p>\n\n<p>(a) deliver to the Competition Sponsor the final model's software code as used to generate the winning Submission and associated documentation. The delivered software code should follow these documentation guidelines, must be capable of generating the winning Submission, and contain a description of resources required to build and/or run the executable code successfully. If the final model’s software code includes generally commercially available software that is not owned by you, but that can be procured by the Competition Sponsor without undue expense, then instead of delivering the code for that software to the Competition Sponsor, you must identify that software, method for procuring it, and any parameters or other information necessary to replicate the winning Submission;</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 914407,
      "author_name": "Dipanwita",
      "author_url": "",
      "post_date": "2020-07-03T20:56:29.957000",
      "content": "<p>hi, how can i load the datasets in google colab ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 910692,
      "author_name": "Patrick",
      "author_url": "",
      "post_date": "2020-07-01T10:02:00.943000",
      "content": "<p>Hello to everyone. I am new to Kaggle and this almost the first time. It seems to me that the test.csv file will be used to create predictions that will be submitted to the competition. The test file does not have 'target' column, which is of course the column to be predicted. My question is: How do you evaluate your models? Do you have to split again train.csv into train and test sets?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 910708,
          "author_name": "David Mbatuegwu",
          "author_url": "",
          "post_date": "2020-07-01T10:14:58.713000",
          "content": "<p>I'm new too, but I think I can help you, please give me details of your problem and i'll try to help you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 910968,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-01T13:24:21.637000",
          "content": "<p>You are on the right track. The train data is typically split into \"train\" and \"validation\" subsets. You train on the \"train\" data, and validate with the \"validation\" subset. The use of \"train\" in two different meanings can be confusing. Some datasets will come with pre-defined \"train\", \"validate\" and \"test\" datasets, so everybody can be consistent on what is the \"validate\" data. However, splitting out the validate data is part of the art of machine learning.</p>\n\n<p>It is common to split out 20% of the data for the validation subset. Often you will make five splits, each with a different 20% of the data as the validation and the other 80% of the data for training. Then you average or otherwise combine your 5 sets of results to get a submission. So if you put your data into 5 buckets, you would have:</p>\n\n<p>fold 1 - T T T T V\nfold 2 - T T T V T\nfold 3 - T T V T T\nfold 4 - T V T T T \nfold 5 - V T T T T</p>\n\n<p>Where you train on the T's in each fold and validate on the V's.</p>\n\n<p>For this contest you might want to make sure a given patient is not split between the train and validate datasets, since that could be a source of leakage.</p>\n\n<p>Separately, if you use External data, you might want the External data only in the Train dataset and not in the Validate dataset. Otherwise, your validate dataset results will be artificially improved. So, I split out my validation data before I add in the External data.</p>\n\n<p>Best of luck</p>\n\n<p>-Rich</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 899050,
      "author_name": "Khizar Hussain",
      "author_url": "",
      "post_date": "2020-06-23T23:31:24.713000",
      "content": "<p>Hi, can I do \"Fake\" submission on kaggle to see the distribution of melanoma and benign classes to get a rough idea ? The classes are highly imbalanced in terms of data. I was wondering If I can do a fake submission will all zeros to see  the rough percentage estimation of benign images. Is it allowed ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 906902,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-29T15:51:17.927000",
          "content": "<p><a href=\"/khizarhussain\">@khizarhussain</a> There is nothing that prevents you from making an all zeros submission. You may not hand-label the test set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 907138,
          "author_name": "Khizar Hussain",
          "author_url": "",
          "post_date": "2020-06-29T17:48:39.033000",
          "content": "<p>Okay thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 896250,
      "author_name": "Ram (ji)",
      "author_url": "",
      "post_date": "2020-06-22T02:44:48.170000",
      "content": "<p>Hi Everyone,</p>\n\n<p>I have a question, In Leaderboard it says Public leaderboard calculates the score based on 30% of test data. What does that mean? Is it like if we have submitted for 100 test images, the score will be calculated for 30 test images and the remaining 70 images will be scored in the private leaderboard?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 896459,
          "author_name": "AR",
          "author_url": "",
          "post_date": "2020-06-22T07:36:14.900000",
          "content": "<p>Quoting <a href=\"/juliaelliott\">@juliaelliott</a> - </p>\n\n<p>\"\"\nThis is a prediction submission competition, which means the test set that we're sharing on the Data page constitutes the entire test set. For your submission to Kaggle, you simply need to upload a CSV in the format of the sample submission on the Data page (containing the test set's observations and your corresponding predictions). The models you build locally, on Kaggle, or using your own methods are reserved for sharing if you are in winners' standing. We then have a hidden split of the test set's public and private id's, which provide you a public score, and reserve the private score for when the competition ends. \"\" </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 897026,
          "author_name": "Ram (ji)",
          "author_url": "",
          "post_date": "2020-06-22T15:22:45.223000",
          "content": "<p>Thank you <a href=\"/akashram\">@akashram</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 884092,
      "author_name": "Brad Hanks",
      "author_url": "",
      "post_date": "2020-06-13T07:29:44.330000",
      "content": "<p>I'm going through fast.ai right now. First competition ever. I figure I'll learn a lot faster keeping things project-based. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 881460,
      "author_name": "Digvijay Yadav",
      "author_url": "",
      "post_date": "2020-06-11T04:04:29.910000",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Can I use the datasets here for my research paper publication?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 883593,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-12T18:58:19.973000",
          "content": "<p>Per the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">competition rules</a>, the dataset is licensed as <a href=\"https://creativecommons.org/licenses/by-nc/4.0/legalcode.txt\">CC BY-NC 4.0</a> for non-commercial use. So please be sure you include the correct attribution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 883989,
          "author_name": "Digvijay Yadav",
          "author_url": "",
          "post_date": "2020-06-13T06:42:14.293000",
          "content": "<p>Can you tell me the correct attribution?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885078,
          "author_name": "G_R_S",
          "author_url": "",
          "post_date": "2020-06-13T21:20:37.033000",
          "content": "<p>She means that you have to give credit to the author of the database. For example: \"All credit for this dataset goes to Mr. John Doe. I haven't done anything in creating it, so give credit to him and him only, okay?!?\" You get the picture.\nGo to this link: (<a href=\"https://en.wikipedia.org/wiki/Creative_Commons_license\">https://en.wikipedia.org/wiki/Creative_Commons_license</a>)\nScroll down until you see the table with icons. You'll understand.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885361,
          "author_name": "Digvijay Yadav",
          "author_url": "",
          "post_date": "2020-06-14T06:07:55.077000",
          "content": "<p><a href=\"/danoozy44\">@danoozy44</a> Thanks for clearing that out</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 879619,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-06-09T15:32:06.537000",
      "content": "<p>Hi Julia!\nHow is the final LB value calculated?\nIf one ensembles submissions, how do you use this information on the private test data?</p>\n\n<p>best Roman</p>",
      "votes": 0,
      "replies": [
        {
          "id": 879671,
          "author_name": "yash chaudhary",
          "author_url": "",
          "post_date": "2020-06-09T16:15:40.217000",
          "content": "<p>Final LB is calculated from a private dataset which isn't shared to anyone , this is done to prevent people for overfitting due to competitiveness but you get a fair idea from the public LB which the test CSV is already in the dataset.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 879705,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-09T16:48:17.223000",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> This is a prediction submission competition, which means the test set that we're sharing on the Data page constitutes the entire test set. For your submission to Kaggle, you simply need to upload a CSV in the format of the sample submission on the Data page (containing the test set's observations and your corresponding predictions). The models you build locally, on Kaggle, or using your own methods are reserved for sharing if you are in winners' standing. We then have a hidden split of the test set's public and private id's, which provide you a public score, and reserve the private score for when the competition ends.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 879805,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-06-09T18:24:08.280000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a>  I see :)\nthx</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 877975,
      "author_name": "Rahul Kumar",
      "author_url": "",
      "post_date": "2020-06-08T06:58:34.873000",
      "content": "<p>Hi,</p>\n\n<p>For SIIM-ISIC Melanoma Classification competition , It has 180 GB of data.</p>\n\n<p>Do I need to download it on my local before I start working on it ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 878004,
          "author_name": "yash chaudhary",
          "author_url": "",
          "post_date": "2020-06-08T07:23:02.003000",
          "content": "<p>Why not use kaggle kernels instead ? It would save time as you don't need to download and it has all the libraries</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 877971,
      "author_name": "Rahul Kumar",
      "author_url": "",
      "post_date": "2020-06-08T06:54:40.433000",
      "content": "<p>Hello Everyone,</p>\n\n<p>I am new to kaggle,</p>\n\n<p>Thanks\nRahul</p>",
      "votes": 0,
      "replies": [
        {
          "id": 964780,
          "author_name": "Atri Saxena",
          "author_url": "",
          "post_date": "2020-08-10T06:55:09.630000",
          "content": "<p>Welcome, Rahul. Tell me if you need any help. </p>\n\n<p>You can check peoples kernels to learn or goto kaggle.com/learn</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 903651,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-27T02:59:55.117000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 903729,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-27T04:20:02.030000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 906906,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-29T15:52:35.793000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "864318": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).\n\nReady to dive into this competition? Review the [Overview Description](https://www.kaggle.com/c/siim-isic-melanoma-classification/overview) and start to work with the [Data](https://www.kaggle.com/c/siim-isic-melanoma-classification/data)!",
    "876951": "I am new to kaggle but have done many projects in DL.",
    "964778": "Not new but trying to do kaggle at least one hour each day. Need Motivations. Thanks.",
    "952564": "This is my first cv competition. Got to learn a lot of stuff.",
    "945810": "HI. I'm a beginner of competitions using images  and deep learning. I want to know if there are any public notebooks that kindly teaches the basics of modules and how it works?",
    "935375": "Hi                        \n",
    "930499": "hi @juliaelliott \nI have a question regarding public LB.\nI submitted several submissions with a 0.957 score in public LB, but got a different rank.\nI went from 27th to 22nd to 17th rank. Is there a hidden 4th post decimal value we don't see or is the private lb score used to rank the values with the same public LB score?\n\nbest Roman",
    "869395": "How can I download the data on a server? Is there a link I can wget from",
    "864789": "Hello,\nI'm new to kaggle! \nRegards,\nLittisha",
    "935866": "New to kaggle. I am confused with why the GPU is not getting used when I am training my model. \n\nI have the GPU session enabled, but the GPU is not working as evident by the GPU bar in the top right of the notebook.\n\nHow do I enable GPU if I already am in a GPU session? Thanks in advance.",
    "869718": "Hi, I am a recent joinee too. I have started a new series on Kaggle named **\"Beginners' Mistakes\"** over [here](https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned). Do check it out  I hope it helps you. \nI would appreciate your views, perspectives, and support to continue it :)",
    "866300": "Hi im new to kaggle.",
    "983164": "Hello here !\n\nI'm new too. I began by learning with the micro-courses.",
    "982653": "Hi everyone! I'm relatively new to kaggle and I want to know how to find the images present in the HAM10000 dataset in this dataset for an upcoming project. Can someone please explain to me how they would go about doing that? I would really appreciate it.",
    "968052": "Hello, newbie here. Just started a few python programming courses after using Matlab for engineering during several years",
    "952648": "Hi ",
    "935449": "Hi everyone, quite a newbie student here on my first kaggle project. \nI have a question regaring the evaluation of a submission:\nI've seen different submission file with different ranges of values, some go up to 0.8 and some have a maximal value of 0.2. How is the threshold then set for the classification, I mean if my maximal value of the submission file is 0.6 and someone else's is 0.2? Is the entire column normalized and then sliced in the middle?\n\nThanks",
    "932702": "Hi\n\nI just ran into a problem with a self-created dataset I wanted to use in my kernel.\n\nI get an error message: Unexpected response from the service. Response: {'errors': ['Google Cloud SDK must be authorized before copying private datasets.'], 'error': {'code': 9}, 'wasSuccessful': False}.\n\nwhen accessing the dataset, which has been imported successfully.\n\nGCS_PATH_PSEUDO =  KaggleDatasets().get_gcs_path('pseudolabeled')\n\nAny idea?",
    "916815": "newbie here ",
    "916802": "Hello",
    "916227": "Hello everyone. I'm new to Kaggle. I'd like to know if I can use any programming language to solve this, since the submission is only the predictions. ",
    "914407": "hi, how can i load the datasets in google colab ?",
    "910692": "Hello to everyone. I am new to Kaggle and this almost the first time. It seems to me that the test.csv file will be used to create predictions that will be submitted to the competition. The test file does not have 'target' column, which is of course the column to be predicted. My question is: How do you evaluate your models? Do you have to split again train.csv into train and test sets?",
    "899050": "Hi, can I do \"Fake\" submission on kaggle to see the distribution of melanoma and benign classes to get a rough idea ? The classes are highly imbalanced in terms of data. I was wondering If I can do a fake submission will all zeros to see  the rough percentage estimation of benign images. Is it allowed ?",
    "896250": "Hi Everyone,\n\nI have a question, In Leaderboard it says Public leaderboard calculates the score based on 30% of test data. What does that mean? Is it like if we have submitted for 100 test images, the score will be calculated for 30 test images and the remaining 70 images will be scored in the private leaderboard?",
    "884092": "I'm going through fast.ai right now. First competition ever. I figure I'll learn a lot faster keeping things project-based. ",
    "881460": "@juliaelliott Can I use the datasets here for my research paper publication?",
    "879619": "Hi Julia!\nHow is the final LB value calculated?\nIf one ensembles submissions, how do you use this information on the private test data?\n\nbest Roman",
    "877975": "Hi,\n\nFor SIIM-ISIC Melanoma Classification competition , It has 180 GB of data.\n\nDo I need to download it on my local before I start working on it ?",
    "877971": "Hello Everyone,\n\nI am new to kaggle,\n\n\nThanks\nRahul",
    "903651": ""
  }
}