{
  "id": 195283,
  "title": "Questions about submission process and scoring",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195283",
  "author_name": "",
  "post_date": "2020-11-04T12:55:14.835939Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I might have a misconception about how the submission works, and i have no idea, how scoring is done. Perhaps someone can clarify this?</p>\n<p>Here is how it works, or at least what I think: The submission code calls for each given dataframe the function env.predict() with a dataframe in the structure of sample_prediction_df as parameter. All the predicted data is stored in submission.csv.</p>\n<p>Scoring is done by comparing predicted and actual answers and divide the number of correct predictions by the number of questions.</p>\n<p>Obviously it is not that simple. After submission i get a score of 0.500 for 2 different Models (random and fix-value 1 - just for knowing, how submission works). This score is way to smooth to be calculated in the way, I think it is.</p>\n<p>Am I missing a step? What am I supposed to do with \"submission.csv\"?</p>\n<p>Any help is highly appreceated :)</p>\n<p>Daniel</p>",
  "messages": [
    {
      "id": "1069431",
      "postDate": "11/04/2020 12:55:14",
      "content": "<p>Hi,</p>\n<p>I might have a misconception about how the submission works, and i have no idea, how scoring is done. Perhaps someone can clarify this?</p>\n<p>Here is how it works, or at least what I think: The submission code calls for each given dataframe the function env.predict() with a dataframe in the structure of sample_prediction_df as parameter. All the predicted data is stored in submission.csv.</p>\n<p>Scoring is done by comparing predicted and actual answers and divide the number of correct predictions by the number of questions.</p>\n<p>Obviously it is not that simple. After submission i get a score of 0.500 for 2 different Models (random and fix-value 1 - just for knowing, how submission works). This score is way to smooth to be calculated in the way, I think it is.</p>\n<p>Am I missing a step? What am I supposed to do with \"submission.csv\"?</p>\n<p>Any help is highly appreceated :)</p>\n<p>Daniel</p>",
      "rawMarkdown": "Hi,\n\nI might have a misconception about how the submission works, and i have no idea, how scoring is done. Perhaps someone can clarify this?\n\nHere is how it works, or at least what I think: The submission code calls for each given dataframe the function env.predict() with a dataframe in the structure of sample_prediction_df as parameter. All the predicted data is stored in submission.csv.\n\nScoring is done by comparing predicted and actual answers and divide the number of correct predictions by the number of questions.\n\nObviously it is not that simple. After submission i get a score of 0.500 for 2 different Models (random and fix-value 1 - just for knowing, how submission works). This score is way to smooth to be calculated in the way, I think it is.\n\nAm I missing a step? What am I supposed to do with \"submission.csv\"?\n\nAny help is highly appreceated :)\n\nDaniel",
      "votes": null
    },
    {
      "id": "1069671",
      "postDate": "11/04/2020 18:34:20",
      "content": "<blockquote>\n  <p>Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.</p>\n</blockquote>\n<p>This is already explained in the Overview of the competition. When you call env.predict it automatically submits your dataframe as the prediction and stores it in submission.csv, so the file is automatically created for you. </p>\n<p>The example_sample_submission.csv is an example on how submission dataframe should be like, and example_test.csv is how the private test set would look like (2.5M rows). The loop loops through this set in small batches.</p>",
      "rawMarkdown": "> Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.\n\nThis is already explained in the Overview of the competition. When you call env.predict it automatically submits your dataframe as the prediction and stores it in submission.csv, so the file is automatically created for you. \n\nThe example_sample_submission.csv is an example on how submission dataframe should be like, and example_test.csv is how the private test set would look like (2.5M rows). The loop loops through this set in small batches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1069671,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "11/04/2020 18:34:20",
      "content": "<blockquote>\n  <p>Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.</p>\n</blockquote>\n<p>This is already explained in the Overview of the competition. When you call env.predict it automatically submits your dataframe as the prediction and stores it in submission.csv, so the file is automatically created for you. </p>\n<p>The example_sample_submission.csv is an example on how submission dataframe should be like, and example_test.csv is how the private test set would look like (2.5M rows). The loop loops through this set in small batches.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1069431": "Hi,\n\nI might have a misconception about how the submission works, and i have no idea, how scoring is done. Perhaps someone can clarify this?\n\nHere is how it works, or at least what I think: The submission code calls for each given dataframe the function env.predict() with a dataframe in the structure of sample_prediction_df as parameter. All the predicted data is stored in submission.csv.\n\nScoring is done by comparing predicted and actual answers and divide the number of correct predictions by the number of questions.\n\nObviously it is not that simple. After submission i get a score of 0.500 for 2 different Models (random and fix-value 1 - just for knowing, how submission works). This score is way to smooth to be calculated in the way, I think it is.\n\nAm I missing a step? What am I supposed to do with \"submission.csv\"?\n\nAny help is highly appreceated :)\n\nDaniel",
    "1069671": "> Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.\n\nThis is already explained in the Overview of the competition. When you call env.predict it automatically submits your dataframe as the prediction and stores it in submission.csv, so the file is automatically created for you. \n\nThe example_sample_submission.csv is an example on how submission dataframe should be like, and example_test.csv is how the private test set would look like (2.5M rows). The loop loops through this set in small batches."
  },
  "source": "meta"
}