{
  "id": 79140,
  "title": "\"This leaderboard is calculated with all of the test data.\" What does that mean?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79140",
  "author_name": "newbiegyc",
  "post_date": "2019-01-31T10:10:24.771000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>What data is used for private LB?</p>",
  "messages": [
    {
      "id": 464421,
      "postDate": "2019-01-31T20:53:40.637Z",
      "content": "<p>At the moment we have a file test.csv that contains ~56k samples. When people submit their solution, Kaggle evaluates all the ~56k predictions (score on LB). After the competition ends, your Kernel is run to make new predictions based on a new test.csv file including ~376 samples.</p>\n\n<p>As stated in the Competition description:</p>\n\n<p>In the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:</p>\n\n<p>test.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.</p>",
      "rawMarkdown": "At the moment we have a file test.csv that contains ~56k samples. When people submit their solution, Kaggle evaluates all the ~56k predictions (score on LB). After the competition ends, your Kernel is run to make new predictions based on a new test.csv file including ~376 samples.\n\nAs stated in the Competition description:\n\nIn the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:\n\ntest.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.",
      "votes": 1
    },
    {
      "id": 464195,
      "postDate": "2019-01-31T10:55:58.857Z",
      "content": "<p>There will stage 2 of the competition with a new test dataset. </p>",
      "rawMarkdown": "There will stage 2 of the competition with a new test dataset. ",
      "votes": 1
    },
    {
      "id": 464181,
      "postDate": "2019-01-31T10:10:24.770Z",
      "content": "<p>What data is used for private LB?</p>",
      "rawMarkdown": "What data is used for private LB?",
      "votes": 1
    },
    {
      "id": 464245,
      "postDate": "2019-01-31T13:12:20.817Z",
      "content": "<p>A few weeks ago <a href=\"/inversion\">@inversion</a> provided some explanation, see the link below\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76792#451902\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76792#451902</a></p>",
      "rawMarkdown": "A few weeks ago @inversion provided some explanation, see the link below\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76792#451902\n",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 464421,
      "author_name": "gsaintcirgue",
      "author_url": "",
      "post_date": "2019-01-31T20:53:40.637000",
      "content": "<p>At the moment we have a file test.csv that contains ~56k samples. When people submit their solution, Kaggle evaluates all the ~56k predictions (score on LB). After the competition ends, your Kernel is run to make new predictions based on a new test.csv file including ~376 samples.</p>\n\n<p>As stated in the Competition description:</p>\n\n<p>In the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:</p>\n\n<p>test.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 464195,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-31T10:55:58.857000",
      "content": "<p>There will stage 2 of the competition with a new test dataset. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 464245,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-31T13:12:20.817000",
      "content": "<p>A few weeks ago <a href=\"/inversion\">@inversion</a> provided some explanation, see the link below\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76792#451902\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76792#451902</a></p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "464421": "At the moment we have a file test.csv that contains ~56k samples. When people submit their solution, Kaggle evaluates all the ~56k predictions (score on LB). After the competition ends, your Kernel is run to make new predictions based on a new test.csv file including ~376 samples.\n\nAs stated in the Competition description:\n\nIn the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:\n\ntest.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.",
    "464195": "There will stage 2 of the competition with a new test dataset. ",
    "464181": "What data is used for private LB?",
    "464245": "A few weeks ago @inversion provided some explanation, see the link below\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76792#451902\n"
  }
}