{
  "id": 405971,
  "title": "Why do I have only 9 questions in my train_labels.csv right now?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/405971",
  "author_name": "",
  "post_date": "2023-04-30T07:48:57.084113700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I downloaded train_labels.csv dataset yesterday but there are only 9 questions there. Why so? </p>",
  "messages": [
    {
      "id": "2240071",
      "postDate": "04/30/2023 07:48:57",
      "content": "<p>I downloaded train_labels.csv dataset yesterday but there are only 9 questions there. Why so? </p>",
      "rawMarkdown": "I downloaded train_labels.csv dataset yesterday but there are only 9 questions there. Why so?",
      "votes": null
    },
    {
      "id": "2242940",
      "postDate": "05/02/2023 15:31:15",
      "content": "<p>There is something wrong how you parse initial session_id column. As we can see there 424116 rows for 18 questions, which result in 424116/18=23562 session which is a true value of real session. Maybe you take only first digit thus 11 and 1 are the same? I would suggest just do<br>\nlabels[['session', 'q']] = labels['session_id'].str.split('_', expand = True)</p>",
      "rawMarkdown": "There is something wrong how you parse initial session_id column. As we can see there 424116 rows for 18 questions, which result in 424116/18=23562 session which is a true value of real session. Maybe you take only first digit thus 11 and 1 are the same? I would suggest just do\nlabels[['session', 'q']] = labels['session_id'].str.split('_', expand = True)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242940,
      "author_name": "glipko",
      "author_url": "",
      "post_date": "05/02/2023 15:31:15",
      "content": "<p>There is something wrong how you parse initial session_id column. As we can see there 424116 rows for 18 questions, which result in 424116/18=23562 session which is a true value of real session. Maybe you take only first digit thus 11 and 1 are the same? I would suggest just do<br>\nlabels[['session', 'q']] = labels['session_id'].str.split('_', expand = True)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2240071": "I downloaded train_labels.csv dataset yesterday but there are only 9 questions there. Why so?",
    "2242940": "There is something wrong how you parse initial session_id column. As we can see there 424116 rows for 18 questions, which result in 424116/18=23562 session which is a true value of real session. Maybe you take only first digit thus 11 and 1 are the same? I would suggest just do\nlabels[['session', 'q']] = labels['session_id'].str.split('_', expand = True)"
  },
  "source": "meta"
}