{
  "id": 189804,
  "title": "EdNet TFRecord dataset (sequential)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189804",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2020-10-08T20:28:11.387000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I published a dataset <a href=\"https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential\" target=\"_blank\">EdNet TFRecords (sequential)</a> for the training dataset in this competition.</p>\n<p>The dataset is stored in <code>TFRecord</code> format. For each user, the dataset gives a dictionary whose keys are the column names in the original competition train.csv file. The corresponding value of each key is the sequence of records of that user for the corresponding attribute.</p>\n<p><strong>Be careful</strong>, for <code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code>, an NaN value in the original train.csv is converted to <code>-1.0</code> and <code>-1</code> in this TFRecord dataset.</p>\n<p>Since the sequences are of different lengths for each user, while loading the dataset, we will use <code>tf.io.RaggedFeature</code>. When the dataset is batched, we obtain <a href=\"https://www.tensorflow.org/api_docs/python/tf/RaggedTensor\" target=\"_blank\">RaggedTensor</a>.</p>\n<p>This dataset might be helpful for people who want to train RNN or Transformer models for this competition. You might still to figure out how to transform the data given by this dataset in order to train models.</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential\" target=\"_blank\">EdNet TFRecords (sequential)</a></p>\n<p>Notebook: <a href=\"https://www.kaggle.com/yihdarshieh/starter-ednet-tfrecords-sequential\" target=\"_blank\">Starter: EdNet TFRecords (sequential)</a></p>",
  "messages": [
    {
      "id": 1043273,
      "postDate": "2020-10-08T20:28:11.387Z",
      "content": "<p>Hi,</p>\n<p>I published a dataset <a href=\"https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential\" target=\"_blank\">EdNet TFRecords (sequential)</a> for the training dataset in this competition.</p>\n<p>The dataset is stored in <code>TFRecord</code> format. For each user, the dataset gives a dictionary whose keys are the column names in the original competition train.csv file. The corresponding value of each key is the sequence of records of that user for the corresponding attribute.</p>\n<p><strong>Be careful</strong>, for <code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code>, an NaN value in the original train.csv is converted to <code>-1.0</code> and <code>-1</code> in this TFRecord dataset.</p>\n<p>Since the sequences are of different lengths for each user, while loading the dataset, we will use <code>tf.io.RaggedFeature</code>. When the dataset is batched, we obtain <a href=\"https://www.tensorflow.org/api_docs/python/tf/RaggedTensor\" target=\"_blank\">RaggedTensor</a>.</p>\n<p>This dataset might be helpful for people who want to train RNN or Transformer models for this competition. You might still to figure out how to transform the data given by this dataset in order to train models.</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential\" target=\"_blank\">EdNet TFRecords (sequential)</a></p>\n<p>Notebook: <a href=\"https://www.kaggle.com/yihdarshieh/starter-ednet-tfrecords-sequential\" target=\"_blank\">Starter: EdNet TFRecords (sequential)</a></p>",
      "rawMarkdown": "Hi,\n\nI published a dataset [EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential) for the training dataset in this competition.\n\nThe dataset is stored in `TFRecord` format. For each user, the dataset gives a dictionary whose keys are the column names in the original competition train.csv file. The corresponding value of each key is the sequence of records of that user for the corresponding attribute.\n\n**Be careful**, for `prior_question_elapsed_time` and `prior_question_had_explanation`, an NaN value in the original train.csv is converted to `-1.0` and `-1` in this TFRecord dataset.\n\n\nSince the sequences are of different lengths for each user, while loading the dataset, we will use `tf.io.RaggedFeature`. When the dataset is batched, we obtain [RaggedTensor](https://www.tensorflow.org/api_docs/python/tf/RaggedTensor).\n\nThis dataset might be helpful for people who want to train RNN or Transformer models for this competition. You might still to figure out how to transform the data given by this dataset in order to train models.\n\nDataset: [EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential)\n\nNotebook: [Starter: EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/starter-ednet-tfrecords-sequential)",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1043273": "Hi,\n\nI published a dataset [EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential) for the training dataset in this competition.\n\nThe dataset is stored in `TFRecord` format. For each user, the dataset gives a dictionary whose keys are the column names in the original competition train.csv file. The corresponding value of each key is the sequence of records of that user for the corresponding attribute.\n\n**Be careful**, for `prior_question_elapsed_time` and `prior_question_had_explanation`, an NaN value in the original train.csv is converted to `-1.0` and `-1` in this TFRecord dataset.\n\n\nSince the sequences are of different lengths for each user, while loading the dataset, we will use `tf.io.RaggedFeature`. When the dataset is batched, we obtain [RaggedTensor](https://www.tensorflow.org/api_docs/python/tf/RaggedTensor).\n\nThis dataset might be helpful for people who want to train RNN or Transformer models for this competition. You might still to figure out how to transform the data given by this dataset in order to train models.\n\nDataset: [EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential)\n\nNotebook: [Starter: EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/starter-ednet-tfrecords-sequential)"
  }
}