{"cells":[{"metadata":{},"cell_type":"markdown","source":"# About this kernel\n\nThis notebook demonstrates how to use this Kaggle dataset [EdNet TFRecords (sequential)](https://www.kaggle.com/yihdarshieh/ednet-tfrecords-sequential). This is the training dataset for the competition [Riiid AIEd Challenge 2020](https://www.kaggle.com/c/riiid-test-answer-prediction/overview).\n\nThe dataset is stored in `TFRecord` format. For each user, the dataset gives a dictionary whose keys are the column names in the original competition train.csv file. The corresponding value of each key is the sequence of records of that user for the corresponding attribute.\n\nSince the sequences are of different lengths for each user, while loading the dataset, we will use `tf.io.RaggedFeature`. When the dataset is batched, we obtain [RaggedTensor](https://www.tensorflow.org/api_docs/python/tf/RaggedTensor).\n\nThis dataset might be helpful for people who want to train RNN or Transformer models for this competition. You might still to figure out how to transform the data given by this dataset in order to train models."},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nimport tensorflow as tf","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## TFRecord files"},{"metadata":{"trusted":true},"cell_type":"code","source":"tfrec_dir = '/kaggle/input/ednet-tfrecords-sequential/'\ntfrec_fns = os.listdir(tfrec_dir)\n\n# Sort the file - For verification purpose\ntfrec_fns = sorted(tfrec_fns, key=lambda x: int(x.replace('EdNet-user-history-', '').replace('.tfrecord', '')))\n\ntfrec_paths = [os.path.join(tfrec_dir, fn) for fn in tfrec_fns]\ntfrec_paths","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Load TFRecord files - tf.io.RaggedFeature\n\n\nGuide for Ragged tensor\n\nhttps://www.tensorflow.org/guide/ragged_tensor"},{"metadata":{"trusted":true},"cell_type":"code","source":"features = {\n    'user_id': tf.io.FixedLenFeature([], dtype=tf.int64),\n    # Zero partitions: returns 1D tf.Tensor for each Example.\n    'row_id': tf.io.RaggedFeature(value_key='row_id', dtype=tf.int64),\n    'timestamp': tf.io.RaggedFeature(value_key='timestamp', dtype=tf.int64),\n    'content_id': tf.io.RaggedFeature(value_key='content_id', dtype=tf.int64),\n    'content_type_id': tf.io.RaggedFeature(value_key='content_type_id', dtype=tf.int64),\n    'task_container_id': tf.io.RaggedFeature(value_key='task_container_id', dtype=tf.int64),\n    'user_answer': tf.io.RaggedFeature(value_key='user_answer', dtype=tf.int64),\n    'answered_correctly': tf.io.RaggedFeature(value_key='answered_correctly', dtype=tf.int64),\n    'prior_question_elapsed_time': tf.io.RaggedFeature(value_key='prior_question_elapsed_time', dtype=tf.float32),\n    'prior_question_had_explanation': tf.io.RaggedFeature(value_key='prior_question_had_explanation', dtype=tf.int64),\n}\n\n\ndef parse_example(example):\n\n    return tf.io.parse_single_example(example, features)\n\n\n# num_parallel_reads=1 to avoid reading in parallel, so the order is kept - for verification purpose here.\n# In real appliation where the order is not important, set it to be > 1 to gain performance.\nds = tf.data.TFRecordDataset(tfrec_paths, num_parallel_reads=1)\nds = ds.map(parse_example)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Check elements in the dataset"},{"metadata":{},"cell_type":"markdown","source":"### unbatched\n\nWe only get the usual `tf.Tensor`, not `tf.RaggedTensor`."},{"metadata":{"trusted":true},"cell_type":"code","source":"for x in ds.take(3):\n    print(x)\n    print('-' * 40)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### batched\n\nNow we have `tf.RaggedTensor` as expected (except for `user_id`). The first batch is for user id `115` and `124`."},{"metadata":{"trusted":true},"cell_type":"code","source":"batched_ds = ds.apply(\n    tf.data.experimental.dense_to_ragged_batch(batch_size=2)\n)\n\nfor x in batched_ds.take(1):\n    print(x)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Next step\n\nTrain a RNN or Transformer model using this dataset, have fun!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}