{"cells":[{"metadata":{},"cell_type":"markdown","source":"# This Notebook is Training word2vec Model with content_id\nFirst of all, Thank you Kaggle, the Competition Host and competitor!!  \nThis competition gives me a lot of learning.  \n  \nAnd I'm impressed with Kaggle's culture of sharing ideas with everyone.\n\nIn this competition, I learned how to use the new idea of word2vec from [this discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209576) and other.  \nThank you @ML_Bear for sharing ideas!!\n\nSo I decided to try implementing this idea.  \n\n\nThis is my first implementation of word2vec, so I may make some mistakes.  \nBut I'll share it in this notebook!!  \n\nIf you have any concerns about memory savings or processing speed, please let me know.  "},{"metadata":{},"cell_type":"markdown","source":"# Import"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!pip install ../input/python-datatable/datatable-0.11.0-cp37-cp37m-manylinux2010_x86_64.whl > /dev/null 2>&1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport random\nimport pandas as pd\nimport datatable as dt\nimport gc\nfrom tqdm.notebook import tqdm\nfrom collections import defaultdict","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Read Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"data_types_dict = {\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'content_type_id':'int8',\n    'answered_correctly': 'int8'\n}\n\ntrain_df = dt.fread('../input/riiid-test-answer-prediction/train.csv', columns=set(data_types_dict.keys())).to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# delete lecture rows\ntrain_df = train_df[~train_df['content_type_id']]\ntrain_df.drop('content_type_id', axis=1, inplace=True)\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data to Sentence"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Added the 6th digit to question_id to distinguish answered_correctly\ntrain_df['word'] = train_df['content_id'] + train_df['answered_correctly'] * 100000","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# memory error\n# sentences = {}\n\n# for _, group in enumerate(tqdm(train_df.groupby('user_id'))):\n#     sentences[group[0]] = list(group[1]['word'].apply(str))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sentences = defaultdict(list)\n\nfor _,row in enumerate(tqdm(train_df[['user_id', 'word']].values)):\n    sentences[row[0]].append(str(row[1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del train_df\nsentences = list(sentences.values())\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# word2vec"},{"metadata":{"trusted":true},"cell_type":"code","source":"from gensim.models import Word2Vec, KeyedVectors\n\nmodel = Word2Vec(sentences,  sg=1, size=100, window=5, min_count=1, sample=0)\nmodel.wv.save_word2vec_format(\"vec.pt\", binary=True)\n\ndel model, sentences\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"wv = KeyedVectors.load_word2vec_format('vec.pt', binary=True)\nprint('most_similar')\nprint(wv.most_similar('105692'))\n\nprint('get_vector')\nprint(wv.get_vector('105692'))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"[This Notebook](https://www.kaggle.com/imazekishota/riiid-lgbm-with-word2vec) is LGBM training code."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}