{"cells":[{"metadata":{"_uuid":"0d35d91d8a54916db80ee1410b4e2733aee00473"},"cell_type":"markdown","source":"Orginally want to try back-translation, but that requires internet connection which is not allowed if I want to submit. So instead, I tried the data augmentation method metioned here https://www.kaggle.com/jpmiller/extending-train-data-with-markov-chains-auc. It is using markov chain to learn the word sequence and generate new ones. However, I saw huge overfiting on my own validation set. So I didn't continue on this method, but still, it is interesting to see the machine generated sentences."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import markovify as mk\nfrom joblib import Parallel, delayed","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"07e8e0bb1b760caeda01286ee10f1921e39c3e68"},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"41dbc53c9b4da56e842bf7d277e0d922619ef31c"},"cell_type":"code","source":"insincere = train.loc[train.target==1, ['question_text', 'target']]\ninsincere.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"94f8b768907f28bb32ecdd0926c3a82b1af8cdea"},"cell_type":"code","source":"nchar = int(insincere.question_text.str.len().median())\nnchar","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a8f73aaa2d19dc0ca2c0af5c5af6c0899e117ca9"},"cell_type":"code","source":"text_model = mk.Text(insincere['question_text'].tolist())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"66b4b212ea140060f15dbdcfbe168c0f811c8e2f"},"cell_type":"code","source":"def data_augment():\n    return text_model.make_short_sentence(nchar)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4e5bf9323235200aa034c513dfedd219484cfdbc"},"cell_type":"code","source":"parallel = Parallel(-1, backend=\"threading\", verbose=5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4faf12d0b3b99f1ddf17f06499259c2dc980e2c"},"cell_type":"code","source":"count = 1000\naug_data = parallel(delayed(data_augment)() for _ in range(count))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"53ec01edb70cef3bffcd01d93008171fa0f7d442"},"cell_type":"code","source":"aug_data[:5]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3b790f4ae48ed173989db47121eb9f115b4737bf"},"cell_type":"markdown","source":"I then tried my same model on the augemented training data, and see huge overfitting on my validation set. I guess the reason is that it does not change the vocabulary at all, and does not change the word order very much. Thus it does not provide much diversity to the data."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}