{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"One of the most important things to establish for every Kaggle competition is whether there is a significatn differnece in distributions of the train and test sets. So far the CV validation scores for kernels and for public LB have been pretty close, suggesting that the two distributions are pretty similar. However, it would be interesting, and potentially very valuable, to find out in a more quantitative and specific way how do these distributions compare. For that purpose we'll build an adverserial validation scheme - we'll run a CV classifier that tries to predict if any given question belongs to the train or the test set. ","metadata":{"_uuid":"e522fb0f302f8511df02808b0013cce499f64333"}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nfrom sklearn.metrics import f1_score, roc_auc_score\n\nimport gc\n\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import cross_val_score\nfrom scipy.sparse import hstack\nfrom sklearn.metrics import f1_score\nfrom sklearn.model_selection import KFold\nfrom scipy.sparse import hstack\nfrom scipy.sparse import coo_matrix\nfrom tqdm import tqdm\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')\ntest = pd.read_csv('../input/test.csv')\ntrain.head()","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['target'] = 0\ntest['target'] = 1","metadata":{"_uuid":"77b37611cd82590eebf153e4ec734c59ee967164","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test = pd.concat([train, test], axis =0)","metadata":{"_uuid":"cdcd6078f24fa93c6b886ddd9efce6cfae0a7c09","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test.tail()","metadata":{"_uuid":"478c9f6e5077bd808d08329dec4c23a575a96d53","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target = train_test['target'].values","metadata":{"_uuid":"7441f08f77c996086e07333e7d23be2f8eb4d2d1","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target = train_test['target'].values\n\ntext = train_test['question_text']\n\n\ndel train, test, train_test\ngc.collect()\n\n\nword_vectorizer = TfidfVectorizer(\n    sublinear_tf=True,\n    strip_accents='unicode',\n    analyzer='word',\n    token_pattern=r'\\w{1,}',\n    stop_words='english',\n    ngram_range=(1, 1),\n    max_features=7500)\nword_vectorizer.fit(text)\nword_features = word_vectorizer.transform(text)\n\ndel word_vectorizer\ngc.collect()\n","metadata":{"_uuid":"9eea99e486eb8002a0cf2eb335bbcabfdc14495f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kf = KFold(n_splits=5, shuffle=True, random_state=43)\noof_pred = np.zeros([target.shape[0],])\n\nfor i, (train_index, val_index) in tqdm(enumerate(kf.split(target))):\n    x_train, x_val = word_features[list(train_index)], word_features[list(val_index)]\n    y_train, y_val = target[train_index], target[val_index]\n    classifier = LogisticRegression(C=5, solver='sag')\n    classifier.fit(x_train, y_train)\n    val_preds = classifier.predict_proba(x_val)[:,1]\n    oof_pred[val_index] = val_preds\n    print(f1_score(y_val, val_preds > 0.1))\n","metadata":{"_uuid":"bb8a28c77d20c37dbc1931e8c2b14682e334faa9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score = 0\nthresh = .5\nfor i in np.arange(0.1, 1.001, 0.01):\n    temp_score = f1_score(target, (oof_pred > i))\n    if(temp_score > score):\n        score = temp_score\n        thresh = i\n\nprint(\"CV: {}, Threshold: {}\".format(score, thresh))","metadata":{"_uuid":"5530ca9d4ae5c66f92bb7dfb9652a544a5a291ce","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roc_auc_score(target, oof_pred)","metadata":{"_uuid":"906db33fa5c122d6fa6251cd4da30024a310e094","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So with F1 score at about 0.01 and AUC at almost 0.5, it would seem that the two distributions are pretty similar, at least as far as can be determined by the distribution of individual words.","metadata":{"_uuid":"102999fc84cb1446853ff92f50cfbd4e6c43a4a9"}},{"cell_type":"code","source":"","metadata":{"_uuid":"46b74b9431ca1aac7b94820641008fe455b03194","trusted":true},"execution_count":null,"outputs":[]}]}