{"cells":[{"metadata":{"_uuid":"6e549e4cfd768fb713a841a2e67ae0ae7cd90086"},"cell_type":"markdown","source":"# Linear model using Vowpal Wabbit\n\n<img src=\"https://cdn.dribbble.com/users/261617/screenshots/3146111/vw-dribbble.png\" alt=\"drawing\" width=\"400\"/>\n\n## What is Vowpal Wabbit?\nThis is a package to perform fast training of linear models. It is a very sophysticated tool, that allows to use many different advanced algorithms. I encourage you to check out their [Wiki on github](https://github.com/VowpalWabbit/vowpal_wabbit/wiki)\n\nFor the baseline the two key issues are:\n\n- it is **fast**. In fact, the package allows online learning, i.e. sequential processing of training examples one-by-one, similar to what SGDClassifier/SGDRegressor models in sklearn aim to achieve. But in command-line mode one truelly reads only a single line  from an input file into memory, thus one can train a model on a dataset that does not fit into memory.\n- it **applies hashing on text features**. \n   - This means that we do not need to run much of pre-processing and can let the machine to do the learning. That's what is implemented in this baseline- we directly feed the message text into the training removing punctuation and stopwords only (the latter is not needed in fact). \n   - This also means that the text features are stored in a more compact form that the naive OHE (=BoW) representation, thus memory footprint is reduced.\n   \nIn the following kernel the sklearn API of VW is used. This allows to use the same data and the same methods to be used in VW as well as in other ML tools. A small technical note: the dataset had to be slightly processed, as the VW internal functions can not properly handle text inputs in a DataFrame.\n\nThe ideas of using class weighting for disbalance problem, threshold optimisation and ngrams come from https://www.kaggle.com/hippskill/vowpal-wabbit-starter-pack (check it out- it is a solid piece of work)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=Warning)\n\nimport os\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7cfb9f38171e59bf3cd41ecff0414e57897e1e1e"},"cell_type":"markdown","source":"Read in the data"},{"metadata":{"trusted":true,"_uuid":"cf009a2f6a9c37d4aebbdaea9c0d1574b99ee145"},"cell_type":"code","source":"df_trn = pd.read_csv('../input/train.csv')\ndf_tst = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e24a13ed421319ac6a00139d077adfccf17b5e88"},"cell_type":"markdown","source":"Print data stats"},{"metadata":{"trusted":true,"_uuid":"c29613f72c0224c9ff0c329570120243411be166","_kg_hide-input":true},"cell_type":"code","source":"print('Train and test shapes are: {}, {}'.format(df_trn.shape, df_tst.shape))\nprint('Train and test memory footprint: {:.2f} MB, {:.2f} MB'\n      .format(df_trn.memory_usage(deep=True).sum()/ 1024**2,\n              df_tst.memory_usage(deep=True).sum()/ 1024**2)\n     )\nw_pos = df_trn['target'].sum()/df_trn.shape[0]\nprint('Fraction of positive target (insencere) = {:.4f}'.format(w_pos))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c428fb2a2c381ca2cfe4a203d32da374696ab2ba"},"cell_type":"markdown","source":"Display a couple of first entries"},{"metadata":{"trusted":true,"_uuid":"b3ef2c039acc454462ceb4d59ae77adf16e1e58c"},"cell_type":"code","source":"df_trn.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d53f5e4e2f5ea71281dbbc5a57cfd1976eeba35a"},"cell_type":"markdown","source":"Stop words"},{"metadata":{"trusted":true,"_uuid":"16d4712aac5fe619d1056fa64761cb6b8fa1b758"},"cell_type":"code","source":"from nltk.corpus import stopwords\nstops = set(stopwords.words(\"english\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7e9294be8e12fb8a55390dfdb498b07cb3a018dc"},"cell_type":"markdown","source":"Extract the training features and the target variable"},{"metadata":{"trusted":true,"_uuid":"07dd24dfc7b6c7bc0c8eae53d0cccf1838b3a1ad","_kg_hide-input":true},"cell_type":"code","source":"import string\ndef remove_punctuation(s):\n    s = ''.join([i for i in s if i not in frozenset(string.punctuation)])\n    return s\n\nX_trn = (df_trn['question_text']\n         .apply(remove_punctuation)\n         .apply(lambda x: ' '.join([w.lower() for w in x.split(' ') if w.lower() not in stops]))\n        )\nX_trn2 = (df_trn['question_text']\n         .apply(remove_punctuation)\n         #.apply(lambda x: ' '.join([w.lower() for w in x.split(' ') if w.lower() not in stops]))\n        )\nX_tst = (df_tst['question_text']\n         .apply(remove_punctuation)\n         #.apply(lambda x: ' '.join([w.lower() for w in x.split(' ') if w.lower() not in stops]))\n        )\ny_trn = df_trn['target']\n\ndel df_trn, df_tst","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"601b2358bd23d0dbe8fb47003d5e31ac8fc9ce84"},"cell_type":"code","source":"X_trn.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"Helper functions for Vowpal Wabbit"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"0c919e32f2aa2e49aab0deb4baf8afa95396f2a4"},"cell_type":"code","source":"import vowpalwabbit as vw\nfrom vowpalwabbit.sklearn_vw import VWClassifier\n\n# VW uses 1/-1 target variables for classification instead of 1/0, so we need to apply mapping\ndef convert_labels_sklearn_to_vw(y_sklearn):\n    return y_sklearn.map({1:1, 0:-1})\n\n# The function to create VW-compatible inputs from the text features and the target\ndef to_vw(X, y=None, namespace='Name', w=None):\n    labels = '1' if y is None else y.astype(str)\n    if w is not None:\n        labels = labels + ' ' + np.round(y.map({1: w, -1: 1}),5).astype(str)\n    prefix = labels + ' |' + namespace + ' '\n    if isinstance(X, pd.DataFrame):\n        return prefix + X.apply(lambda x: ' '.join(x), axis=1)\n    elif isinstance(X, pd.Series):\n        return prefix + X","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4464a90caa8e647ca5f36a2c9341644399172d3a"},"cell_type":"markdown","source":"Define the model that will be trained: our `VW_passes3` model will do 3 iterations (=passes) over the data using the prepared text format as the input. **The threshold to apply for 0/1 label assignment was tuned on the data to get the best F1 score**"},{"metadata":{"trusted":true,"_uuid":"2169b5c08ae3c555cc2a49443b097821763b0b6d","_kg_hide-input":true},"cell_type":"code","source":"mdl_inputs = {\n# VW1x is analogous to the configuration from the kernel cited in the intro\n#                 'VW1x': [VWClassifier(quiet=False, convert_to_vw=False, \n#                                      passes=1, link='logistic',\n#                                      random_seed=314),\n#                              {'pos_threshold':0.5, \n#                               'b':29, 'ngram':2, 'skips': 1, \n#                               'l1':3.4742122764e-09, 'l2':1.24232077629e-11},\n#                              {},\n#                              None,\n#                              None,\n#                              1./w_pos\n#                             ],\n                'VW1': [VWClassifier(quiet=False, convert_to_vw=False, \n                                     passes=3, link='logistic',\n                                     random_seed=314),\n                             {'pos_threshold':0.5},\n                             {},\n                             None,\n                             None,\n                             1./w_pos\n                            ],\n#                 'VW2': [VWClassifier(quiet=False, convert_to_vw=False, \n#                                      passes=5, link='logistic',\n#                                      random_seed=314),\n#                              {'pos_threshold':0.5},\n#                              {},\n#                              None,\n#                              None,\n#                              1./w_pos\n#                             ],\n#                 'VW3': [VWClassifier(quiet=False, convert_to_vw=False, \n#                                      passes=10, link='logistic',\n#                                      random_seed=314),\n#                              {'pos_threshold':0.5},\n#                              {},\n#                              None,\n#                              None,\n#                              1./w_pos\n#                             ],\n         }\n\n# for i in [22]:\n#     mdl_inputs['VW_passes3_thrs{}'.format(i)] = mdl_inputs['VW_passes3'].copy()\n#     mdl_inputs['VW_passes3_thrs{}'.format(i)][1] = {'pos_threshold':i/100.}","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"829498bab9ca1c4611e8b1c18bd43cb841614009"},"cell_type":"markdown","source":"The function to do training in a cross-validation loop and evaluate performance of the model"},{"metadata":{"trusted":true,"_uuid":"30e13ea3dd7988d821384b1f4ed65f61fa34aec7","_kg_hide-input":true},"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold, KFold, GroupKFold\nfrom sklearn.base import clone, ClassifierMixin, RegressorMixin\n\ndef train_single_model(clf_, X_, y_, random_state_=314, opt_parameters_={}, fit_params_={}):\n    '''\n    A wrapper to train a model with particular parameters\n    '''\n    c = clone(clf_)\n    \n    param_dict = {}\n    if 'VW' in type(c).__name__:\n        # we need to get ALL parameters, as the VW instance is destroyed on set_params\n        param_dict = c.get_params()\n        # the threshold is lost in the cloning\n        param_dict['pos_threshold'] = clf_.pos_threshold\n        param_dict.update(opt_parameters_)\n        # the random_state is random_seed so far\n        param_dict.update({'random_seed': random_state_})\n        if hasattr(c, 'fit_'):\n            # reset VW if it has already been trained\n            c.get_vw().finish()\n            c.vw_ = None \n    else:\n        param_dict = opt_parameters_\n        param_dict['random_state'] = random_state_\n    # Set pre-configured parameters\n    c.set_params(**param_dict)\n    #print('Threshold = ',c.pos_threshold)\n    \n    return c.fit(X_, y_, **fit_params_)\n\ndef train_model_in_CV(model, X, y, metric, metric_args={},\n                            model_name='xmodel',\n                            seed=31416, n=5,\n                            opt_parameters_={}, fit_params_={},\n                            verbose=True,\n                            groups=None, \n                            y_eval=None,\n                            w_=1.):\n    # the list of classifiers for voting ensable\n    clfs = []\n    # performance \n    perf_eval = {'score_i_oof': 0,\n                 'score_i_ave': 0,\n                 'score_i_std': 0,\n                 'score_i': []\n                }\n\n    cv = KFold(n, shuffle=True, random_state=seed) #Stratified\n\n    scores = []\n    clfs = []\n\n    for n_fold, (trn_idx, val_idx) in enumerate(cv.split(X, (y!=0).astype(np.int8), groups=groups)):\n        X_trn, y_trn = X.iloc[trn_idx], y.iloc[trn_idx]\n        X_val, y_val = X.iloc[val_idx], y.iloc[val_idx]\n        X_trn_vw = to_vw(X_trn, convert_labels_sklearn_to_vw(y_trn), w=w_).values\n        X_val_vw = to_vw(X_val, convert_labels_sklearn_to_vw(y_val), w=w_).values\n\n        #display(y_trn.head())\n        clf = train_single_model(model, X_trn_vw, None, 314+n_fold, opt_parameters_, fit_params_)\n        #plt.hist(clf.decision_function(X_val_vw), bins=50)\n        \n        if 'VW' in type(clf).__name__:\n            x_thres = np.linspace(0.05, 0.95, num=37)\n            y_f1    = []\n            for thres in x_thres:\n                # predict on the validation sample\n                y_pred_tmp = (clf.decision_function(X_val_vw) > thres).astype(int)\n                y_f1.append(metric(y_val, y_pred_tmp, **metric_args))\n            i_opt = np.argmax(y_f1)\n\n            clf.pos_threshold = x_thres[i_opt]\n            #print('Optimal threshold = {:.4f}'.format(clf.pos_threshold))\n        \n        # predict on the validation sample\n        y_pred_tmp = (clf.decision_function(X_val_vw) > clf.pos_threshold).astype(int)\n        #store evaluated metric\n        scores.append(metric(y_val, y_pred_tmp, **metric_args))\n        \n        # store the model\n        clfs.append(('{}{}'.format(model_name,n_fold), clf))\n        \n        #cleanup\n        del X_trn, y_trn, X_val, y_val, y_pred_tmp, X_trn_vw, X_val_vw\n\n    #plt.show()\n    perf_eval['score_i_oof'] = 0\n    perf_eval['score_i'] = scores            \n    perf_eval['score_i_ave'] = np.mean(scores)\n    perf_eval['score_i_std'] = np.std(scores)\n\n    return clfs, perf_eval, None\n\ndef print_perf_clf(name, perf_eval, fmt='.4f'):\n    print('Performance of the model:')    \n    print('Mean(Val) score inner {} Classifier: {:{fmt}}+-{:{fmt}}'.format(name, \n                                                                       perf_eval['score_i_ave'],\n                                                                       perf_eval['score_i_std'],\n                                                                       fmt=fmt\n                                                                     ))\n    print('Min/max scores on folds: {:{fmt}} / {:{fmt}}'.format(np.min(perf_eval['score_i']),\n                                                            np.max(perf_eval['score_i']),\n                                                            fmt=fmt\n                                                           ))\n    print('OOF score inner {} Classifier: {:{fmt}}'.format(name, perf_eval['score_i_oof'], fmt=fmt))\n    print('Scores in individual folds: [{}]'\n          .format(' '.join(['{:{fmt}}'.format(c, fmt=fmt) \n                            for c in perf_eval['score_i']\n                           ])\n                 )\n         )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a1175f8605e20735bc144d557bac89eed246855f"},"cell_type":"markdown","source":"Actual training of the model"},{"metadata":{"_uuid":"d5c5dec8cf927fc188a5db17a7fb43e9ddaf6169"},"cell_type":"markdown","source":"--------------- VW1 -----------\nPerformance of the model:\nMean(Val) score inner VW1 Classifier: 0.4984+-0.0122\nMin/max scores on folds: 0.4783 / 0.5154\nOOF score inner VW1 Classifier: 0.0000\nScores in individual folds: [0.4952 0.4978 0.4783 0.5154 0.5054]\n--------------- VW1_l1_1em7 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l1_1em7 Classifier: 0.4987+-0.0123\nMin/max scores on folds: 0.4787 / 0.5161\nOOF score inner VW1_l1_1em7 Classifier: 0.0000\nScores in individual folds: [0.4956 0.4981 0.4787 0.5161 0.5050]\n--------------- VW1_l1_1em5 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l1_1em5 Classifier: 0.5009+-0.0104\nMin/max scores on folds: 0.4853 / 0.5149\nOOF score inner VW1_l1_1em5 Classifier: 0.0000\nScores in individual folds: [0.4983 0.4853 0.4964 0.5097 0.5149]\n--------------- VW1_l2_1em7 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l2_1em7 Classifier: 0.4985+-0.0121\nMin/max scores on folds: 0.4787 / 0.5154\nOOF score inner VW1_l2_1em7 Classifier: 0.0000\nScores in individual folds: [0.4952 0.4978 0.4787 0.5154 0.5054]\n--------------- VW1_l2_1em5 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l2_1em5 Classifier: 0.5037+-0.0108\nMin/max scores on folds: 0.4875 / 0.5200\nOOF score inner VW1_l2_1em5 Classifier: 0.0000\nScores in individual folds: [0.5030 0.4990 0.4875 0.5200 0.5094]\n--------------- VW2 -----------\nPerformance of the model:\nMean(Val) score inner VW2 Classifier: 0.4430+-0.0058\nMin/max scores on folds: 0.4373 / 0.4536\nOOF score inner VW2 Classifier: 0.0000\nScores in individual folds: [0.4390 0.4373 0.4444 0.4406 0.4536]\n--------------- VW2_l1_1em7 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l1_1em7 Classifier: 0.4461+-0.0059\nMin/max scores on folds: 0.4401 / 0.4569\nOOF score inner VW2_l1_1em7 Classifier: 0.0000\nScores in individual folds: [0.4401 0.4413 0.4450 0.4469 0.4569]\n--------------- VW2_l1_1em5 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l1_1em5 Classifier: 0.4768+-0.0102\nMin/max scores on folds: 0.4640 / 0.4936\nOOF score inner VW2_l1_1em5 Classifier: 0.0000\nScores in individual folds: [0.4701 0.4744 0.4640 0.4820 0.4936]\n--------------- VW2_l2_1em7 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l2_1em7 Classifier: 0.4432+-0.0070\nMin/max scores on folds: 0.4361 / 0.4555\nOOF score inner VW2_l2_1em7 Classifier: 0.0000\nScores in individual folds: [0.4361 0.4371 0.4420 0.4452 0.4555]\n--------------- VW2_l2_1em5 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l2_1em5 Classifier: 0.4501+-0.0124\nMin/max scores on folds: 0.4330 / 0.4707\nOOF score inner VW2_l2_1em5 Classifier: 0.0000\nScores in individual folds: [0.4330 0.4495 0.4439 0.4532 0.4707]"},{"metadata":{"_uuid":"1ff3f5ecd4b9529f6f23fbd2dc57465519a3afeb"},"cell_type":"markdown","source":"--------------- VW1_l1_1em4 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l1_1em4 Classifier: 0.3987+-0.0173\nMin/max scores on folds: 0.3705 / 0.4185\nOOF score inner VW1_l1_1em4 Classifier: 0.0000\nScores in individual folds: [0.4185 0.3911 0.3705 0.3993 0.4142]\n--------------- VW1_l1_1em3 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l1_1em3 Classifier: 0.1171+-0.0024\nMin/max scores on folds: 0.1132 / 0.1203\nOOF score inner VW1_l1_1em3 Classifier: 0.0000\nScores in individual folds: [0.1203 0.1169 0.1132 0.1185 0.1164]\n--------------- VW1_l2_1em4 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l2_1em4 Classifier: 0.4889+-0.0128\nMin/max scores on folds: 0.4659 / 0.5017\nOOF score inner VW1_l2_1em4 Classifier: 0.0000\nScores in individual folds: [0.4848 0.4981 0.4659 0.5017 0.4939]\n--------------- VW1_l2_1em3 -----------\nPerformance of the model:\nMean(Val) score inner VW1_l2_1em3 Classifier: 0.3869+-0.0157\nMin/max scores on folds: 0.3606 / 0.4039\nOOF score inner VW1_l2_1em3 Classifier: 0.0000\nScores in individual folds: [0.3950 0.3777 0.3606 0.3972 0.4039]\n--------------- VW2_l1_1em4 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l1_1em4 Classifier: 0.1171+-0.0024\nMin/max scores on folds: 0.1132 / 0.1203\nOOF score inner VW2_l1_1em4 Classifier: 0.0000\nScores in individual folds: [0.1203 0.1169 0.1132 0.1185 0.1164]\n--------------- VW2_l1_1em3 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l1_1em3 Classifier: 0.1171+-0.0024\nMin/max scores on folds: 0.1132 / 0.1203\nOOF score inner VW2_l1_1em3 Classifier: 0.0000\nScores in individual folds: [0.1203 0.1169 0.1132 0.1185 0.1164]\n--------------- VW2_l2_1em4 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l2_1em4 Classifier: 0.4441+-0.0153\nMin/max scores on folds: 0.4260 / 0.4705\nOOF score inner VW2_l2_1em4 Classifier: 0.0000\nScores in individual folds: [0.4367 0.4501 0.4260 0.4370 0.4705]\n--------------- VW2_l2_1em3 -----------\nPerformance of the model:\nMean(Val) score inner VW2_l2_1em3 Classifier: 0.2886+-0.0157\nMin/max scores on folds: 0.2693 / 0.3157\nOOF score inner VW2_l2_1em3 Classifier: 0.0000\nScores in individual folds: [0.2866 0.3157 0.2783 0.2932 0.2693]\n"},{"metadata":{"trusted":true,"_uuid":"eda8f9ceedab4066d34dc74133b03f8e1a550e55"},"cell_type":"code","source":"%%time\nfrom sklearn.metrics import f1_score\n\nmdls = {}\nresults = {}\ny_oofs = {}\nfor name, (mdl, mdl_pars, fit_pars, y_, g_, w_) in mdl_inputs.items():\n    print('--------------- {} -----------'.format(name))\n    mdl_, perf_eval_, y_oof_ = train_model_in_CV(mdl, X_trn2.iloc[:],\n                                                  y_trn.iloc[:], f1_score, \n                                                  metric_args={},\n                                                  model_name=name, \n                                                  opt_parameters_=mdl_pars,\n                                                  fit_params_=fit_pars, \n                                                  n=5,\n                                                  verbose=500, \n                                                  groups=g_, \n                                                  y_eval=None if 'LGBMRanker' not in type(mdl).__name__ else y_rnk_eval,\n                                                  w_=w_\n                                                )\n    results[name] = perf_eval_\n    mdls[name] = mdl_\n    y_oofs[name] = y_oof_\n    print_perf_clf(name, perf_eval_)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bf2ca1da8961e67390bde97254187a14fff00b6f"},"cell_type":"markdown","source":"## Prepare submission"},{"metadata":{"_uuid":"1065df7abccbcdc49af6ccffe4546651d250707f"},"cell_type":"markdown","source":"Models, that were trained on k sets of k-1 folds are averaged before application of the decision threshold"},{"metadata":{"trusted":true,"_uuid":"0a49c9f1123c393a78706cba6bca808dbd9b7254"},"cell_type":"code","source":"%%time\ny_subs= {}\nX_tst_vw = to_vw(X_tst, None).values\nfor c in mdl_inputs:\n    mdls_= mdls[c]\n    y_sub = np.zeros(X_tst_vw.shape[0])\n    for mdl_ in mdls_:\n        y_sub += mdl_[1].decision_function(X_tst_vw)\n    y_sub /= len(mdls_)\n    \n    y_subs[c] = y_sub","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d42b17d3fffadf519d4e0415f53638a78c098e19"},"cell_type":"code","source":"df_sub = pd.read_csv('../input/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cdedab5e0ed6977ba3728a49ade1e9c1d439c204"},"cell_type":"code","source":"s= 'VW1'#VW_passes3_w\ndf_sub['prediction'] = (y_subs[s] > np.median([mdl_[1].pos_threshold for mdl_ in mdls[s]])).astype(int)\ndf_sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a774989b21f4b26b3876b95e2927103d210e305a"},"cell_type":"code","source":"!head submission.csv","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ce2e3a6f847847d626e277c35731cc8619aff160"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}