{"cells":[{"metadata":{"_uuid":"fd76c67d05fd6acd901ab6c02538fad60daf4421"},"cell_type":"markdown","source":"This is a bare bones kernel that shows that even a very simple NN pipeline is non deterministic on kernels. \n\nThis is without any Cudnn usage so that is not the reason\n\n Once everything is seeded properly I am unable to get non deterministic behaviour locally. So it must be something in the kernel environment.\n \n Even Cudnn behaves deterministically in my local environment.  Would love if others can verify running this kernel locally on a GPU.\n \n Things to keep in mind when you try to make your code deterministic. Cudnn has 2 initializers that need to be seeded. Also .fit has shuffle by default so I turn that off and shuffle manually. Also dropouts take a seed value. With all these things I can get deterministic behaviour (locally) even when running a loop to try hyper parameter changes.\n\n For a bit I thought it was the versioning since I thought I had non deterministic behaviour locally on 2.2.4 but I have not been able to duplicate that so seems to be a deadend. At this point I am throwing in the towel. Maybe someone else has some more ideas on the cause and remedy. At least we should stop blamming Cudnn.\n\n The biggest issue with this is that they presumablly will only be running these one time for final submissions. With this much variance I think it will put alot of luck in play.**\n"},{"metadata":{"trusted":true,"_uuid":"9292863ac4c5966abee9869784c10cfe20feea60"},"cell_type":"code","source":"#import sys\n#!{sys.executable} -m conda install tensorflow-gpu -y\n#!{sys.executable} -m pip install keras==2.1.4","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"53c1d2041dfe310060d5f7e00990b736093dbc32"},"cell_type":"code","source":"import keras\nimport tensorflow as tf\nprint(keras.__version__)\nprint(tf.__version__)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e31d6e126881ee56a1de3efe02fcf309e900ef00"},"cell_type":"code","source":"## some config values \nembed_size = 300 # how big is each word vector\nmax_features = 50000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 50 # max number of words in a question to use\n\nRANDOM_SEED = 42\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"522d9790478f62193ea5c315372a2ab9cbe9b27f"},"cell_type":"markdown","source":"**Load packages and data**"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import os\nimport time\nimport numpy as np # linear algebra\nimport random as rn\n\nnp.random.seed(RANDOM_SEED)\nrn.seed(RANDOM_SEED)\n\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nimport math\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\nfrom sklearn.model_selection import GridSearchCV, StratifiedKFold\nfrom sklearn.metrics import f1_score, roc_auc_score\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import *\nfrom keras.optimizers import *\nfrom keras.models import Model\nfrom keras import backend as K\nfrom keras.engine.topology import Layer\nfrom keras import initializers, regularizers, constraints, optimizers, layers\nfrom keras.callbacks import *\nfrom keras.initializers import *\n\nfrom tensorflow import set_random_seed\n \nset_random_seed(RANDOM_SEED)\n\nsession_conf = tf.ConfigProto(intra_op_parallelism_threads=1, inter_op_parallelism_threads=1)\nsess = tf.Session(graph=tf.get_default_graph(), config=session_conf)\nK.set_session(sess)\n\nfrom tensorflow.python.client import device_lib\ndevice_lib.list_local_devices()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5cdc95950037613c690c49b27930ae0f59eb23c3"},"cell_type":"code","source":"def load_and_prec():\n    train_df = pd.read_csv(\"../input/train.csv\")\n    test_df = pd.read_csv(\"../input/test.csv\")\n    print(\"Train shape : \",train_df.shape)\n    print(\"Test shape : \",test_df.shape)\n    \n    ## fill up the missing values\n    train_X = train_df[\"question_text\"].fillna(\"_##_\").values\n    test_X = test_df[\"question_text\"].fillna(\"_##_\").values\n\n    ## Tokenize the sentences\n    tokenizer = Tokenizer(num_words=max_features)\n    tokenizer.fit_on_texts(list(train_X))\n    train_X = tokenizer.texts_to_sequences(train_X)\n    test_X = tokenizer.texts_to_sequences(test_X)\n\n    ## Pad the sentences \n    train_X = pad_sequences(train_X, maxlen=maxlen)\n    test_X = pad_sequences(test_X, maxlen=maxlen)\n\n    ## Get the target values\n    train_y = train_df['target'].values\n    \n    #shuffling the data\n    np.random.seed(RANDOM_SEED)\n    trn_idx = np.random.permutation(len(train_X))\n\n    train_X = train_X[trn_idx]\n    train_y = train_y[trn_idx]\n    \n    return train_X, test_X, train_y, tokenizer.word_index","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dba1893c267a1e7536bbf720636647d85c7e349c"},"cell_type":"markdown","source":"**Load embeddings**"},{"metadata":{"trusted":true,"_uuid":"a662716cc5fbbcc0c84019a87c52332ed8912e8d"},"cell_type":"code","source":"def load_glove(word_index):\n    EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding='utf-8'))\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    np.random.seed(RANDOM_SEED)\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix \n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d96793d88c22274d985436e192f62970c227c324"},"cell_type":"markdown","source":"**LSTM models**"},{"metadata":{"trusted":true,"_uuid":"05164d541a0c35cae727d0338548d156efe21427"},"cell_type":"code","source":"def get_model(embedding_matrix):\n    \n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    #x = SpatialDropout1D(0.1, seed = RANDOM_SEED)(x)\n    \n    #x = Bidirectional(CuDNNLSTM(40, return_sequences=True, kernel_initializer=glorot_uniform(seed = RANDOM_SEED), \\\n    #                                recurrent_initializer=Orthogonal(seed = RANDOM_SEED)))(x)\n        \n    avg_pool = GlobalAveragePooling1D()(x)\n    max_pool = GlobalMaxPooling1D()(x)\n    \n    conc = concatenate([avg_pool, max_pool])\n    \n    x = Dense(16, activation=\"relu\", kernel_initializer=he_uniform(seed=RANDOM_SEED))(conc)\n    #x = Dropout(0.1, seed = RANDOM_SEED)(x)\n    outp = Dense(1, activation=\"sigmoid\", kernel_initializer=he_uniform(seed=RANDOM_SEED))(x)    \n\n    model = Model(inputs=inp, outputs=outp)\n    optimizer = Adam()\n    model.compile(loss='binary_crossentropy', optimizer=optimizer, metrics=['accuracy'])\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a8c857424e9c9f1703a71c1c0ade28713314dd29"},"cell_type":"markdown","source":"**Train and predict**"},{"metadata":{"trusted":true,"_uuid":"e8523d876b6eae762e673b777cc7af4d7f085792"},"cell_type":"code","source":"# https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\ndef train_pred(model, train_X, train_y, val_X, val_y, epochs=2, callback=None):\n    for e in range(epochs):\n        \n        np.random.seed(RANDOM_SEED + e)\n        trn_idx = np.random.permutation(len(train_X))\n        #print('trn_idx[:1000].sum()', trn_idx[:10000].sum())\n        X = train_X[trn_idx]\n        y = train_y[trn_idx]\n        \n        model.fit(X, y, batch_size=512, epochs=1, validation_data=(val_X, val_y), callbacks = callback, verbose=0, shuffle = False)\n        pred_val_y = model.predict([val_X], batch_size=1024, verbose=0)\n        \n        best_score = 0.0\n        best_score = metrics.f1_score(val_y, (pred_val_y > 0.33).astype(int))\n        print(\"Epoch: \", e, \"-    Val F1 Score: {:.4f}\".format(best_score))\n\n    pred_test_y = model.predict([test_X], batch_size=1024, verbose=0)\n    print('='*100)\n    return pred_val_y, pred_test_y, best_score","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f79081928ca032fbfe3b90c6d3ce91cf57d443d8"},"cell_type":"markdown","source":"**Main part: load, train, pred and blend**"},{"metadata":{"trusted":true,"_uuid":"99d03d2eb63600f1b222522616eab3fa35819f37"},"cell_type":"code","source":"train_X, test_X, train_y, word_index = load_and_prec()\nembedding_matrix = load_glove(word_index)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d951e5240c66363ba6148658c05fb73247276d66"},"cell_type":"code","source":"epochs = 5","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1e700ba1ef2301b99ff37bbfe7d8bfd39c9a2dd7"},"cell_type":"code","source":"%%time\nset_random_seed(RANDOM_SEED)\n\nsplits = list(StratifiedKFold(n_splits=4, shuffle=True, random_state=RANDOM_SEED).split(train_X, train_y))\nfor idx, (train_idx, valid_idx) in list(enumerate(splits))[:1]:\n        X_train = train_X[train_idx]\n        y_train = train_y[train_idx]\n        X_val = train_X[valid_idx]\n        y_val = train_y[valid_idx]\n        model = get_model(embedding_matrix)\n        \n        pred_val_y, pred_test_y, best_score = train_pred(model, X_train, y_train, X_val, y_val, epochs = epochs)\n        \n        for l in model.layers:\n            for w in l.get_weights():\n                print(w.shape, w.sum())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d05ed5b42434dfec8b85ff350a3e1eea4d3e6680"},"cell_type":"markdown","source":"# This should output the same weights if everything is deterministic. \n# This works fine for me on local kernels even after adding more depth/epochs/Cudnn etc\n\nThis actually just ran deterministically for the first time I have seen in many many tries. Right before this and I was going to publish it. Good thing I checked."},{"metadata":{"trusted":true,"_uuid":"9a22ad6a4aded7b93cc95b5b1d5853a66a9892b7"},"cell_type":"code","source":"set_random_seed(RANDOM_SEED)\n\nsplits = list(StratifiedKFold(n_splits=4, shuffle=True, random_state=RANDOM_SEED).split(train_X, train_y))\nfor idx, (train_idx, valid_idx) in list(enumerate(splits))[:1]:\n        X_train = train_X[train_idx]\n        y_train = train_y[train_idx]\n        X_val = train_X[valid_idx]\n        y_val = train_y[valid_idx]\n        model = get_model(embedding_matrix)\n        \n        pred_val_y, pred_test_y, best_score = train_pred(model, X_train, y_train, X_val, y_val, epochs = epochs)\n        \n        for l in model.layers:\n            for w in l.get_weights():\n                print(w.shape, w.sum())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e8e7dd69375fe041ddd59f56d241eb3c9066b109"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}