{"cells":[{"metadata":{},"cell_type":"markdown","source":"#### The following steps would help you get started. To **RUN THIS CODE HERE**, you need to *fork* a new branch (see the blue button on the top-right corner), and then execute each cell by pressing *Shift+Enter* or clicking on the blue botton on the left side of each cell.\n\n* Let's import a few handy toolds."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom json import JSONDecoder, JSONDecodeError  # for reading the JSON data files\nimport re  # for regular expressions\nimport os  # for os related operations","execution_count":1,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Input data files are available in the '../input/' directory.\nAny results you write to the current directory will be saved here as output.\n\n* We can list all files in this directory:"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(os.listdir(\"../input\"))","execution_count":2,"outputs":[{"output_type":"stream","text":"['fold1Training.json', 'fold3Training.json', 'fold2Training.json', 'testSet.json', 'sampleSubmission.csv']\n","name":"stdout"}]},{"metadata":{},"cell_type":"markdown","source":"* To be able to read the data in json format, we need to have a decoder as follows:"},{"metadata":{"trusted":true},"cell_type":"code","source":"def decode_obj(line, pos=0, decoder=JSONDecoder()):\n    no_white_space_regex = re.compile(r'[^\\s]')\n    while True:\n        match = no_white_space_regex.search(line, pos)\n        if not match:\n            return\n        pos = match.start()\n        try:\n            obj, pos = decoder.raw_decode(line, pos)\n        except JSONDecodeError as err:\n            print('Oops! something went wrong. Error: {}'.format(err))\n        yield obj","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* As an example, let's implement a method that gets the last values of a multi-variate time series corresponding to each observation window."},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_obj_with_last_n_val(line, n):\n    obj = next(decode_obj(line))  # type:dict\n    id = obj['id']\n    class_label = obj['classNum']\n\n    data = pd.DataFrame.from_dict(obj['values'])  # type:pd.DataFrame\n    data.set_index(data.index.astype(int), inplace=True)\n    last_n_indices = np.arange(0, 60)[-n:]\n    data = data.loc[last_n_indices]\n\n    return {'id': id, 'classType': class_label, 'values': data}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The above methods allow us to load the data as Pandas.DataFrame, or even save them in CSV format. Let's define a new method that does this. Note that you can uncomment the part that stores the data in CSV format if you want."},{"metadata":{"trusted":true},"cell_type":"code","source":"def convert_json_data_to_csv(data_dir: str, file_name: str):\n    \"\"\"\n    Generates a dataframe by concatenating the last values of each\n    multi-variate time series. This method is designed as an example\n    to show how a json object can be converted into a csv file.\n    :param data_dir: the path to the data directory.\n    :param file_name: name of the file to be read, with the extension.\n    :return: the generated dataframe.\n    \"\"\"\n    fname = os.path.join(data_dir, file_name)\n\n    all_df, labels, ids = [], [], []\n    with open(fname, 'r') as infile: # Open the file for reading\n        for line in infile:  # Each 'line' is one MVTS with its single label (0 or 1).\n            obj = get_obj_with_last_n_val(line, 1)\n            all_df.append(obj['values'])\n            labels.append(obj['classType'])\n            ids.append(obj['id'])\n\n    df = pd.concat(all_df).reset_index(drop=True)\n    df = df.assign(LABEL=pd.Series(labels))\n    df = df.assign(ID=pd.Series(ids))\n    df.set_index([pd.Index(ids)])\n    # Uncomment if you want to save this as CSV\n    # df.to_csv(file_name + '_last_vals.csv', index=False)\n    return df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Now we are ready to load data. We try loading 'fold3Training.json' as an example. This should result in a dataframe with 27006 rows and 27 columns (i.e., all 25 physical parameters, plus two additional columns: ID and LABEL)"},{"metadata":{"trusted":true},"cell_type":"code","source":"path_to_data = \"../input\"\nfile_name = \"fold3Training.json\"\n\ndf = convert_json_data_to_csv(path_to_data, file_name)  # shape: 27006 X 27\nprint('df.shape = {}'.format(df.shape))\n# print(list(df))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* There are many ways to deal with missing values. The simplest approach would be to drop all rows which contain any missing values."},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df.dropna()  # shape: 26666 X 27\nprint('df.shape = {}'.format(df.shape))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* To train a simple classifier, we first need to have training and validation sets. For simplicity, let's assign the first two-third of this fold to our training set, and use the rest as a validation set."},{"metadata":{"trusted":true},"cell_type":"code","source":"t = (2/3) * df.shape[0]\ndf_train = df[df['ID'] <= t]  # shape: 18004 X 27\ndf_val = df[df['ID'] > t]  # shape: 9002 X 27\nprint('df_train.shape = {}'.format(df_train.shape))\nprint('df_val.shape = {}'.format(df_val.shape))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Finally, time to train a model. Of course, we should import some packages first."},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn import svm\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import f1_score","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Note** that the training phase may take a few minutes."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Separate values and labels columns\ndf_train_data = df_train.iloc[:, :-2]  # all columns excluding 'ID' and 'LABEL'\ndf_train_labels = pd.DataFrame(df_train.LABEL)  # only 'LABEL' column\n\ndf_val_data = df_val.iloc[:, :-2]  # all columns excluding 'ID' and 'LABEL'\ndf_val_labels = pd.DataFrame(df_val.LABEL)  # only 'LABEL' column\n\n# Train a simple SVM as an example\nsvm_c = 1000\nsvm_gamma = 0.01\nclf = svm.SVC(gamma=svm_gamma, C=svm_c, max_iter=-1, verbose=1, shrinking=True, random_state=42)\nclf.fit(df_train_data, np.ravel(df_train_labels))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Our model is now ready for prediction. Let's see how good it performs. (We measure its performance using f1-score)."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Test the model against the validation set\npred_labels = clf.predict(df_val_data)\n\n# Evaluate the predictions\nscores = confusion_matrix(df_val_labels, pred_labels).ravel()\ntn, fp, fn, tp = scores\nprint('TN:{}\\tFP:{}\\tFN:{}\\tTP:{}'.format(tn, fp, fn, tp))\nf1 = f1_score(df_val_labels, pred_labels, average='binary', labels=[0, 1])\nprint('f1-score = {}'.format(f1))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Oops! It seems that the model is not good at all! Maybe the model needs tuning! Maybe the data needs more preprocessing! Or, maybe the \"last values\", as we used in this example code, is not such a good predicator for solar flares!\n#### You can pick up from here. It's all in your hands now."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}