{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# An Approach Without ML Libraries\n\n**How the model trains:**\n1. Find all unique words from all markdown (MD) cells in the training dataset.\n1. For each unique word, find the mean \"rank\" of all MD cells that contain that word. The rank is a value between [0,1] and represents the location of a cell within the notebook. 0 is the start and 1 is the end of the notebook.\n1. Update a dictionary with that unique word and its mean rank.\n\n**How the model predicts the rank of a given MD cell:**\n1. Find all unique words in the MD cell.\n1. For each unique word in the cell, look up its mean rank from our dictionary. If the word is not in our dictionary, ignore it.\n1. The predicted rank for this cell is the mean of all the means we looked up from the dictionary.\n1. If no words in the cell are found in the dictionary, assign the cell a random rank between 0.25 and 0.75.\n\n**How the model constructs a final ordering of a notebook:**\n1. Predict a rank for all MD cells in the notebook.\n1. Order the MD cells by predicted rank from least to greatest. Order the code cells in the same way too. (The order of the code cells is known)\n1. Construct a final ordering by alternating between MD and code cells until all cells have been accounted for.\n\n**Assumptions made by the model:**\n* The presence of unique words in a cell is the only determinant of the rank of the cell. (All context is disregarded)\n* The text within code cells is not used for predicting a final ordering. (We only clean the text from code cells for EDA purposes)\n* Each unique word in a cell has an equal influence on the cell's rank.\n* The frequency or location of a word in a cell does not impact the cell's rank.\n\n**Other limitations:**\n* My code for extracting the most useful words from MD and code cells is not perfect.\n* The model seems slow to train.\n\n**Next steps:**\n* Explore the relationship between the total number of words in a cell and the cell's rank. Perhaps shorter cells often go somewhere other than longer cells.\n* Perfect the code for cleaning the cell text.","metadata":{}},{"cell_type":"code","source":"import re\nimport json \nimport string\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport random\nfrom pathlib import Path\nfrom tqdm import tqdm\nfrom nltk.corpus import stopwords\n\nstop_words = set(stopwords.words('english'))\n_RE_COMBINE_WHITESPACE = re.compile(r\"(?a:\\s+)\")\n\nletters = string.ascii_uppercase\ndigits = string.digits\npunctuation = string.punctuation\nwhitespace = string.whitespace","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_files = 2000\ntrain_size = 0.8\n\nTRAIN_ORDERS_PATH = '/kaggle/input/AI4Code/train_orders.csv'\nTRAIN_DATA_PATH = '/kaggle/input/AI4Code/train'\nTEST_DATA_PATH = '/kaggle/input/AI4Code/test'\n\ntrain_orders = pd.read_csv(TRAIN_ORDERS_PATH, index_col=\"id\")\ntrain_orders = train_orders.head(n_files)\n\nnum_train = len(train_orders)\nvalidation_files = train_orders[int(num_train*train_size):].index\n\ntrain_orders.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T13:29:16.544168Z","iopub.execute_input":"2022-08-02T13:29:16.544568Z","iopub.status.idle":"2022-08-02T13:29:17.379364Z","shell.execute_reply.started":"2022-08-02T13:29:16.544532Z","shell.execute_reply":"2022-08-02T13:29:17.378138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Functions to Evaluate Model Predictions","metadata":{}},{"cell_type":"code","source":"from bisect import bisect\n\ndef count_inversions(a):\n    inversions = 0\n    sorted_so_far = []\n    for i, u in enumerate(a):\n        j = bisect(sorted_so_far, u)\n        inversions += i - j\n        sorted_so_far.insert(j, u)\n    return inversions\n\ndef kendall_tau(ground_truth, predictions):\n    total_inversions = 0\n    total_2max = 0  # twice the maximum possible inversions across all instances\n    for gt, pred in zip(ground_truth, predictions):\n        ranks = [gt.index(x) for x in pred]  # rank predicted order in terms of ground truth\n        total_inversions += count_inversions(ranks)\n        n = len(gt)\n        total_2max += n * (n - 1)\n    return 1 - 4 * total_inversions / total_2max","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:15:51.935291Z","iopub.execute_input":"2022-08-02T12:15:51.935677Z","iopub.status.idle":"2022-08-02T12:15:51.946161Z","shell.execute_reply.started":"2022-08-02T12:15:51.935644Z","shell.execute_reply":"2022-08-02T12:15:51.944861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Functions to Clean Cell Text\n\n**Note:** These functions cleanse both code and markdown cells, but for my model to work, you only need to cleanse the markdown cells.","metadata":{}},{"cell_type":"code","source":"def flatten(xss):\n    return [x for xs in xss for x in xs]\n\ndef cleanse_str(s):\n    \n    if s is None:\n        return s\n    \n    s = re.sub('\\s+',' ',s)\n    \n    # Remove punctuation, then digits, then tabs etc\n    s_clean = \"\".join(['' if l in punctuation else l for l in s.lower()])\n    s_clean = \"\".join(['' if l in digits else l for l in s_clean])\n    s_clean = \"\".join([' ' if l in '\\t\\n\\r\\x0b\\x0c/' else l for l in s_clean])\n    \n    return s_clean\n\ndef cleanse_code(code):\n    \n    cleanCode = \"\".join([' ' if l in '\\t\\n\\r\\x0b\\x0c/' else l for l in code])\n    \n    code_split = cleanCode.split(' ')\n    words = []\n    \n    # Add all imported libraries to 'words'\n    for i in range(len(code_split)-1):\n        if (code_split[i] == 'from') |  (code_split[i] == 'import'):\n            for word in code_split[i+1].split('.'):\n                words.append(word)\n    \n    # Add all functions to 'words'\n    for i in range(len(code_split)):\n        if '.' in code_split[i]:\n            for s in code_split[i].split('.'):\n                for word in s.split('('):\n                    words.append(cleanse_str(word))\n        # Add new variable names when they are defined\n        elif '=' in code_split[i]:\n            words.append(code_split[i].split('=')[0])\n    \n    # Remove all empty strings (\"\")\n    for _ in range(words.count(\"\")):\n        words.remove(\"\")\n    \n    return [word for word in words if len(word) > 1]\n\ndef cleanse_markdown(cell):\n    \n    # Use cleanse function on all markdown blocks to find count of each word used, whose length>=3\n    \n    cell = cleanse_str(cell).split(' ')\n    cell = [w if len(w) >= 3 else \"\" for w in cell]\n    cell = [w if not w in stop_words else \"\" for w in cell]\n\n    # Remove all empty strings (\"\")\n    for _ in range(cell.count(\"\")):\n            cell.remove(\"\")\n    \n    return cell\n\ndef add_cleansed_column(df):\n    \n    # This function creates a new column in your DataFrame for all the cleansed text to go.\n    \n    markdown_ix = df[\"cell_type\"] == \"markdown\"\n    code_ix = df['cell_type'] == 'code'\n    \n    df['source_cleansed'] = ['' for i in range(len(df))]\n\n    df.loc[markdown_ix, 'source_cleansed'] = df.loc[markdown_ix, 'source'].apply(cleanse_markdown)\n    \n    df.loc[code_ix, 'source_cleansed'] = df.loc[code_ix, 'source'].apply(cleanse_code)\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-02T14:05:55.204086Z","iopub.execute_input":"2022-08-02T14:05:55.204486Z","iopub.status.idle":"2022-08-02T14:05:55.220878Z","shell.execute_reply.started":"2022-08-02T14:05:55.204453Z","shell.execute_reply":"2022-08-02T14:05:55.219746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_notebook(path, is_train_set=False):\n    \n    df = (pd.read_json(path, dtype={'cell_type': 'category', 'source': 'str'})\n          .rename_axis('id'))\n          \n    if is_train_set:\n        df = df.reindex(train_orders.loc[path.stem, 'cell_order'].split(' ')) # Order the cells correctly\n        df = df.reset_index()\n        df['cell_rank'] = df.index / df.index.max()    \n    else:\n        df = df.reset_index()\n    \n    file_id = pd.Series(path.stem, index=range(len(df)))\n    \n    df = pd.concat([file_id, df], axis=1)\n    \n    df.columns.values[0] = 'file_id'\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:15:51.971309Z","iopub.execute_input":"2022-08-02T12:15:51.971786Z","iopub.status.idle":"2022-08-02T12:15:51.987371Z","shell.execute_reply.started":"2022-08-02T12:15:51.971739Z","shell.execute_reply":"2022-08-02T12:15:51.986073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read training set and cleanse the cell text","metadata":{}},{"cell_type":"code","source":"# Get train data and cleanse its cells\n\npaths_train = list((Path(TRAIN_DATA_PATH)).glob('*.json'))\n\nnotebooks_train = []\n\nnotebook_ix = 0\n\nfor ix, path in enumerate(tqdm(paths_train, desc='Test NBs')):\n    \n    if path.stem in train_orders.index:    \n        notebooks_train.append(read_notebook(path, is_train_set = True))\n    \ndf_train = pd.concat(notebooks_train, ignore_index=True)\ndf_train['is_validation'] = df_train['file_id'].apply(lambda x: 1 if x in validation_files else 0)\ndf_train = add_cleansed_column(df_train.sort_values('is_validation'))\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-08-02T14:05:57.254136Z","iopub.execute_input":"2022-08-02T14:05:57.255637Z","iopub.status.idle":"2022-08-02T14:06:26.554083Z","shell.execute_reply.started":"2022-08-02T14:05:57.255518Z","shell.execute_reply":"2022-08-02T14:06:26.552916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get test data and cleanse its cells\n\npaths_test = list((Path(TEST_DATA_PATH)).glob('*.json'))\n\nnotebooks_test = [\n    read_notebook(path) for path in tqdm(paths_test, desc='Test NBs')\n]\n\ndf_test = pd.concat(notebooks_test, ignore_index=True)\ndf_test = add_cleansed_column(df_test)\ndf_test","metadata":{"execution":{"iopub.status.busy":"2022-08-02T14:06:26.555895Z","iopub.execute_input":"2022-08-02T14:06:26.556225Z","iopub.status.idle":"2022-08-02T14:06:26.640349Z","shell.execute_reply.started":"2022-08-02T14:06:26.556194Z","shell.execute_reply":"2022-08-02T14:06:26.639055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Analysis","metadata":{}},{"cell_type":"code","source":"cell_types = ['markdown', 'code']\n\nplt.rcParams.update({'font.size': 22})\nfig,ax = plt.subplots(1,1, figsize=[14, 9])\n\nfor cell_type in cell_types:\n    cell_ranks = df_train.loc[df_train['cell_type'] == cell_type, 'cell_rank']\n    sns.kdeplot(cell_ranks).set(title=f'Density of cells')\n\nplt.legend(labels = cell_types)\nplt.xlim(0,1)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:16:32.831712Z","iopub.execute_input":"2022-08-02T12:16:32.832073Z","iopub.status.idle":"2022-08-02T12:16:33.536035Z","shell.execute_reply.started":"2022-08-02T12:16:32.832039Z","shell.execute_reply":"2022-08-02T12:16:33.534792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose a word and plot its distribution throughout code and markdown cells\n\nword = 'plot'\n\ncell_word_ranks = []\nfig,ax = plt.subplots(1,1, figsize=[14, 9])\n\nfor ix, cell_type in enumerate(cell_types):\n    word_ix = df_train.loc[df_train['cell_type'] == cell_type, 'source_cleansed'].apply(lambda x: word in x)\n    cell_word_ranks.append(df_train.loc[(df_train['cell_type'] == cell_type) & word_ix, 'cell_rank'])\n    sns.kdeplot(cell_word_ranks[ix]).set(title=f'Distribution of cells containing the word \\'{word}\\'')\n\nplt.legend(labels = [cell_types[ix] for ix in range(len(cell_types)) if not cell_word_ranks[ix].empty])\nplt.xlim((0, 1))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:37:54.448087Z","iopub.execute_input":"2022-08-02T12:37:54.449275Z","iopub.status.idle":"2022-08-02T12:37:54.871123Z","shell.execute_reply.started":"2022-08-02T12:37:54.449232Z","shell.execute_reply":"2022-08-02T12:37:54.869972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split train set between train and validation","metadata":{}},{"cell_type":"code","source":"is_validation_ix = df_train['is_validation'] == 1\n\nX_train, X_valid = df_train[~is_validation_ix], df_train[is_validation_ix]\n\nvalidation_orders = X_valid.groupby(X_valid['file_id'])['id'].apply(list)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T13:33:05.325898Z","iopub.execute_input":"2022-08-02T13:33:05.32636Z","iopub.status.idle":"2022-08-02T13:33:05.372671Z","shell.execute_reply.started":"2022-08-02T13:33:05.326321Z","shell.execute_reply":"2022-08-02T13:33:05.371427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create dictionary to power the model","metadata":{}},{"cell_type":"code","source":"dictionary_md = {}\n\nX_train_md = X_train[X_train['cell_type'] == 'markdown']\n\nunique_words_md = pd.Series(flatten(X_train_md['source_cleansed'])).unique()\n    \nfor word in unique_words_md:\n    cells_containing_word = X_train_md['source_cleansed'].apply(lambda x: word in x)\n    mean = np.mean(X_train_md.loc[cells_containing_word, 'cell_rank'])\n    dictionary_md.update({word: mean})","metadata":{"execution":{"iopub.status.busy":"2022-08-02T13:33:06.585896Z","iopub.execute_input":"2022-08-02T13:33:06.586269Z","iopub.status.idle":"2022-08-02T13:47:09.292857Z","shell.execute_reply.started":"2022-08-02T13:33:06.586237Z","shell.execute_reply":"2022-08-02T13:47:09.29154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Functions to predict a cell's rank and construct final ordering of cells","metadata":{}},{"cell_type":"code","source":"# Predict rank between 0 and 1 for each cell\ndef predict_rank(df):\n    fy = []\n    \n    for ix, row in df.iterrows():\n        unique_words = pd.Series([word for word in row['source_cleansed']]).unique()\n        is_md = row['cell_type'] == 'markdown'\n\n        # Look up words in dictionary\n        word_ranks = 0\n        num_words_in_dic = 0\n\n        if is_md:\n            for word in unique_words:\n                if word in dictionary_md.keys():\n                    word_ranks += dictionary_md[word]\n                    num_words_in_dic += 1\n        else:\n            for word in unique_words:\n                if word in dictionary_code.keys():\n                    word_ranks += dictionary_code[word]\n                    num_words_in_dic += 1\n\n        if num_words_in_dic > 0:\n            rank_pred = word_ranks / num_words_in_dic\n        else:\n            rank_pred = random.uniform(0.25, 0.75)\n\n        fy.append(rank_pred)\n        \n    return fy\n\n# Construct a final predicted order of the cells\ndef predict_order(df):\n    \n    df_md_cells = df[df['cell_type'] == 'markdown'].copy()\n    \n    df_md_cells['rank_pred'] = predict_rank(df_md_cells)\n    \n    df_code_orders = (df[df['cell_type'] == 'code']\n                      .groupby('file_id')['id']\n                      .apply(list))\n\n    df_md_orders_preds = (df_md_cells\n                          .sort_values(['file_id', 'rank_pred'])\n                          .groupby('file_id')['id']\n                          .apply(list))\n\n    df_orders_preds = []\n\n    for file_id in df_code_orders.index:\n        \n        new_pred = []\n        \n        num_cells_in_file = len(df_code_orders[file_id]) + len(df_md_orders_preds[file_id])\n        \n        while len(new_pred) < (num_cells_in_file):\n            \n            if len(df_md_orders_preds[file_id]) > 0:\n                new_pred.append(df_md_orders_preds[file_id].pop(0))\n\n            if len(df_code_orders[file_id]) > 0:\n                new_pred.append(df_code_orders[file_id].pop(0))\n        \n        df_orders_preds.append(new_pred)\n\n    df_orders_preds = pd.Series(df_orders_preds, index = df_code_orders.index)\n    \n    return df_orders_preds\n","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:32:54.84861Z","iopub.execute_input":"2022-08-02T12:32:54.848973Z","iopub.status.idle":"2022-08-02T12:32:54.86177Z","shell.execute_reply.started":"2022-08-02T12:32:54.848942Z","shell.execute_reply":"2022-08-02T12:32:54.86087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test model on validation set","metadata":{}},{"cell_type":"code","source":"validation_preds = predict_order(X_valid)\n\nkendall_tau(validation_orders, validation_preds)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T13:47:09.294931Z","iopub.execute_input":"2022-08-02T13:47:09.295288Z","iopub.status.idle":"2022-08-02T13:47:10.99193Z","shell.execute_reply.started":"2022-08-02T13:47:09.295255Z","shell.execute_reply":"2022-08-02T13:47:10.990694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make Submission","metadata":{}},{"cell_type":"code","source":"# Make preds\n\ndf_preds = predict_order(df_test)\ndf_preds","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:41:34.140576Z","iopub.execute_input":"2022-08-02T12:41:34.141009Z","iopub.status.idle":"2022-08-02T12:41:34.172492Z","shell.execute_reply.started":"2022-08-02T12:41:34.140977Z","shell.execute_reply":"2022-08-02T12:41:34.171327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert preds to correct format\n\ny_submit = (\n    df_preds\n    .apply(' '.join)  # list of ids -> string of ids\n    .rename_axis('id')\n    .rename('cell_order')\n)\ny_submit","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:41:34.174545Z","iopub.execute_input":"2022-08-02T12:41:34.175004Z","iopub.status.idle":"2022-08-02T12:41:34.184116Z","shell.execute_reply.started":"2022-08-02T12:41:34.174969Z","shell.execute_reply":"2022-08-02T12:41:34.182845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_submit.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:41:34.185221Z","iopub.execute_input":"2022-08-02T12:41:34.185566Z","iopub.status.idle":"2022-08-02T12:41:34.196849Z","shell.execute_reply.started":"2022-08-02T12:41:34.185526Z","shell.execute_reply":"2022-08-02T12:41:34.195897Z"},"trusted":true},"execution_count":null,"outputs":[]}]}