{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Google AI4Code Reconstruct the order of Kaggle notebooks\n\n![Kaggle Python Notebook](https://corngk.github.io/images/KagglePython.png)\n<br>\n\nIn this notebook we are learning how machine learning can be used to solve the challenge in this competition which to reconstruct Kaggle notebooks whose cells have been shuffled.\n\nI use the idea and approach a lot from this awesome notebook [Getting Started with AI4Code](https://www.kaggle.com/code/ryanholbrook/getting-started-with-ai4code/notebook) so we can learn how to achieve the goal. Please read the [Competition Pages](https://www.kaggle.com/competitions/AI4Code/overview) for the detail of this competition. We will add ideas to the notebook a long the way as we are learning from it.","metadata":{}},{"cell_type":"markdown","source":"# Setup","metadata":{}},{"cell_type":"markdown","source":"First thing first. We need to load the required libraries and given data.","metadata":{}},{"cell_type":"code","source":"import json\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nfrom scipy import sparse\nfrom tqdm import tqdm\n\npd.options.display.width = 180\npd.options.display.max_colwidth = 120\n\ndata_dir = Path('../input/AI4Code')","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T05:49:19.759314Z","iopub.execute_input":"2022-07-31T05:49:19.759656Z","iopub.status.idle":"2022-07-31T05:49:19.766537Z","shell.execute_reply.started":"2022-07-31T05:49:19.759619Z","shell.execute_reply":"2022-07-31T05:49:19.765607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data #","metadata":{}},{"cell_type":"markdown","source":"Then, let's load the given data.","metadata":{}},{"cell_type":"code","source":"NUM_TRAIN = 10000\n\ndef read_notebook(path):\n    return (\n        pd.read_json(\n            path,\n            dtype={'cell_type': 'category', 'source': 'str'})\n        .assign(id=path.stem)\n        .rename_axis('cell_id')\n    )\n\npaths_train = list((data_dir / 'train').glob('*.json'))[:NUM_TRAIN]\nnotebooks_train = [\n    read_notebook(path) for path in tqdm(paths_train, desc='Train NBs')\n]\n\ndf = (\n    pd.concat(notebooks_train)\n    .set_index('id', append=True)\n    .swaplevel()\n    .sort_index(level='id', sort_remaining=False)\n)\n\ndf","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T05:49:24.665431Z","iopub.execute_input":"2022-07-31T05:49:24.665741Z","iopub.status.idle":"2022-07-31T05:51:08.411234Z","shell.execute_reply.started":"2022-07-31T05:49:24.665705Z","shell.execute_reply":"2022-07-31T05:51:08.410206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compare the codes versus markdowns cells.","metadata":{}},{"cell_type":"code","source":"import plotly.io as pio\npio.renderers.default='notebook'\nimport plotly.express as px\n\ndf_temp = df.reset_index()\npie_data = df_temp[\"cell_type\"].value_counts().reset_index()\npie_data.columns = [\"cell_type\", \"count\"]\n\nfig = px.pie(pie_data, values='count', names='cell_type', title='Code vs Markdown')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:04.583876Z","iopub.execute_input":"2022-07-31T05:52:04.584908Z","iopub.status.idle":"2022-07-31T05:52:04.858875Z","shell.execute_reply.started":"2022-07-31T05:52:04.584850Z","shell.execute_reply":"2022-07-31T05:52:04.857692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see in the pie chart above overall, we have 34% comment cells and 66% code cells respectively.","metadata":{}},{"cell_type":"markdown","source":"## Scatter Plot","metadata":{}},{"cell_type":"code","source":"cell_analysis = df_temp.groupby([\"id\", \"cell_type\"])[\"cell_id\"].count().reset_index()\nscatter_data = pd.pivot(data=cell_analysis, index=\"id\", columns=\"cell_type\", values=\"cell_id\")\nscatter_data[\"size\"] = 30\n\nfig = px.scatter(scatter_data, x=\"code\", y=\"markdown\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:10.422204Z","iopub.execute_input":"2022-07-31T05:52:10.423191Z","iopub.status.idle":"2022-07-31T05:52:11.023147Z","shell.execute_reply.started":"2022-07-31T05:52:10.423136Z","shell.execute_reply":"2022-07-31T05:52:11.021797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that in general, Kagglers have 66% to 34% proportion between codes and comments/markdown. See the dots are more crowded under 100 codes/comments, some are between 100-300 and very limited have more than 300.\n\nIn Kaggle we developers are more focus on the codes, while for readers and visitors who many of them are note developers are more interested to the story or the markdown and comments. They are looking for the result and less interested to codes. ","metadata":{}},{"cell_type":"markdown","source":"## Distribution","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(scatter_data, x=\"code\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:31.755668Z","iopub.execute_input":"2022-07-31T05:52:31.756434Z","iopub.status.idle":"2022-07-31T05:52:31.892029Z","shell.execute_reply.started":"2022-07-31T05:52:31.756395Z","shell.execute_reply":"2022-07-31T05:52:31.890904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of notebooks have 100 codes.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(scatter_data, x=\"markdown\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:35.849872Z","iopub.execute_input":"2022-07-31T05:52:35.850802Z","iopub.status.idle":"2022-07-31T05:52:35.953888Z","shell.execute_reply.started":"2022-07-31T05:52:35.850756Z","shell.execute_reply":"2022-07-31T05:52:35.952964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of markdown have less than 50.","metadata":{}},{"cell_type":"markdown","source":"Now we will see how the code and comments are originally ordered (or disordered).","metadata":{}},{"cell_type":"code","source":"# Get an example notebook\nnb_id = df.index.unique('id')[6]\nprint('Notebook:', nb_id)\n\nprint(\"The disordered notebook:\")\nnb = df.loc[nb_id, :]\ndisplay(nb)\nprint()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:43.704441Z","iopub.execute_input":"2022-07-31T05:52:43.705061Z","iopub.status.idle":"2022-07-31T05:52:43.730867Z","shell.execute_reply.started":"2022-07-31T05:52:43.704999Z","shell.execute_reply":"2022-07-31T05:52:43.730285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ordering the Cells #","metadata":{}},{"cell_type":"code","source":"df_orders = pd.read_csv(\n    data_dir / 'train_orders.csv',\n    index_col='id',\n    squeeze=True,\n).str.split()  # Split the string representation of cell_ids into a list\n\ndf_orders.head(10)","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T05:52:48.401072Z","iopub.execute_input":"2022-07-31T05:52:48.401508Z","iopub.status.idle":"2022-07-31T05:52:51.336837Z","shell.execute_reply.started":"2022-07-31T05:52:48.401470Z","shell.execute_reply":"2022-07-31T05:52:51.335960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the correct order\ncell_order = df_orders.loc[nb_id]\n\nprint(\"The ordered notebook:\")\nnb.loc[cell_order, :]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:53.611464Z","iopub.execute_input":"2022-07-31T05:52:53.611809Z","iopub.status.idle":"2022-07-31T05:52:53.656360Z","shell.execute_reply.started":"2022-07-31T05:52:53.611773Z","shell.execute_reply":"2022-07-31T05:52:53.655671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_ranks(base, derived):\n    return [base.index(d) for d in derived]\n\ncell_ranks = get_ranks(cell_order, list(nb.index))\nnb.insert(0, 'rank', cell_ranks)\n\nnb","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:56.445821Z","iopub.execute_input":"2022-07-31T05:52:56.446129Z","iopub.status.idle":"2022-07-31T05:52:56.461715Z","shell.execute_reply.started":"2022-07-31T05:52:56.446093Z","shell.execute_reply":"2022-07-31T05:52:56.460656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pandas.testing import assert_frame_equal\n\nassert_frame_equal(nb.loc[cell_order, :], nb.sort_values('rank'))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:52:59.841224Z","iopub.execute_input":"2022-07-31T05:52:59.841535Z","iopub.status.idle":"2022-07-31T05:52:59.852928Z","shell.execute_reply.started":"2022-07-31T05:52:59.841498Z","shell.execute_reply":"2022-07-31T05:52:59.851705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_orders_ = df_orders.to_frame().join(\n    df.reset_index('cell_id').groupby('id')['cell_id'].apply(list),\n    how='right',\n)\n\nranks = {}\nfor id_, cell_order, cell_id in df_orders_.itertuples():\n    ranks[id_] = {'cell_id': cell_id, 'rank': get_ranks(cell_order, cell_id)}\n\ndf_ranks = (\n    pd.DataFrame\n    .from_dict(ranks, orient='index')\n    .rename_axis('id')\n    .apply(pd.Series.explode)\n    .set_index('cell_id', append=True)\n)\n\ndf_ranks","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:53:01.836110Z","iopub.execute_input":"2022-07-31T05:53:01.836747Z","iopub.status.idle":"2022-07-31T05:53:05.331399Z","shell.execute_reply.started":"2022-07-31T05:53:01.836706Z","shell.execute_reply":"2022-07-31T05:53:05.330361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Splits #","metadata":{}},{"cell_type":"code","source":"df_ancestors = pd.read_csv(data_dir / 'train_ancestors.csv', index_col='id')\ndf_ancestors","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T05:53:08.625479Z","iopub.execute_input":"2022-07-31T05:53:08.625803Z","iopub.status.idle":"2022-07-31T05:53:08.893602Z","shell.execute_reply.started":"2022-07-31T05:53:08.625771Z","shell.execute_reply":"2022-07-31T05:53:08.892436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GroupShuffleSplit\n\nNVALID = 0.1  # size of validation set\n\nsplitter = GroupShuffleSplit(n_splits=1, test_size=NVALID, random_state=0)\n\n# Split, keeping notebooks with a common origin (ancestor_id) together\nids = df.index.unique('id')\nancestors = df_ancestors.loc[ids, 'ancestor_id']\nids_train, ids_valid = next(splitter.split(ids, groups=ancestors))\nids_train, ids_valid = ids[ids_train], ids[ids_valid]\n\ndf_train = df.loc[ids_train, :]\ndf_valid = df.loc[ids_valid, :]","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:00:01.068758Z","iopub.execute_input":"2022-07-31T06:00:01.069098Z","iopub.status.idle":"2022-07-31T06:00:01.604722Z","shell.execute_reply.started":"2022-07-31T06:00:01.069064Z","shell.execute_reply":"2022-07-31T06:00:01.603626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:00:04.886575Z","iopub.execute_input":"2022-07-31T06:00:04.886954Z","iopub.status.idle":"2022-07-31T06:00:04.900094Z","shell.execute_reply.started":"2022-07-31T06:00:04.886920Z","shell.execute_reply":"2022-07-31T06:00:04.899426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering #\n\nLet's generate [tf-idf features](https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.TfidfTransformer.html#sklearn.feature_extraction.text.TfidfTransformer) to use with our ranking model. These features will help our model learn what kinds of words tend to occur most often at various positions within a notebook.","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_extraction.text import TfidfVectorizer\n\n# Training set\ntfidf = TfidfVectorizer(min_df=0.01)\nX_train = tfidf.fit_transform(df_train['source'].astype(str))\n# Rank of each cell within the notebook\ny_train = df_ranks.loc[ids_train].to_numpy()\n# Number of cells in each notebook\ngroups = df_ranks.loc[ids_train].groupby('id').size().to_numpy()","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:00:08.288851Z","iopub.execute_input":"2022-07-31T06:00:08.289647Z","iopub.status.idle":"2022-07-31T06:00:27.793297Z","shell.execute_reply.started":"2022-07-31T06:00:08.289609Z","shell.execute_reply":"2022-07-31T06:00:27.792173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add code cell ordering\nX_train = sparse.hstack((\n    X_train,\n    np.where(\n        df_train['cell_type'] == 'code',\n        df_train.groupby(['id', 'cell_type']).cumcount().to_numpy() + 1,\n        0,\n    ).reshape(-1, 1)\n))\nprint(X_train.shape)","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:01:54.579107Z","iopub.execute_input":"2022-07-31T06:01:54.579475Z","iopub.status.idle":"2022-07-31T06:01:54.958841Z","shell.execute_reply.started":"2022-07-31T06:01:54.579432Z","shell.execute_reply":"2022-07-31T06:01:54.958003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train #","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBRanker\n\nmodel = XGBRanker(\n    min_child_weight=10,\n    subsample=0.5,\n    tree_method='hist',\n)\nmodel.fit(X_train, y_train, group=groups)","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:02:15.103353Z","iopub.execute_input":"2022-07-31T06:02:15.103669Z","iopub.status.idle":"2022-07-31T06:02:29.540543Z","shell.execute_reply.started":"2022-07-31T06:02:15.103635Z","shell.execute_reply":"2022-07-31T06:02:29.539686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluate #","metadata":{}},{"cell_type":"markdown","source":"## Validation set ##","metadata":{}},{"cell_type":"code","source":"# Validation set\nX_valid = tfidf.transform(df_valid['source'].astype(str))\n\n# The metric uses cell ids\ny_valid = df_orders.loc[ids_valid]\n\nX_valid = sparse.hstack((\n    X_valid,\n    np.where(\n        df_valid['cell_type'] == 'code',\n        df_valid.groupby(['id', 'cell_type']).cumcount().to_numpy() + 1,\n        0,\n    ).reshape(-1, 1)\n))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:02:41.341205Z","iopub.execute_input":"2022-07-31T06:02:41.341541Z","iopub.status.idle":"2022-07-31T06:02:43.213111Z","shell.execute_reply.started":"2022-07-31T06:02:41.341507Z","shell.execute_reply":"2022-07-31T06:02:43.211934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = pd.DataFrame({'rank': model.predict(X_valid)}, index=df_valid.index)\ny_pred = (\n    y_pred\n    .sort_values(['id', 'rank'])  # Sort the cells in each notebook by their rank.\n                                  # The cell_ids are now in the order the model predicted.\n    .reset_index('cell_id')  # Convert the cell_id index into a column.\n    .groupby('id')['cell_id'].apply(list)  # Group the cell_ids for each notebook into a list.\n)\ny_pred.head(10)","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:02:50.484627Z","iopub.execute_input":"2022-07-31T06:02:50.484984Z","iopub.status.idle":"2022-07-31T06:02:50.692157Z","shell.execute_reply.started":"2022-07-31T06:02:50.484949Z","shell.execute_reply":"2022-07-31T06:02:50.691145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nb_id = df_valid.index.get_level_values('id').unique()[8]\ndisplay(df.loc[nb_id])\ndisplay(df.loc[nb_id].loc[y_pred.loc[nb_id]])","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:02:55.026224Z","iopub.execute_input":"2022-07-31T06:02:55.026570Z","iopub.status.idle":"2022-07-31T06:02:55.059753Z","shell.execute_reply.started":"2022-07-31T06:02:55.026536Z","shell.execute_reply":"2022-07-31T06:02:55.058675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Metric ##\n\nThis competition uses a variant of the [Kendall tau correlation](https://www.kaggle.com/competitions/AI4Code/overview/evaluation), which will measure how close to the correct order our predicted orderings are. See this notebook for more on this metric: [Competition Metric - Kendall Tau Correlation](https://www.kaggle.com/code/ryanholbrook/competition-metric-kendall-tau-correlation/notebook).","metadata":{}},{"cell_type":"code","source":"from bisect import bisect\n\ndef count_inversions(a):\n    inversions = 0\n    sorted_so_far = []\n    for i, u in enumerate(a):\n        j = bisect(sorted_so_far, u)\n        inversions += i - j\n        sorted_so_far.insert(j, u)\n    return inversions\n\ndef kendall_tau(ground_truth, predictions):\n    total_inversions = 0\n    total_2max = 0  # twice the maximum possible inversions across all instances\n    for gt, pred in zip(ground_truth, predictions):\n        ranks = [gt.index(x) for x in pred]  # rank predicted order in terms of ground truth\n        total_inversions += count_inversions(ranks)\n        n = len(gt)\n        total_2max += n * (n - 1)\n    return 1 - 4 * total_inversions / total_2max","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:02:59.659762Z","iopub.execute_input":"2022-07-31T06:02:59.660081Z","iopub.status.idle":"2022-07-31T06:02:59.668533Z","shell.execute_reply.started":"2022-07-31T06:02:59.660049Z","shell.execute_reply":"2022-07-31T06:02:59.667375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's test the metric with a dummy submission created from the ids of the shuffled notebooks.","metadata":{}},{"cell_type":"code","source":"y_dummy = df_valid.reset_index('cell_id').groupby('id')['cell_id'].apply(list)\nkendall_tau(y_valid, y_dummy)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:03:03.196244Z","iopub.execute_input":"2022-07-31T06:03:03.196640Z","iopub.status.idle":"2022-07-31T06:03:03.363800Z","shell.execute_reply.started":"2022-07-31T06:03:03.196579Z","shell.execute_reply":"2022-07-31T06:03:03.362689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kendall_tau(y_valid, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:03:11.936504Z","iopub.execute_input":"2022-07-31T06:03:11.936857Z","iopub.status.idle":"2022-07-31T06:03:12.041677Z","shell.execute_reply.started":"2022-07-31T06:03:11.936822Z","shell.execute_reply":"2022-07-31T06:03:12.040665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission #","metadata":{}},{"cell_type":"code","source":"paths_test = list((data_dir / 'test').glob('*.json'))\nnotebooks_test = [\n    read_notebook(path) for path in tqdm(paths_test, desc='Test NBs')\n]\ndf_test = (\n    pd.concat(notebooks_test)\n    .set_index('id', append=True)\n    .swaplevel()\n    .sort_index(level='id', sort_remaining=False)\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:03:23.938940Z","iopub.execute_input":"2022-07-31T06:03:23.939254Z","iopub.status.idle":"2022-07-31T06:03:23.993479Z","shell.execute_reply.started":"2022-07-31T06:03:23.939221Z","shell.execute_reply":"2022-07-31T06:03:23.992624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = tfidf.transform(df_test['source'].astype(str))\nX_test = sparse.hstack((\n    X_test,\n    np.where(\n        df_test['cell_type'] == 'code',\n        df_test.groupby(['id', 'cell_type']).cumcount().to_numpy() + 1,\n        0,\n    ).reshape(-1, 1)\n))","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:03:27.706665Z","iopub.execute_input":"2022-07-31T06:03:27.707110Z","iopub.status.idle":"2022-07-31T06:03:27.724783Z","shell.execute_reply.started":"2022-07-31T06:03:27.707057Z","shell.execute_reply":"2022-07-31T06:03:27.724082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_infer = pd.DataFrame({'rank': model.predict(X_test)}, index=df_test.index)\ny_infer = y_infer.sort_values(['id', 'rank']).reset_index('cell_id').groupby('id')['cell_id'].apply(list)\ny_infer","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2022-07-31T06:03:30.443839Z","iopub.execute_input":"2022-07-31T06:03:30.444361Z","iopub.status.idle":"2022-07-31T06:03:30.477249Z","shell.execute_reply.started":"2022-07-31T06:03:30.444308Z","shell.execute_reply":"2022-07-31T06:03:30.476101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_sample = pd.read_csv(data_dir / 'sample_submission.csv', index_col='id', squeeze=True)\ny_sample","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:03:33.204885Z","iopub.execute_input":"2022-07-31T06:03:33.205303Z","iopub.status.idle":"2022-07-31T06:03:33.220378Z","shell.execute_reply.started":"2022-07-31T06:03:33.205271Z","shell.execute_reply":"2022-07-31T06:03:33.219675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_submit = (\n    y_infer\n    .apply(' '.join)  # list of ids -> string of ids\n    .rename_axis('id')\n    .rename('cell_order')\n)\ny_submit","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:03:35.366051Z","iopub.execute_input":"2022-07-31T06:03:35.366343Z","iopub.status.idle":"2022-07-31T06:03:35.375423Z","shell.execute_reply.started":"2022-07-31T06:03:35.366309Z","shell.execute_reply":"2022-07-31T06:03:35.374485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's submit for the firsttime.","metadata":{}},{"cell_type":"code","source":"y_submit.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-31T06:03:38.387684Z","iopub.execute_input":"2022-07-31T06:03:38.388305Z","iopub.status.idle":"2022-07-31T06:03:38.397339Z","shell.execute_reply.started":"2022-07-31T06:03:38.388252Z","shell.execute_reply":"2022-07-31T06:03:38.396623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Thank you for visiting this notebook!\n\nThanks for reading this notebook. If you have any feedback or comments please write it down the comment section below.","metadata":{}},{"cell_type":"markdown","source":"# References\n1. [Getting Started with AI4Code by RYAN HOLBROOK](https://www.kaggle.com/code/ryanholbrook/getting-started-with-ai4code/notebook)<br>\n2. [Kendall rank correlation coefficient](https://en.wikipedia.org/wiki/Kendall_rank_correlation_coefficient)<br>\n3. [AI4Code Detailed EDA SANSKAR HASIJA](https://www.kaggle.com/code/odins0n/ai4code-detailed-eda)\n4. [Plotly Express in Python](https://plotly.com/python/plotly-express/)\n","metadata":{}}]}