{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#  1) Import libraries and packages","metadata":{"execution":{"iopub.status.busy":"2023-02-23T12:13:49.470781Z","iopub.execute_input":"2023-02-23T12:13:49.471220Z","iopub.status.idle":"2023-02-23T12:13:49.476878Z","shell.execute_reply.started":"2023-02-23T12:13:49.471183Z","shell.execute_reply":"2023-02-23T12:13:49.475743Z"}}},{"cell_type":"code","source":"!pip install autoviz\n!pip install emot\n!pip install symspellpy","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:27:31.697064Z","iopub.execute_input":"2023-02-26T20:27:31.697512Z","iopub.status.idle":"2023-02-26T20:28:07.122461Z","shell.execute_reply.started":"2023-02-26T20:27:31.697475Z","shell.execute_reply":"2023-02-26T20:28:07.121080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport emoji\nimport re\nfrom autoviz.classify_method import data_cleaning_suggestions ,data_suggestions\nfrom bs4 import BeautifulSoup\nfrom symspellpy import SymSpell, Verbosity\nimport string\nfrom emot.emo_unicode import UNICODE_EMOJI # For emojis\nfrom emot.emo_unicode import EMOTICONS_EMO # For EMOTICONS\nimport pkg_resources\nimport nltk\nfrom nltk.corpus import stopwords\nfrom sklearn.model_selection import train_test_split\n\n%matplotlib inline\nsns.set_style(\"whitegrid\")\nplt.style.use(\"fivethirtyeight\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.639140Z","iopub.execute_input":"2023-02-26T20:23:58.639962Z","iopub.status.idle":"2023-02-26T20:23:58.650820Z","shell.execute_reply.started":"2023-02-26T20:23:58.639908Z","shell.execute_reply":"2023-02-26T20:23:58.649803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2) Import and show data","metadata":{}},{"cell_type":"code","source":"# import train data into dataframe\ntrain_df = pd.read_csv('/kaggle/input/nlp-getting-started/train.csv')\n\n# import test data into dataframe\ntest_df = pd.read_csv('/kaggle/input/nlp-getting-started/test.csv')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.653367Z","iopub.execute_input":"2023-02-26T20:23:58.653868Z","iopub.status.idle":"2023-02-26T20:23:58.705044Z","shell.execute_reply.started":"2023-02-26T20:23:58.653823Z","shell.execute_reply":"2023-02-26T20:23:58.703332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.sample(7)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.706833Z","iopub.execute_input":"2023-02-26T20:23:58.707336Z","iopub.status.idle":"2023-02-26T20:23:58.721367Z","shell.execute_reply.started":"2023-02-26T20:23:58.707297Z","shell.execute_reply":"2023-02-26T20:23:58.720152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At first glance we can see that each tweet has its respecitve keywords, a location from which it was posted and the 'text' of the tweet itself. We can already see some strange combinations of words or symbols, which we are going to try and clean in the the next chapter, the Explorative Data Analysis and pre-processing.","metadata":{}},{"cell_type":"markdown","source":"# 3) EDA and Pre-processing","metadata":{"execution":{"iopub.status.busy":"2023-02-23T12:20:17.240777Z","iopub.execute_input":"2023-02-23T12:20:17.241173Z","iopub.status.idle":"2023-02-23T12:20:17.246480Z","shell.execute_reply.started":"2023-02-23T12:20:17.241140Z","shell.execute_reply":"2023-02-23T12:20:17.245174Z"}}},{"cell_type":"markdown","source":"We begin this data analysis and pre-processing part of the notebook by taking a deeper look into the dataset´s columns, their datatypes and what we can improve on each one of them for a better algorithim performance later on.  ","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.725145Z","iopub.execute_input":"2023-02-26T20:23:58.725546Z","iopub.status.idle":"2023-02-26T20:23:58.744383Z","shell.execute_reply.started":"2023-02-26T20:23:58.725508Z","shell.execute_reply":"2023-02-26T20:23:58.743065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are dealing with 3 columns that have the 'object' datatype assigned. They are string variables, and we have to optimize those in order to get better results with our NLP models. ","metadata":{}},{"cell_type":"markdown","source":"## Data cleaning ","metadata":{}},{"cell_type":"code","source":"data_cleaning_suggestions(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.770585Z","iopub.execute_input":"2023-02-26T20:23:58.771400Z","iopub.status.idle":"2023-02-26T20:23:58.854045Z","shell.execute_reply.started":"2023-02-26T20:23:58.771343Z","shell.execute_reply":"2023-02-26T20:23:58.852734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the table above we can see some pre-analysis of our columns that tells us some things about our data. We can see, for instance, the number of unique values, nulls, percentage of null and unique values, as well as some improvement suggestions.\n\nIn our case, we are not digging too much into the suggestions regarding the number or percentage of unique values, as these are expected for tweet´s location, text and keyword values. We are expecting most of the values in those columns to be different from one another, otherwise it would mean that we either have too many duplicates or many similar tweets from automated bots.     \n\n**What shuld we then clean?** \n- For location and keyword we have plenty of **null values** that can be filled so that they don´t interfere in our further calculations. ","metadata":{}},{"cell_type":"markdown","source":"### Handling null-values","metadata":{"execution":{"iopub.status.busy":"2023-02-23T12:45:20.345169Z","iopub.execute_input":"2023-02-23T12:45:20.345577Z","iopub.status.idle":"2023-02-23T12:45:20.350535Z","shell.execute_reply.started":"2023-02-23T12:45:20.345543Z","shell.execute_reply":"2023-02-23T12:45:20.349432Z"}}},{"cell_type":"markdown","source":"**Location:** for the location of the tweets, we cannot simply delete the affected rows (as they ammount to ~33% of the data), or assign the expected values per proportion. The most reasonable approach, in my opinion, would be to create a new location value with the name 'Not given', so that this can also be a value to consider in the algorithm.   \n\n","metadata":{}},{"cell_type":"code","source":"# LOCATION\n## change null values to \"Not given\"\ntrain_df.location = train_df.location.fillna('Not given')\ntest_df.location = test_df.location.fillna('Not given')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.856454Z","iopub.execute_input":"2023-02-26T20:23:58.856941Z","iopub.status.idle":"2023-02-26T20:23:58.866230Z","shell.execute_reply.started":"2023-02-26T20:23:58.856891Z","shell.execute_reply":"2023-02-26T20:23:58.864784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Keyword:** for the keyword column we take a different approach. As keywords are written information that is, similar to the text, also attached to a tweet, we can merge these two columns. This will not result in any confusion for the new column, as the text itself won´t be analyzed as a whole, but on each of its words.   ","metadata":{}},{"cell_type":"code","source":"# KEYWORD\n## merge columns 'keyword' and 'text\ntrain_df.keyword = train_df.keyword.fillna('')\ntrain_df.text = train_df.keyword + \" \" + train_df.text\n\n## drop 'keyword' column\ntrain_df = train_df.drop(['keyword'], axis = 1)\n\n## merge columns 'keyword' and 'text\ntest_df.keyword = test_df.keyword.fillna('')\ntest_df.text = test_df.keyword + \" \" + test_df.text\n\n## drop 'keyword' column\ntest_df = test_df.drop(['keyword'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.868096Z","iopub.execute_input":"2023-02-26T20:23:58.868871Z","iopub.status.idle":"2023-02-26T20:23:58.887666Z","shell.execute_reply.started":"2023-02-26T20:23:58.868817Z","shell.execute_reply":"2023-02-26T20:23:58.886149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.sample(10)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.896880Z","iopub.execute_input":"2023-02-26T20:23:58.897269Z","iopub.status.idle":"2023-02-26T20:23:58.911823Z","shell.execute_reply.started":"2023-02-26T20:23:58.897232Z","shell.execute_reply":"2023-02-26T20:23:58.910067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nsym_spell = SymSpell(max_dictionary_edit_distance=2, prefix_length=7)\ndictionary_path = pkg_resources.resource_filename(\n    \"symspellpy\", \"frequency_dictionary_en_82_765.txt\"\n)\nsym_spell.load_dictionary(dictionary_path, term_index=0, count_index=1)\n\n\nbigram_path = pkg_resources.resource_filename(\n    \"symspellpy\", \"frequency_bigramdictionary_en_243_342.txt\"\n)\nsym_spell.load_bigram_dictionary(bigram_path, term_index=0, count_index=2)\n\n\ndef preProcessing(text):\n    # 1) converting emojis\n    #for emot in UNICODE_EMOJI:\n     #   text = text.replace(emot, \"_\".join(UNICODE_EMOJI[emot].replace(\",\",\"\").replace(\":\",\"\").split()))\n    text = emoji.demojize(text)\n\n    \n    # 2) removing html code\n    soup = BeautifulSoup(text, 'lxml')\n    text = soup.get_text()\n\n    \n    # 3) convert all capital letters to lower (to make sure both instances such as 'Quake' and 'quake' are considered the same word)\n    text = text.lower()\n\n    \n    # 4) removing urls\n    pattern = re.compile(r'https?://(www\\.)?(\\w+)(\\.\\w+)(/\\w*)?')\n    text = re.sub(pattern, \"\", text)\n\n    \n    # 5) removing mentions to other twitter accounts ('@' follow by username / handle)\n    pattern = re.compile(r\"@\\w+\")\n    text = re.sub(pattern, \"\", text)\n\n    \n    # 6) removing unicode chars\n    text = text.encode(\"ascii\", \"ignore\").decode()\n\n    \n    # 7) removing punctuations\n    string.punctuation\n    text = re.sub('[%s]' % re.escape(string.punctuation), \" \",text)\n\n    \n    # 8) remove extra spaces\n    text = re.sub(' +', ' ', text).strip()\n\n    \n    # 9) correcting spelling with 'symspell'\n    words = [\n        sym_spell.lookup(\n            word, \n            Verbosity.CLOSEST, \n            max_edit_distance=2,\n            include_unknown=True\n            )[0].term \n        for word in text.split()] \n    text = \" \".join(words)\n\n    \n    # 10) Correcting componded words with bigram\n\n    words = [\n        sym_spell.lookup_compound(\n            word, \n            max_edit_distance=2\n            )[0].term \n        for word in text.split()] \n    text = \" \".join(words)\n\n    \n    # 11) removing stopwords\n    stop_words = set(stopwords.words('english'))\n    text = \" \".join([word for word in str(text).split() if word not in stop_words])\n    \n    return text","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:23:58.932901Z","iopub.execute_input":"2023-02-26T20:23:58.933337Z","iopub.status.idle":"2023-02-26T20:24:03.577647Z","shell.execute_reply.started":"2023-02-26T20:23:58.933297Z","shell.execute_reply":"2023-02-26T20:24:03.576528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.text = train_df.text.apply(preProcessing)\ntest_df.text = test_df.text.apply(preProcessing)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:03.579586Z","iopub.execute_input":"2023-02-26T20:24:03.579986Z","iopub.status.idle":"2023-02-26T20:24:23.665588Z","shell.execute_reply.started":"2023-02-26T20:24:03.579947Z","shell.execute_reply":"2023-02-26T20:24:23.664581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add column text_length to see if there is correlation\ntrain_df['text_length'] = train_df.text.apply(len)\ntest_df['text_length'] = test_df.text.apply(len)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.667515Z","iopub.execute_input":"2023-02-26T20:24:23.668116Z","iopub.status.idle":"2023-02-26T20:24:23.680561Z","shell.execute_reply.started":"2023-02-26T20:24:23.668061Z","shell.execute_reply":"2023-02-26T20:24:23.679375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.sample(10)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.683237Z","iopub.execute_input":"2023-02-26T20:24:23.683568Z","iopub.status.idle":"2023-02-26T20:24:23.701185Z","shell.execute_reply.started":"2023-02-26T20:24:23.683535Z","shell.execute_reply":"2023-02-26T20:24:23.699735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EXPERIMENTATION","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_extraction.text import TfidfVectorizer","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.702670Z","iopub.execute_input":"2023-02-26T20:24:23.703079Z","iopub.status.idle":"2023-02-26T20:24:23.710738Z","shell.execute_reply.started":"2023-02-26T20:24:23.703042Z","shell.execute_reply":"2023-02-26T20:24:23.709113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tfidf_vect = TfidfVectorizer()","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.712343Z","iopub.execute_input":"2023-02-26T20:24:23.712697Z","iopub.status.idle":"2023-02-26T20:24:23.724884Z","shell.execute_reply.started":"2023-02-26T20:24:23.712662Z","shell.execute_reply":"2023-02-26T20:24:23.723559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.text","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.726963Z","iopub.execute_input":"2023-02-26T20:24:23.727315Z","iopub.status.idle":"2023-02-26T20:24:23.740616Z","shell.execute_reply.started":"2023-02-26T20:24:23.727274Z","shell.execute_reply":"2023-02-26T20:24:23.739316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_tfidf = tfidf_vect.fit_transform(train_df.text)\ndf_test_vector = tfidf_vect.transform(test_df.text)\nprint(X_tfidf.shape)\n#print(tfidf_vect.get_feature_names())","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.742122Z","iopub.execute_input":"2023-02-26T20:24:23.742546Z","iopub.status.idle":"2023-02-26T20:24:23.916460Z","shell.execute_reply.started":"2023-02-26T20:24:23.742492Z","shell.execute_reply":"2023-02-26T20:24:23.915015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = test_df.set_index('id')\ntrain_df = train_df.set_index('id')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.918483Z","iopub.execute_input":"2023-02-26T20:24:23.919338Z","iopub.status.idle":"2023-02-26T20:24:23.926599Z","shell.execute_reply.started":"2023-02-26T20:24:23.919296Z","shell.execute_reply":"2023-02-26T20:24:23.925449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create splits train-test\nX = X_tfidf\ny = train_df.target\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.930760Z","iopub.execute_input":"2023-02-26T20:24:23.931190Z","iopub.status.idle":"2023-02-26T20:24:23.940393Z","shell.execute_reply.started":"2023-02-26T20:24:23.931153Z","shell.execute_reply":"2023-02-26T20:24:23.939322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# with RANDOM FOREST CLASSIFIER\nfrom sklearn.ensemble import RandomForestClassifier\nrf = RandomForestClassifier()\n\nrf.fit(X_train, y_train)\n  ","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:23.941661Z","iopub.execute_input":"2023-02-26T20:24:23.941997Z","iopub.status.idle":"2023-02-26T20:24:40.865350Z","shell.execute_reply.started":"2023-02-26T20:24:23.941966Z","shell.execute_reply":"2023-02-26T20:24:40.864049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf.score(X_test, y_test)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:40.866839Z","iopub.execute_input":"2023-02-26T20:24:40.867191Z","iopub.status.idle":"2023-02-26T20:24:41.081274Z","shell.execute_reply.started":"2023-02-26T20:24:40.867157Z","shell.execute_reply":"2023-02-26T20:24:41.080161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Hyperparameter tuning\nfrom sklearn.model_selection import RandomizedSearchCV\ngrid_param = {\n    'n_estimators': [100, 150, 300],\n    'max_depth': [None, 90, 120],\n    'max_features' : ['auto', 'log2', 'sqrt'],\n    'bootstrap' : [True, False]\n}\n\ngrid_search = RandomizedSearchCV(estimator = rf, param_distributions = grid_param, cv = 5, n_jobs = 1, verbose = 3)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:41.082538Z","iopub.execute_input":"2023-02-26T20:24:41.083500Z","iopub.status.idle":"2023-02-26T20:24:41.089105Z","shell.execute_reply.started":"2023-02-26T20:24:41.083460Z","shell.execute_reply":"2023-02-26T20:24:41.088131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#grid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:41.090140Z","iopub.execute_input":"2023-02-26T20:24:41.091108Z","iopub.status.idle":"2023-02-26T20:24:41.102522Z","shell.execute_reply.started":"2023-02-26T20:24:41.091072Z","shell.execute_reply":"2023-02-26T20:24:41.101168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:41.103897Z","iopub.execute_input":"2023-02-26T20:24:41.104581Z","iopub.status.idle":"2023-02-26T20:24:41.113922Z","shell.execute_reply.started":"2023-02-26T20:24:41.104540Z","shell.execute_reply":"2023-02-26T20:24:41.113005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_optimized = RandomForestClassifier(max_depth = None, n_estimators = 150, bootstrap = False, max_features = 'log2')\n\nrf_optimized.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:24:41.115080Z","iopub.execute_input":"2023-02-26T20:24:41.116068Z","iopub.status.idle":"2023-02-26T20:25:24.421507Z","shell.execute_reply.started":"2023-02-26T20:24:41.116029Z","shell.execute_reply":"2023-02-26T20:25:24.420189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_optimized.score(X_test, y_test)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:25:24.423163Z","iopub.execute_input":"2023-02-26T20:25:24.425113Z","iopub.status.idle":"2023-02-26T20:25:24.939451Z","shell.execute_reply.started":"2023-02-26T20:25:24.425057Z","shell.execute_reply":"2023-02-26T20:25:24.937896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn import metrics\n\ny_pred_class = rf_optimized.predict(X_test)\n\n\n# calculate accuracy of class predictions\nprint(\"=======Accuracy Score===========\")\nprint(\"Test acc.:\", metrics.accuracy_score(y_test, y_pred_class))\n\n# print the confusion matrix\nprint(\"=======Confision Matrix===========\")\nmetrics.confusion_matrix(y_test, y_pred_class)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:25:24.941056Z","iopub.execute_input":"2023-02-26T20:25:24.941419Z","iopub.status.idle":"2023-02-26T20:25:25.453675Z","shell.execute_reply.started":"2023-02-26T20:25:24.941383Z","shell.execute_reply":"2023-02-26T20:25:25.452367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nfrom sklearn.metrics import roc_auc_score, f1_score, confusion_matrix\n\nprint(classification_report(y_test, y_pred_class))\nprint(confusion_matrix(y_test, y_pred_class))","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:25:25.455631Z","iopub.execute_input":"2023-02-26T20:25:25.456067Z","iopub.status.idle":"2023-02-26T20:25:25.473974Z","shell.execute_reply.started":"2023-02-26T20:25:25.456031Z","shell.execute_reply":"2023-02-26T20:25:25.472728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PREDICT WITH TEST SET","metadata":{}},{"cell_type":"code","source":"test_results= rf_optimized.predict(df_test_vector)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:25:25.476070Z","iopub.execute_input":"2023-02-26T20:25:25.477091Z","iopub.status.idle":"2023-02-26T20:25:26.450634Z","shell.execute_reply.started":"2023-02-26T20:25:25.477039Z","shell.execute_reply":"2023-02-26T20:25:26.449305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.DataFrame({'id' : test_df.index.values, 'target' : test_results})\n\nsubmission_df.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:25:26.452734Z","iopub.execute_input":"2023-02-26T20:25:26.453480Z","iopub.status.idle":"2023-02-26T20:25:26.464240Z","shell.execute_reply.started":"2023-02-26T20:25:26.453438Z","shell.execute_reply":"2023-02-26T20:25:26.463024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.sample(10)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:25:26.466815Z","iopub.execute_input":"2023-02-26T20:25:26.467762Z","iopub.status.idle":"2023-02-26T20:25:26.479643Z","shell.execute_reply.started":"2023-02-26T20:25:26.467682Z","shell.execute_reply":"2023-02-26T20:25:26.478143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}