{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border-radius:20px;\n            border : black solid;\n            background-color: ##FFFFFF;\n            font-size:200%;\n            text-align: left\">\n\n<h1 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:green'><center> Natural Language Processing with Disaster Tweets </center></h1>","metadata":{}},{"cell_type":"markdown","source":"<center><img src=\"https://github.com/Isharaneranjana/kaggle_gif/blob/main/NLP%20WITH%20DISASTER%20TWEETS.gif?raw=true\"></center>","metadata":{}},{"cell_type":"markdown","source":"<div style='font-size:200%;'>\n    <a id='nan'></a>\n    <h1 style='color: red; font-weight: bold; font-family: Cascadia code;'> Contents </h1>\n</div>\n\n- [Importing necessary libraries](#import)\n- [Importing the data](#data)\n- [Exploratory Data Analysis](#eda)\n    - [Disaster vs Non-disaster distribution](#disa)\n    - [NaN values heat-map](#heatmap)\n    - [Word distribution](#dist)\n    - [Location count](#con)\n    - [Word cloud](#cloud)\n    - [Map](#map)\n- [Data Pre-processing](#preprocess)\n    - [Removing unnecessary characters and special symbols](#re)\n    - [Fitting tokenizer on texts](#fit)\n- [Model Creation and Evalutatio](#classify)\n    - [Diagram of our sequential model](#dia)\n- [Submission](#submit)","metadata":{}},{"cell_type":"markdown","source":"**Competition Description¶**\n\nTwitter has become an important communication channel in times of emergency. The ubiquitousness of smartphones enables people to announce an emergency they’re observing in real-time. Because of this, more agencies are interested in programatically monitoring Twitter (i.e. disaster relief organizations and news agencies).\n\nBut, it’s not always clear whether a person’s words are actually announcing a disaster. Take this example:\n\nThe author explicitly uses the word “ABLAZE” but means it metaphorically. This is clear to a human right away, especially with the visual aid. But it’s less clear to a machine.\n\nIn this competition, you’re challenged to build a machine learning model that predicts which Tweets are about real disasters and which one’s aren’t. You’ll have access to a dataset of 10,000 tweets that were hand classified. If this is your first time working on an NLP problem, we've created a quick tutorial to get you up and running.\n\nDisclaimer: The dataset for this competition contains text that may be considered profane, vulgar, or offensive.\n\nAcknowledgments This dataset was created by the company figure-eight and originally shared on their ‘Data For Everyone’ website here.\n\nTweet source: https://twitter.com/AnyOtherAnnaK/status/629195955506708480","metadata":{}},{"cell_type":"markdown","source":"<div style='font-size:200%;'>\n    <a id='import'></a>\n    <h1 style='color: green; font-weight: bold; font-family: Cascadia code;'>\n        <center> Importing libraries 📚 </center>\n    </h1>\n  \n</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport re\nfrom collections import defaultdict\n","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:51:57.373543Z","iopub.execute_input":"2022-08-10T07:51:57.374322Z","iopub.status.idle":"2022-08-10T07:52:07.176946Z","shell.execute_reply.started":"2022-08-10T07:51:57.374211Z","shell.execute_reply":"2022-08-10T07:52:07.175886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:07.178885Z","iopub.execute_input":"2022-08-10T07:52:07.179468Z","iopub.status.idle":"2022-08-10T07:52:07.186267Z","shell.execute_reply.started":"2022-08-10T07:52:07.179434Z","shell.execute_reply":"2022-08-10T07:52:07.185439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Train = pd.read_csv('/kaggle/input/nlp-getting-started/train.csv')\nTest = pd.read_csv('/kaggle/input/nlp-getting-started/test.csv')\nTest\n","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:07.187555Z","iopub.execute_input":"2022-08-10T07:52:07.188177Z","iopub.status.idle":"2022-08-10T07:52:07.304413Z","shell.execute_reply.started":"2022-08-10T07:52:07.188143Z","shell.execute_reply":"2022-08-10T07:52:07.303552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> Exploratory Data Analysis 📊 </center></h2><a id=\"eda\"></a>","metadata":{}},{"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/1000/1*Owa2rsDG6Rwv1IM_RdsL3A.gif)","metadata":{}},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='disa'><b>Disaster vs Non-disaster distribution<b></a></h2>","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(ncols=2, figsize=(12, 4))\nplt.tight_layout()\n\nTrain.groupby('target').count()['id'].plot(kind='pie', ax=axes[0], labels=['Not Disaster (57%)', 'Disaster (43%)'],colors=['lightcoral','lightskyblue'])\nsns.countplot(x=Train['target'], hue=Train['target'], ax=axes[1], palette=\"RdBu\")\n\naxes[0].set_ylabel('')\naxes[1].set_ylabel('')\naxes[1].set_xticklabels(['Not Disaster (4342)', 'Disaster (3271)'])\naxes[0].tick_params(axis='x', labelsize=12)\naxes[0].tick_params(axis='y', labelsize=12)\naxes[1].tick_params(axis='x', labelsize=12)\naxes[1].tick_params(axis='y', labelsize=12)\n\naxes[0].set_title('Target Distribution in Training Set', fontsize=13)\naxes[1].set_title('Target Count in Training Set', fontsize=13)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:07.306590Z","iopub.execute_input":"2022-08-10T07:52:07.307169Z","iopub.status.idle":"2022-08-10T07:52:07.715798Z","shell.execute_reply.started":"2022-08-10T07:52:07.307134Z","shell.execute_reply":"2022-08-10T07:52:07.714360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='heatmap'><b>Null values heat-map<b></a></h2>","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, sharex=True, figsize=(20,10))\nsns.heatmap(ax=axes[0], yticklabels=False, data=Train.isnull(), cbar=False, cmap=\"viridis\")\nsns.heatmap(ax=axes[1], yticklabels=False, data=Test.isnull(), cbar=False, cmap=\"tab20c\")\naxes[0].set_title('Heatmap of missing values in training data')\naxes[1].set_title('Heatmap of missing values in testing data')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:07.717991Z","iopub.execute_input":"2022-08-10T07:52:07.718918Z","iopub.status.idle":"2022-08-10T07:52:08.056423Z","shell.execute_reply.started":"2022-08-10T07:52:07.718866Z","shell.execute_reply":"2022-08-10T07:52:08.055266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='con'><b>Location Count<b></a></h2>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (9, 6))\nax = plt.axes()\nax.set_facecolor('white')\nax = ((Train.location.value_counts())[:10]).plot(kind = 'bar', color = 'lightcoral', linewidth = 2, edgecolor = 'white')\nplt.title('Location Count', fontsize = 14)\nplt.xlabel('Location', fontsize = 12)\nplt.ylabel('Count', fontsize = 12)\nax.xaxis.set_tick_params(labelsize = 12, rotation = 30)\nax.yaxis.set_tick_params(labelsize = 12)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:08.058038Z","iopub.execute_input":"2022-08-10T07:52:08.058497Z","iopub.status.idle":"2022-08-10T07:52:08.242273Z","shell.execute_reply.started":"2022-08-10T07:52:08.058451Z","shell.execute_reply":"2022-08-10T07:52:08.241088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='dist'><b>Word distribution<b></a></h2>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18, 8))\nplt.subplot(1, 2, 1)\nax1 = Train.query(\"target==1\").text.map(lambda x: len(x.split())).plot(kind=\"hist\",\n                                                                    color=\"cyan\",\n                                                                    title=\"Disaster tweets\",\n                                                                    edgecolor='white');\nplt.subplot(1, 2, 2)\nax2 = Train.query(\"target==0\").text.map(lambda x: len(x.split())).plot(kind=\"hist\",\n                                                                    color=\"orange\",\n                                                                    title=\"Non-Disaster tweets\",\n                                                                    edgecolor='white');\n\nax1.grid(visible = True, color ='grey',linestyle ='-.', linewidth = 0.5,alpha = 0.6)\nax2.grid(visible = True, color ='grey',linestyle ='-.', linewidth = 0.5,alpha = 0.6)\nplt.suptitle('Word distribution in tweets')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:08.243868Z","iopub.execute_input":"2022-08-10T07:52:08.244241Z","iopub.status.idle":"2022-08-10T07:52:08.633168Z","shell.execute_reply.started":"2022-08-10T07:52:08.244206Z","shell.execute_reply":"2022-08-10T07:52:08.632015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='cloud'><b>Word cloud ☁ <b></a></h2>","metadata":{}},{"cell_type":"code","source":"from nltk.corpus import stopwords\nfrom wordcloud import WordCloud\nstop_words= set(stopwords.words(\"english\"))\n\nstop_words.update(['https', 'http', 'amp', 'CO', 't', 'u', 'new', \"I'm\", \"would\"])\n\nwc = WordCloud(width=800,\n               height=400,\n               max_words=200,\n               stopwords=stop_words,\n               background_color='white',\n               max_font_size=150)\ndisaster_tweets_text = Train.query(\"target==1\").text\nconcat_disaster_tweets_text = disaster_tweets_text.str.cat(sep=\" \")\n\nnon_disaster_tweets_text = Train.query(\"target==0\").text\nconcat_non_disaster_tweets_text = non_disaster_tweets_text.str.cat(sep=\" \")\n\n\n\nprint('\\n\\nWord Cloud for Non-Disaster Tweets\\n\\n')\nwc.generate(concat_non_disaster_tweets_text)\nplt.figure(figsize=(16, 8))\nplt.imshow(wc, interpolation='bilinear')\nplt.axis('off')\nplt.show()\n\nprint('\\n\\nWord Cloud for Disaster Tweets\\n\\n')\nwc.generate(concat_disaster_tweets_text)\nplt.figure(figsize=(16, 8))\nplt.imshow(wc, interpolation='bilinear')\nplt.axis('off')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:08.634673Z","iopub.execute_input":"2022-08-10T07:52:08.635034Z","iopub.status.idle":"2022-08-10T07:52:11.773712Z","shell.execute_reply.started":"2022-08-10T07:52:08.635000Z","shell.execute_reply":"2022-08-10T07:52:11.772382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='map'><b>MAP<b></a></h2>","metadata":{}},{"cell_type":"code","source":"from geopy.geocoders import Nominatim\nfrom geopy.extra.rate_limiter import RateLimiter\nimport folium \nfrom folium import plugins \n\nnew_df = pd.DataFrame()\nnew_df['location'] = ((Train['location'].value_counts())[:10]).index\nnew_df['count'] = ((Train['location'].value_counts())[:10]).values\ngeolocator = Nominatim(user_agent = 'Rahil')\ngeocode = RateLimiter(geolocator.geocode, min_delay_seconds = 0.5)\nlat = {}\nlong = {}\nfor i in new_df['location']:\n    location = geocode(i)\n    lat[i] = location.latitude\n    long[i] = location.longitude\nnew_df['latitude'] = new_df['location'].map(lat)\nnew_df['longitude'] = new_df['location'].map(long)\nmap = folium.Map(location = [10.0, 10.0], tiles = 'CartoDB dark_matter', zoom_start = 1.5)\nmarkers = []\ntitle = '''<h1 align = \"center\" style = \"font-size: 15px\"><b>Top 10 Tweet Locations</b></h1>'''\nfor i, r in new_df.iterrows():\n    loss = r['count']\n    if r['count'] > 0:\n        counts = r['count'] * 0.4\n        folium.CircleMarker([float(r['latitude']), float(r['longitude'])], radius = float(counts), color = 'lightcoral', fill = True).add_to(map)\nmap.get_root().html.add_child(folium.Element(title))\nmap","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:11.775485Z","iopub.execute_input":"2022-08-10T07:52:11.775998Z","iopub.status.idle":"2022-08-10T07:52:17.749398Z","shell.execute_reply.started":"2022-08-10T07:52:11.775958Z","shell.execute_reply":"2022-08-10T07:52:17.748160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h1 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center>  Data Pre-processing ⌛</center></h1> <a id='preprocess'></a>","metadata":{}},{"cell_type":"markdown","source":"![](https://cdn.dribbble.com/users/2017910/screenshots/5102683/ai_trends_dribbble_shot.gif)","metadata":{}},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='re'><b>Removing unnecessary characters and special symbols<b></a></h2>","metadata":{}},{"cell_type":"code","source":"def cleanText(text):\n    whitespace = re.compile(r\"\\s+\")\n    web_address = re.compile(r\"(?i)http(s):\\/\\/[a-z0-9.~_\\-\\/]+\")\n    user = re.compile(r\"(?i)@[a-z0-9_]+\")\n    text = whitespace.sub(' ', text)\n    text = web_address.sub('', text)\n    text = user.sub('', text)\n    text = re.sub(r\"\\[[^()]*\\]\", \"\", text)\n    text = re.sub(\"\\d+\", \"\", text)\n    text = re.sub(r'[^\\w\\s]','',text)\n    text = re.sub(r\"(?:@\\S*|#\\S*|http(?=.*://)\\S*)\", \"\", text)\n    return text.lower()\n\nTrain.text = [cleanText(item) for item in Train.text]\nTrain","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:17.753023Z","iopub.execute_input":"2022-08-10T07:52:17.753410Z","iopub.status.idle":"2022-08-10T07:52:17.943835Z","shell.execute_reply.started":"2022-08-10T07:52:17.753372Z","shell.execute_reply":"2022-08-10T07:52:17.942419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install pyspellchecker","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:17.945413Z","iopub.execute_input":"2022-08-10T07:52:17.946134Z","iopub.status.idle":"2022-08-10T07:52:33.008547Z","shell.execute_reply.started":"2022-08-10T07:52:17.946095Z","shell.execute_reply":"2022-08-10T07:52:33.007245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from spellchecker import SpellChecker\n\nspell = SpellChecker()\ndef correct_spellings(text):\n    corrected_text = []\n    misspelled_words = spell.unknown(text.split())\n    for word in text.split():\n        if word in misspelled_words:\n            corrected_text.append(spell.correction(word))\n        else:\n            corrected_text.append(word)\n    return \" \".join(corrected_text)\n        \ntext = \"corect me plese\"\ncorrect_spellings(text)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:33.010474Z","iopub.execute_input":"2022-08-10T07:52:33.010881Z","iopub.status.idle":"2022-08-10T07:52:33.192754Z","shell.execute_reply.started":"2022-08-10T07:52:33.010841Z","shell.execute_reply":"2022-08-10T07:52:33.191372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='fit'><b>Fitting tokenizer on texts<b></a></h2>","metadata":{}},{"cell_type":"code","source":"tokenizer = tf.keras.preprocessing.text.Tokenizer()\ntokenizer.oov_token = '<oovToken>'\ntokenizer.fit_on_texts(Train.text)\nvocab = tokenizer.word_index\nvocabCount = len(vocab)+1\n\nvocabCount","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:33.194119Z","iopub.execute_input":"2022-08-10T07:52:33.194578Z","iopub.status.idle":"2022-08-10T07:52:34.574141Z","shell.execute_reply.started":"2022-08-10T07:52:33.194531Z","shell.execute_reply":"2022-08-10T07:52:34.572843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_Train = tf.keras.preprocessing.sequence.pad_sequences(tokenizer.texts_to_sequences(Train.text.to_numpy()), padding='pre')\ny_Train = Train.target.to_numpy()\n\nx_Train.shape, y_Train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:52:34.576108Z","iopub.execute_input":"2022-08-10T07:52:34.576505Z","iopub.status.idle":"2022-08-10T07:52:35.012014Z","shell.execute_reply.started":"2022-08-10T07:52:34.576461Z","shell.execute_reply":"2022-08-10T07:52:35.010718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> Model Creation and Evalutation </center></h2><a id=\"model\"></a>","metadata":{}},{"cell_type":"markdown","source":"![](https://miro.medium.com/max/1280/1*czcdGNhz6jvyxSRvmuxlSQ.gif)","metadata":{}},{"cell_type":"markdown","source":"## **Sequential model**\n\n#### A Sequential model is appropriate for a plain stack of layers where each layer has exactly one input tensor and one output tensor.\n\nThe Embedding layer takes the integer-encoded vocabulary and looks up the embedding vector for each word-index. \n\nGlobal Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding category of the classification task in the last mlpconv layer. Instead of adding fully connected layers on top of the feature maps, we take the average of each feature map.\n\nA dense layer is a layer that is deeply connected with its preceding layer which means the neurons of the layer are connected to every neuron of its preceding layer. This layer is the most commonly used layer in artificial neural network networks.","metadata":{}},{"cell_type":"code","source":"model = tf.keras.Sequential()\nmodel.add(tf.keras.layers.Embedding(input_dim=vocabCount+1, output_dim=64, input_length=31))\nmodel.add(tf.keras.layers.GlobalAveragePooling1D())\nmodel.add(tf.keras.layers.Dense(256, activation='relu'))\nmodel.add(tf.keras.layers.Dense(128, activation='relu'))\nmodel.add(tf.keras.layers.Dense(64, activation='relu'))\nmodel.add(tf.keras.layers.Dense(32, activation='relu'))\nmodel.add(tf.keras.layers.Dense(1, activation='sigmoid'))\n\n\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:57:22.800193Z","iopub.execute_input":"2022-08-10T07:57:22.800644Z","iopub.status.idle":"2022-08-10T07:57:22.882195Z","shell.execute_reply.started":"2022-08-10T07:57:22.800597Z","shell.execute_reply":"2022-08-10T07:57:22.881019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 align=\"center\" ><a id='dia'><b>Diagram of our sequential model<b></a></h2>","metadata":{}},{"cell_type":"code","source":"tf.keras.utils.plot_model(model)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:57:29.811075Z","iopub.execute_input":"2022-08-10T07:57:29.811507Z","iopub.status.idle":"2022-08-10T07:57:29.924427Z","shell.execute_reply.started":"2022-08-10T07:57:29.811471Z","shell.execute_reply":"2022-08-10T07:57:29.923352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(x_Train, y_Train, epochs=10, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:57:35.557468Z","iopub.execute_input":"2022-08-10T07:57:35.557958Z","iopub.status.idle":"2022-08-10T07:58:02.665598Z","shell.execute_reply.started":"2022-08-10T07:57:35.557911Z","shell.execute_reply":"2022-08-10T07:58:02.664404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:58:14.591441Z","iopub.execute_input":"2022-08-10T07:58:14.591898Z","iopub.status.idle":"2022-08-10T07:58:14.607032Z","shell.execute_reply.started":"2022-08-10T07:58:14.591860Z","shell.execute_reply":"2022-08-10T07:58:14.605789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center>Submission file </center></h2><a id=\"sv\"></a>\n","metadata":{}},{"cell_type":"markdown","source":"![](https://dashtechinc.com/wp-content/uploads/2020/01/Machine-Learning-Hero-Banner.png)","metadata":{}},{"cell_type":"code","source":"idCol = Test['id'].to_numpy()\nTest = Test.drop(columns=['keyword', 'location', 'id'])\nTest.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:58:24.891186Z","iopub.execute_input":"2022-08-10T07:58:24.891677Z","iopub.status.idle":"2022-08-10T07:58:25.003281Z","shell.execute_reply.started":"2022-08-10T07:58:24.891609Z","shell.execute_reply":"2022-08-10T07:58:25.001600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Test.text = [cleanText(item) for item in Test.text]\nTest","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:58:27.806647Z","iopub.execute_input":"2022-08-10T07:58:27.807076Z","iopub.status.idle":"2022-08-10T07:58:27.881256Z","shell.execute_reply.started":"2022-08-10T07:58:27.807038Z","shell.execute_reply":"2022-08-10T07:58:27.880112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_Test = tf.keras.preprocessing.sequence.pad_sequences(tokenizer.texts_to_sequences(Test.text.to_numpy()), padding='pre', maxlen=31)\n\nx_Test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:53:02.180839Z","iopub.execute_input":"2022-08-10T07:53:02.181572Z","iopub.status.idle":"2022-08-10T07:53:02.259815Z","shell.execute_reply.started":"2022-08-10T07:53:02.181525Z","shell.execute_reply":"2022-08-10T07:53:02.258709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = model.predict(x_Test)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:58:31.454034Z","iopub.execute_input":"2022-08-10T07:58:31.454468Z","iopub.status.idle":"2022-08-10T07:58:31.747761Z","shell.execute_reply.started":"2022-08-10T07:58:31.454430Z","shell.execute_reply":"2022-08-10T07:58:31.746540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = preds.reshape(len(preds))\nfor i in range(len(preds)):\n    if preds[i]>0.9:\n        preds[i] = 1\n    else:\n        preds[i] = 0\n\nsubmission = pd.DataFrame({'id': idCol, 'target': preds})\nsubmission.target = submission.target.astype(int)\nsubmission.set_index('id')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:58:33.021303Z","iopub.execute_input":"2022-08-10T07:58:33.022045Z","iopub.status.idle":"2022-08-10T07:58:33.045216Z","shell.execute_reply.started":"2022-08-10T07:58:33.022000Z","shell.execute_reply":"2022-08-10T07:58:33.044307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T07:58:37.002510Z","iopub.execute_input":"2022-08-10T07:58:37.002931Z","iopub.status.idle":"2022-08-10T07:58:37.015162Z","shell.execute_reply.started":"2022-08-10T07:58:37.002896Z","shell.execute_reply":"2022-08-10T07:58:37.013860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"output.to_csv('submission.csv',index=False)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color: #DA70D6;\n            font-size:200%;\n            text-align: left\">\n\n<h1 style='; border:0; border-radius: 10px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> YOUR FEEDBACKS IS SO VALUABLE FOR ME </center></h1>","metadata":{}}]}