{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":10737,"databundleVersionId":290346,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:27:53.694944Z","iopub.execute_input":"2024-12-12T08:27:53.695499Z","iopub.status.idle":"2024-12-12T08:27:55.076998Z","shell.execute_reply.started":"2024-12-12T08:27:53.695444Z","shell.execute_reply":"2024-12-12T08:27:55.075278Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"إيه فكرة المشروع ده؟ # \nالمشروع ده غالبًا بيستخدم LSTM (ودي نوع من الشبكات العصبية) علشان تصنف نصوص زي الأسئلة على موقع Quora، وده يعني إنك تعرف لو السؤالين مكررّين أو لو السؤال بيخص موضوع معين.","metadata":{}},{"cell_type":"markdown","source":"# steps\n**1- UnZip for Folder**\n\n**2- Open folder (folder name(glove))**\n\n**3- Read text File for embedding (folder name)**\n\n4- Preprocessing:\n   * Tokenizer===convert long text to word \n   * Embadding===convert words to vectors of numbers using (pre-trained model)\n\n5- LSTM(LSTM and dropout, ANN)\n\n6- Compile(optimizer = adam , loss = 'binary_crossentropy' , metrics ['accurecy'])\n\n7- Train the model (fit  epoches , batch size) on  (X_train , Y_train) + validation (x_val , y_val)\n\n8- Evaluation (accuracy , precion , recall , f1score) on validation (x_val, y_val )\n\n9- Prediction on test_data to creat sumbission file (qid , predistion )\n","metadata":{}},{"cell_type":"markdown","source":"### Steps:\n \n1. **Extract the Embedding File:**\n   - Unzip the file (embedding.zip).\n   - Open the folder (glove.840B.300d).\n   - Read the text file (glove.840B.300d) to load the GloVe embedding matrix.\n \n---\n \n### Preprocessing:\n \n- **Tokenizer:**\n   - Convert the long text into individual words (tokens).\n- **Embedding:**\n   - Convert the words (tokens) into numerical vectors (word vectors) using the pre-trained GloVe model.\n \n---\n \n### Building the Model:\n \n1. **Choose the Model Type:**\n   - Use LSTM with Dropout and ANN layers for better performance.\n2. **Compile the Model:**\n   - Choose the optimizer: `optimizer='adam'`.\n   - Choose the loss function: `loss='binary_crossentropy'`.\n   - Choose the metrics: `metrics=['accuracy']`.\n \n3. **Train the Model:**\n   - Use the `fit` function to train the model on the training data (`X_train`, `y_train`), specifying the number of epochs and batch size.\n   - Set the validation data during training: (`X_val`, `y_val`).\n \n4. **Evaluate the Model:**\n   - Measure performance using metrics like accuracy, precision, recall, and F1-score on the validation data (`X_val`, `y_val`).\n \n5. **Make Predictions:**\n   - Make predictions on the test data (`X_test`) and create the submission file containing the columns: `qid`, `prediction`.\n \n---","metadata":{}},{"cell_type":"code","source":"import zipfile\nimport os\nimport numpy as np\nimport pandas as pd","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:27:55.078999Z","iopub.execute_input":"2024-12-12T08:27:55.079596Z","iopub.status.idle":"2024-12-12T08:27:55.086030Z","shell.execute_reply.started":"2024-12-12T08:27:55.079542Z","shell.execute_reply":"2024-12-12T08:27:55.084463Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 1- UnZip for Folder","metadata":{}},{"cell_type":"code","source":"zip_file_path = '/kaggle/input/quora-insincere-questions-classification/embeddings.zip'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:27:55.087419Z","iopub.execute_input":"2024-12-12T08:27:55.087790Z","iopub.status.idle":"2024-12-12T08:27:55.110597Z","shell.execute_reply.started":"2024-12-12T08:27:55.087753Z","shell.execute_reply":"2024-12-12T08:27:55.109491Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"with zipfile.ZipFile(zip_file_path , 'r') as zip_rf:\n    zip_rf.printdir()\n    zip_rf.extract('glove.840B.300d/glove.840B.300d.txt' , '/kaggle/working/')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:27:55.113389Z","iopub.execute_input":"2024-12-12T08:27:55.113906Z","iopub.status.idle":"2024-12-12T08:29:03.700352Z","shell.execute_reply.started":"2024-12-12T08:27:55.113843Z","shell.execute_reply":"2024-12-12T08:29:03.698096Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"emb_file_path = '/kaggle/working/glove.840B.300d/glove.840B.300d.txt'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:29:03.702710Z","iopub.execute_input":"2024-12-12T08:29:03.703168Z","iopub.status.idle":"2024-12-12T08:29:03.710728Z","shell.execute_reply.started":"2024-12-12T08:29:03.703124Z","shell.execute_reply":"2024-12-12T08:29:03.709212Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"emb_file_path","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:29:03.712464Z","iopub.execute_input":"2024-12-12T08:29:03.713044Z","iopub.status.idle":"2024-12-12T08:29:04.766682Z","shell.execute_reply.started":"2024-12-12T08:29:03.712985Z","shell.execute_reply":"2024-12-12T08:29:04.765254Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2- Read text File for embedding (folder name)","metadata":{}},{"cell_type":"code","source":"def load_glove_embeddings(file_path):\n    embedding_index = {}\n    with open(file_path, 'r', encoding='utf-8') as f:\n        for line in f:\n            values = line.split()\n            word = values[0]\n            try:\n                coeff = np.array(values[1:], dtype='float32')\n                embedding_index[word] = coeff\n            except ValueError as e:\n                #print(f\"Skipping line with invalid data for word '{word}': {line.strip()}\")\n                continue  \n    print(f'Loaded {len(embedding_index)} word vectors!')\n    return embedding_index\n \nembedding_index = load_glove_embeddings(emb_file_path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:29:04.768156Z","iopub.execute_input":"2024-12-12T08:29:04.768525Z","iopub.status.idle":"2024-12-12T08:31:44.718519Z","shell.execute_reply.started":"2024-12-12T08:29:04.768488Z","shell.execute_reply":"2024-12-12T08:31:44.717153Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/train.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:31:44.720382Z","iopub.execute_input":"2024-12-12T08:31:44.720793Z","iopub.status.idle":"2024-12-12T08:31:50.021792Z","shell.execute_reply.started":"2024-12-12T08:31:44.720755Z","shell.execute_reply":"2024-12-12T08:31:50.020508Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:31:50.023517Z","iopub.execute_input":"2024-12-12T08:31:50.023870Z","iopub.status.idle":"2024-12-12T08:31:50.052909Z","shell.execute_reply.started":"2024-12-12T08:31:50.023837Z","shell.execute_reply":"2024-12-12T08:31:50.051425Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"texts=train_data['question_text']\nlabels=train_data['target']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:31:50.057461Z","iopub.execute_input":"2024-12-12T08:31:50.058249Z","iopub.status.idle":"2024-12-12T08:31:50.069720Z","shell.execute_reply.started":"2024-12-12T08:31:50.058203Z","shell.execute_reply":"2024-12-12T08:31:50.068216Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4- Tokenization : Tokenizer===convert long text to word","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.text import Tokenizer\nfrom tensorflow.keras.preprocessing.sequence import pad_sequences\n \nMax_num_tokens = 20000 \nMax_number_seq = 8\n \ntokenizer = Tokenizer(num_words=Max_num_tokens)\ntokenizer.fit_on_texts(texts)\n \nsequences = tokenizer.texts_to_sequences(texts)\n \npadded_sequences = pad_sequences(sequences, maxlen=Max_number_seq)\n \nprint(f'Tokens number: {len(tokenizer.word_index)} | Unique Tokens')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:31:50.071825Z","iopub.execute_input":"2024-12-12T08:31:50.072359Z","iopub.status.idle":"2024-12-12T08:33:01.275084Z","shell.execute_reply.started":"2024-12-12T08:31:50.072306Z","shell.execute_reply":"2024-12-12T08:33:01.273821Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5-  Embadding===convert words to vectors of numbers using (pre-trained model)","metadata":{}},{"cell_type":"code","source":"#embedding_index\n#tokenizer.word_index\nEmbedding_dim = 300\n\nword_index = tokenizer.word_index\nnum_words =len(word_index)+1\n\nembedding_matrix = np.zeros((num_words ,Embedding_dim ))\n\nfor word , i in word_index.items():\n    embedding_vector = embedding_index.get(word) \n    if embedding_vector is not None:\n        embedding_matrix[i] = embedding_vector\n\nprint(f'{embedding_matrix.shape}')\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:33:01.276301Z","iopub.execute_input":"2024-12-12T08:33:01.276905Z","iopub.status.idle":"2024-12-12T08:33:02.651657Z","shell.execute_reply.started":"2024-12-12T08:33:01.276868Z","shell.execute_reply":"2024-12-12T08:33:02.650270Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"embedding_matrix[i]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:33:02.653210Z","iopub.execute_input":"2024-12-12T08:33:02.653547Z","iopub.status.idle":"2024-12-12T08:33:02.666468Z","shell.execute_reply.started":"2024-12-12T08:33:02.653515Z","shell.execute_reply":"2024-12-12T08:33:02.665300Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6- Building LSTM NETWORK","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Embedding , LSTM , Dense , Dropout\n\nmodel = Sequential()\nmodel.add(Embedding(input_dim = num_words,\n                   output_dim = Embedding_dim,\n                   weights = [embedding_matrix],\n                   input_length = Max_number_seq,\n                   trainable = False\n                   \n                   \n))\n\nmodel.add(LSTM(128, return_sequences = True))\nmodel.add(Dropout(0.5))\n\nmodel.add(LSTM(64))\nmodel.add(Dropout(0.5))\n\nmodel.add(Dense(64,activation = 'relu'))\nmodel.add(Dropout(0.5))\n\nmodel.add(Dense(1, activation = 'sigmoid'))\n\nmodel.compile(optimizer = 'adam', loss = 'binary_crossentropy', metrics = ['accuracy'])\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:33:02.667887Z","iopub.execute_input":"2024-12-12T08:33:02.668267Z","iopub.status.idle":"2024-12-12T08:33:04.317569Z","shell.execute_reply.started":"2024-12-12T08:33:02.668225Z","shell.execute_reply":"2024-12-12T08:33:04.316289Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nx_train , x_val , y_train , y_val = train_test_split(padded_sequences , labels , test_size= 0.3, random_state=42)\n\nhistory = model.fit(x_train , y_train,\n                   validation_data=(x_val, y_val),\n                   epochs = 5,\n                   batch_size = 512)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:33:04.319163Z","iopub.execute_input":"2024-12-12T08:33:04.319598Z","iopub.status.idle":"2024-12-12T08:49:29.123243Z","shell.execute_reply.started":"2024-12-12T08:33:04.319563Z","shell.execute_reply":"2024-12-12T08:49:29.118341Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nplt.figure(figsize=(12, 6))\n\n#plot mae\nplt.subplot(1, 2, 1)\nplt.plot(history.history['accuracy'], label='Training Accuracy')\nplt.plot(history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Accuracy during Training')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.legend()\n\nplt.subplot(1, 2, 2)\nplt.plot(history.history['loss'], label='Training Loss')\nplt.plot(history.history['val_loss'], label='Validation Loss')\nplt.title('Loss during Training')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:54:08.197120Z","iopub.execute_input":"2024-12-12T08:54:08.197656Z","iopub.status.idle":"2024-12-12T08:54:08.691195Z","shell.execute_reply.started":"2024-12-12T08:54:08.197619Z","shell.execute_reply":"2024-12-12T08:54:08.690072Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_data = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/test.csv')\ntest_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:54:37.328225Z","iopub.execute_input":"2024-12-12T08:54:37.328652Z","iopub.status.idle":"2024-12-12T08:54:38.847570Z","shell.execute_reply.started":"2024-12-12T08:54:37.328614Z","shell.execute_reply":"2024-12-12T08:54:38.846208Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/sample_submission.csv')\nsubmission","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:54:51.213261Z","iopub.execute_input":"2024-12-12T08:54:51.213685Z","iopub.status.idle":"2024-12-12T08:54:51.685505Z","shell.execute_reply.started":"2024-12-12T08:54:51.213648Z","shell.execute_reply":"2024-12-12T08:54:51.684395Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_text = test_data['question_text']\n  \nsequences = tokenizer.texts_to_sequences(test_text)\n \ntest_padded_sequences = pad_sequences(sequences, maxlen=Max_number_seq)\n# padded_sequences : This will be the input of the Model (Neural Network)\n\nprint(f'Tokens number: {len(tokenizer.word_index)} | Unique Tokens')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:55:33.520296Z","iopub.execute_input":"2024-12-12T08:55:33.521003Z","iopub.status.idle":"2024-12-12T08:55:41.025475Z","shell.execute_reply.started":"2024-12-12T08:55:33.520961Z","shell.execute_reply":"2024-12-12T08:55:41.024291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"predictions = (model.predict(test_padded_sequences) > 0.5).astype(\"int32\")\npredictions","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:56:02.755077Z","iopub.execute_input":"2024-12-12T08:56:02.755589Z","iopub.status.idle":"2024-12-12T08:57:37.154792Z","shell.execute_reply.started":"2024-12-12T08:56:02.755550Z","shell.execute_reply":"2024-12-12T08:57:37.153558Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.DataFrame({'qid': test_data['qid'], 'prediction': predictions.flatten()})\nsubmission.to_csv('/kaggle/working/submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:58:19.629798Z","iopub.execute_input":"2024-12-12T08:58:19.630908Z","iopub.status.idle":"2024-12-12T08:58:20.050794Z","shell.execute_reply.started":"2024-12-12T08:58:19.630853Z","shell.execute_reply":"2024-12-12T08:58:20.049635Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission_file = pd.read_csv('/kaggle/working/submission.csv')\nsubmission_file.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T08:58:23.640190Z","iopub.execute_input":"2024-12-12T08:58:23.640616Z","iopub.status.idle":"2024-12-12T08:58:24.001229Z","shell.execute_reply.started":"2024-12-12T08:58:23.640580Z","shell.execute_reply":"2024-12-12T08:58:23.999872Z"}},"outputs":[],"execution_count":null}]}