{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30716,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-06-02T17:02:44.491451Z","iopub.execute_input":"2024-06-02T17:02:44.492200Z","iopub.status.idle":"2024-06-02T17:02:45.483602Z","shell.execute_reply.started":"2024-06-02T17:02:44.492166Z","shell.execute_reply":"2024-06-02T17:02:45.482548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Installing packages\n!pip install duckdb\n!pip install watermark\n!pip install pysmiles\n!pip install rdkit","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:02:45.485724Z","iopub.execute_input":"2024-06-02T17:02:45.486181Z","iopub.status.idle":"2024-06-02T17:03:39.913336Z","shell.execute_reply.started":"2024-06-02T17:02:45.486146Z","shell.execute_reply":"2024-06-02T17:03:39.912357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import of libraries\n\n# System libraries\nimport re\nimport os\nimport unicodedata\nimport itertools\n\n# Library for file manipulation\nimport pandas as pd\nimport numpy as np\nimport pandas\n\n# Database\nimport duckdb\n\n# Data visualization\nimport pysmiles\nimport plotly\nimport seaborn as sns\nimport matplotlib.pylab as pl\nimport matplotlib as m\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nfrom matplotlib import pyplot as plt\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw, AllChem\nfrom rdkit import RDLogger\nfrom rdkit import Chem\nfrom rdkit.Chem.Draw import IPythonConsole\nfrom rdkit.Chem import Draw\nfrom rdkit.Chem.Draw import rdMolDraw2D\n\n# Python version\nfrom IPython.display import SVG\nIPythonConsole.ipython_useSVG=True\n\n# Configuration for graph width and layout\nsns.set_theme(style='whitegrid')\npalette='viridis'\n\n# Warnings remove alerts\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Python version\nfrom platform import python_version\nprint('Python version in this Jupyter Notebook:', python_version())\n\n# Load library versions\nimport watermark\n\n# Library versions\n%reload_ext watermark\n%watermark -a \"Library versions\" --iversions","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:03:39.922475Z","iopub.execute_input":"2024-06-02T17:03:39.922769Z","iopub.status.idle":"2024-06-02T17:03:41.716779Z","shell.execute_reply.started":"2024-06-02T17:03:39.922737Z","shell.execute_reply":"2024-06-02T17:03:41.715830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **1.Database**","metadata":{}},{"cell_type":"code","source":"# Database - Parquet format\ndata_train = '/kaggle/input/leash-BELKA/train.parquet'\ntest_path = '/kaggle/input/leash-BELKA/test.parquet'\n\n# Bank connection\ncon = duckdb.connect()\n\n# Query\ndata = con.query(f\"\"\"(SELECT * FROM parquet_scan('{data_train}') \nWHERE binds = 0\nORDER BY random()\nLIMIT 30000)\nUNION ALL\n(SELECT * FROM parquet_scan('{data_train}')\nWHERE binds = 1\nORDER BY random()\nLIMIT 30000)\"\"\").df()\n\n# Closing database\ncon.close()\n\n# Saving dataset\ndata.to_csv(\"/kaggle/working/dataset.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:03:41.717887Z","iopub.execute_input":"2024-06-02T17:03:41.718419Z","iopub.status.idle":"2024-06-02T17:04:31.908631Z","shell.execute_reply.started":"2024-06-02T17:03:41.718389Z","shell.execute_reply":"2024-06-02T17:04:31.907592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**1. Defines the file paths:**\n\ndata_train and test_path are variables that store the paths to the train and test Parquet files, respectively. These paths are specific to the Kaggle environment.\n\n**2. Connects to the DuckDB database:**\n\ncon = duckdb.connect() establishes a connection to the DuckDB database. DuckDB is a column-oriented database management \nsystem efficient for analytical queries.\n\n**3. Executes an SQL query:**\n\n- The SQL query uses parquet_scan to read data from the Parquet files.\n\n- The query is split into two parts that are combined using UNION ALL.\n\n- The first part selects up to 30,000 random rows (ORDER BY random() LIMIT 30000) where the column binds is equal to 0.\n\n- The second part selects up to 30,000 random rows where the column binds is equal to 1.\n\n- The result of the query is stored in the variable data as a Pandas DataFrame using .df().\n\n**4. Closes the database connection:**\n\ncon.close() closes the connection to the DuckDB database.\n\n**5. Saves the DataFrame as a CSV file:**\n\ndata.to_csv(\"/kaggle/working/dataset.csv\") saves the data DataFrame to a CSV file in the Kaggle working directory.\nIn summary, the code reads data from a Parquet file, randomly selects a balanced set of rows based on the binds column, and saves this data to a CSV file for later use.","metadata":{}},{"cell_type":"code","source":"# Viewing first 5 data\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:04:31.909896Z","iopub.execute_input":"2024-06-02T17:04:31.910196Z","iopub.status.idle":"2024-06-02T17:04:31.931948Z","shell.execute_reply.started":"2024-06-02T17:04:31.910170Z","shell.execute_reply":"2024-06-02T17:04:31.931067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Viewing 5 latest data\ndata.tail()","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:04:31.933260Z","iopub.execute_input":"2024-06-02T17:04:31.933619Z","iopub.status.idle":"2024-06-02T17:04:31.945939Z","shell.execute_reply.started":"2024-06-02T17:04:31.933586Z","shell.execute_reply":"2024-06-02T17:04:31.945048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Rows and columns\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:04:31.947062Z","iopub.execute_input":"2024-06-02T17:04:31.947378Z","iopub.status.idle":"2024-06-02T17:04:31.956387Z","shell.execute_reply.started":"2024-06-02T17:04:31.947350Z","shell.execute_reply":"2024-06-02T17:04:31.955508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Info data\ndata.info()","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:04:31.957520Z","iopub.execute_input":"2024-06-02T17:04:31.957809Z","iopub.status.idle":"2024-06-02T17:04:32.010471Z","shell.execute_reply.started":"2024-06-02T17:04:31.957782Z","shell.execute_reply":"2024-06-02T17:04:32.009566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Data type\ndata.dtypes","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:04:32.012045Z","iopub.execute_input":"2024-06-02T17:04:32.012433Z","iopub.status.idle":"2024-06-02T17:04:32.019716Z","shell.execute_reply.started":"2024-06-02T17:04:32.012398Z","shell.execute_reply":"2024-06-02T17:04:32.018868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **2.Pre processing**\n\n- Here in the preprocessing stage, we perform several essential steps to prepare the data for analysis. First, we apply data cleaning techniques to remove duplicate entries and handle missing values. Next, we transform categorical variables into numerical variables using appropriate encoding methods, such as one-hot encoding or label encoding, depending on the nature of the data. Additionally, we normalize or standardize the numerical variables to ensure they are on the same scale, which is crucial for many machine learning algorithms. These preprocessing steps ensure that the data is in a suitable format for analytical models, enhancing the accuracy and efficiency of subsequent analyses.","metadata":{}},{"cell_type":"code","source":"# Converting column\ndata['molecule'] = data['molecule_smiles'].apply(Chem.MolFromSmiles)\n\n# Creating a onehotencode function\ndef modl(molecule_data, radius=2, bits=1024):\n    if molecule_data is None:\n        return None\n    return list(AllChem.GetMorganFingerprintAsBitVect(molecule_data, radius, nBits=bits))\n\n# Viewing new column\ndata['H1_ecfp'] = data['molecule'].apply(modl)\n\n# New column\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:04:32.020789Z","iopub.execute_input":"2024-06-02T17:04:32.021041Z","iopub.status.idle":"2024-06-02T17:06:10.795088Z","shell.execute_reply.started":"2024-06-02T17:04:32.021018Z","shell.execute_reply":"2024-06-02T17:06:10.794152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## One hot encoder\n\nThe OneHotEncoder is a preprocessing technique used in machine learning and data analysis to convert categorical variables into a format that can be supplied to machine learning algorithms, which generally require numerical inputs.\n\n## How OneHotEncoder Works:\n\n- Input: Categorical variable with n distinct categories.\n- Output: n binary variables (0 or 1).\n\n## Steps of OneHotEncoder:\n\n- **Identify Categories:** Identifies all possible categories in the categorical variable. Create Binary Variables: Creates a new binary variable for each category. Each binary variable corresponds to one of the categories of the original variable.\n\n\n- **Encoding:** For each observation, the binary variable corresponding to that observation's category is set to 1, while all other binary variables are set to 0.\n\n**Example:** Suppose we have a categorical variable \"Color\" with categories [\"Red\", \"Green\", \"Blue\"].\n\n**Original Variable:** \"Red\", \"Green\", \"Blue\", \"Red\"\n\n**One-Hot Encoding:**\n\n- Red: [1, 0, 0]\n\n- Green: [0, 1, 0]\n\n- Blue: [0, 0, 1]\n\n- Red: [1, 0, 0]\n\n## Advantages\n\n**No Implicit Order:**\n\nUnlike label encoding, OneHotEncoder does not impose a numerical order on the categories. Algorithm Compatibility: Many machine learning algorithms work better with numerical variables and may misinterpret categorical variables if encoded otherwise.\n\n## Disadvantages\n\n**High Dimensionality:**\n\n- If the categorical variable has many unique categories, OneHotEncoder can generate a very large number of variables, increasing the dimensionality of the data.\n\n**Efficiency:** With high dimensionality, model training can become less efficient. OneHotEncoder is widely used in machine learning pipelines and can be easily implemented using Python libraries such as scikit-learn.","metadata":{}},{"cell_type":"code","source":"# Importing library\nfrom sklearn.preprocessing import OneHotEncoder\n\n# Creating an instance\nencoder_onehot = OneHotEncoder(sparse_output=False)\n\n# Training\nencoder_onehot_fit = encoder_onehot.fit_transform(data['protein_name'].values.reshape(-1, 1))\n\n# Viewing\nencoder_onehot","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:06:10.796258Z","iopub.execute_input":"2024-06-02T17:06:10.796557Z","iopub.status.idle":"2024-06-02T17:06:10.903176Z","shell.execute_reply.started":"2024-06-02T17:06:10.796531Z","shell.execute_reply":"2024-06-02T17:06:10.902334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Viewing column\ndata.protein_name","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:06:10.907983Z","iopub.execute_input":"2024-06-02T17:06:10.908255Z","iopub.status.idle":"2024-06-02T17:06:10.915411Z","shell.execute_reply.started":"2024-06-02T17:06:10.908231Z","shell.execute_reply":"2024-06-02T17:06:10.914570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Here, we apply the One Hot Encoder to the \"protein_name\" variable. This process transforms the categorical variable into a numerical format, where each unique category is represented by a binary vector. Thus, each value of \"protein_name\" is converted into a combination of 0s and 1s, facilitating the use of this variable in machine learning algorithms that require numerical inputs.","metadata":{}},{"cell_type":"markdown","source":"# 3.Training and testing division","metadata":{}},{"cell_type":"code","source":"# Combine ECFPs and one-hot encoded protein_name\nX = [ecfp + protein for ecfp, protein in zip(data['H1_ecfp'].tolist(), encoder_onehot_fit.tolist())]\ny = data['binds'].tolist()","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:06:10.916455Z","iopub.execute_input":"2024-06-02T17:06:10.916760Z","iopub.status.idle":"2024-06-02T17:06:12.265730Z","shell.execute_reply.started":"2024-06-02T17:06:10.916736Z","shell.execute_reply":"2024-06-02T17:06:12.264947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Here, we perform the division between two variables: \"H1_ecfp\" and the target variable called \"binds.\" This step is crucial for normalizing the data, ensuring that the values of \"H1_ecfp\" are scaled relative to the target variable \"binds.\" Normalization is important to avoid scaling issues that can impair the performance of various machine learning algorithms, especially those based on distance, such as k-nearest neighbors (KNN) and clustering methods moreover, this operation can provide valuable insights into the proportional relationship between \"H1_ecfp\" and \"binds,\" allowing for better interpretation of the model's results. The division can highlight hidden trends or patterns in the data that may be critical for predictive modeling. With proper normalization, we can improve the model's stability and accuracy, ensuring that all variables contribute equally to the learning process.","metadata":{}},{"cell_type":"markdown","source":"# 4. Model training and testing","metadata":{}},{"cell_type":"code","source":"# Importing library\nfrom sklearn.model_selection import train_test_split\n\n# Model training\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:06:12.266833Z","iopub.execute_input":"2024-06-02T17:06:12.267123Z","iopub.status.idle":"2024-06-02T17:06:12.408851Z","shell.execute_reply.started":"2024-06-02T17:06:12.267099Z","shell.execute_reply":"2024-06-02T17:06:12.408058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**X_train and y_train:** These are the training sets for the features (X) and the labels (y) respectively. These are the data that will be used to train the machine learning model.\n\n**X_test and y_test:** These are the testing sets for the features (X) and the labels (y) respectively. These are the data that will be used to evaluate the model's performance after training.\n\n**train_test_split(X, y, test_size=0.2, random_state=42):** This function splits the data into training and testing sets. X represents the features (i.e., the inputs), y represents the labels (i.e., the outputs). The test_size=0.2 parameter indicates that 20% of the data will be used as the test set, while 80% will be used as the training set. The random_state=42 parameter ensures that the split is always the same, which is useful for reproducibility of results.","metadata":{}},{"cell_type":"markdown","source":"# 5.Machine learning model","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Importing libraries\nfrom tqdm import tqdm\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom lightgbm import LGBMClassifier\nfrom xgboost import XGBClassifier, plot_importance as plot_importance_xgb\nfrom lightgbm import LGBMClassifier, plot_importance as plot_importance_lgbm\n\n# Metrics and model evaluation\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import roc_curve, auc, confusion_matrix, accuracy_score, classification_report\n\n# Machine learning model\nmodels = { \n    \n    # model Logistic Regression\n    \"Logistic Regression\": LogisticRegression(),\n    \n    # model Naive bayes\n    \"Naive bayes\": GaussianNB(),\n    \n    # model KNN\n    \"KNN\": KNeighborsClassifier(),\n    \n    # model AdaBoost\n    \"Ada Boost\": AdaBoostClassifier(),\n    \n    # model Gradient Boosting\n    \"Gradient Boosting Classifier\": GradientBoostingClassifier(),\n    \n    # model Decision Tree Classifier - With adjusted internal parameters\n    \"Decision Tree Classifier\": DecisionTreeClassifier(max_depth=5, \n                                                       min_samples_split=2, \n                                                       random_state=105),\n    \n    # model XGBoost - With adjusted internal parameters\n    \"XGBoost\": XGBClassifier(n_estimators=100,\n                             max_depth=250,\n                             learning_rate=0.1,\n                             subsample=0.8,\n                             colsample_bytree=0.8,\n                             objective='multi:softmax',\n                             num_class=3,\n                             random_state=42,\n                             tree_method='gpu_hist'),\n    \n    # model LGBM - with adjusted internal parameters\n    \"LGBM\": LGBMClassifier(boosting_type='gbdt',\n                           bagging_freq=5,\n                           verbose=0,\n                           device='gpu',\n                           num_leaves=31,\n                           max_depth=250,\n                           learning_rate=0.1,\n                           n_estimators=100)}\n# Training models\nfor name, model in tqdm(models.items(), desc=\"Training models\", total=len(models)):\n    \n    # # Model training\n    model.fit(X_train, y_train)\n    \n    # Cross-validation\n    score_training = cross_val_score(model, X_train, y_train, cv=10)\n    \n    # Model forecast\n    pred_model = model.predict(X_test)\n    \n    # Viewing models\n    tqdm.write(\"Model: {} has Accuracy {:.2f}%\".format(model.__class__.__name__, round(score_training.mean(), 2) * 100))\n    print()","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:06:12.410087Z","iopub.execute_input":"2024-06-02T17:06:12.410414Z","iopub.status.idle":"2024-06-02T17:58:45.915288Z","shell.execute_reply.started":"2024-06-02T17:06:12.410388Z","shell.execute_reply":"2024-06-02T17:58:45.914376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Eight machine learning models were generated for this project: Logistic Regression, Naive Bayes, K-Nearest Neighbors (KNN), Decision Tree, AdaBoost, Gradient Boosting, XGBoost, and LightGBM. Each model was trained and evaluated using a specific dataset to identify the best-performing model. After evaluation, the LightGBM model stood out as the most effective, achieving an accuracy of 90%. This model not only showed the best accuracy but also proved to be robust in other performance metrics such as precision, recall, and F1-score, indicating its consistency and reliability in various situations. The second-best performance was by the XGBoost model, which achieved an accuracy of 84%. Although its accuracy is lower than that of LightGBM, XGBoost also demonstrated good results in other evaluation metrics. Additionally, a detailed analysis of each model's performance on different data subsets was conducted to verify their generalization and prevent overfitting. Based on this analysis, LightGBM proved its superiority not only in terms of accuracy but also in terms of generalization capability and stability. Therefore, considering all evaluation criteria, the LightGBM model showed the highest adherence and performance, making it the most recommended choice for future implementations in this context.","metadata":{}},{"cell_type":"markdown","source":"# 6.Model metrics","metadata":{}},{"cell_type":"code","source":"# Iterating over each model\nfor nome, modelo in models.items():\n\n    # Model training\n    modelo.fit(X_train, y_train)\n\n    # Prediction on test set\n    y_pred = modelo.predict(X_test)\n\n    print()\n    print(\"Machine Learning Model:\", nome)\n\n    # ROC curve\n    fpr, tpr, _ = roc_curve(y_test, modelo.predict_proba(X_test)[:,1])\n    roc_auc = auc(fpr, tpr)\n    print()\n    print()\n\n    # Plotting the ROC curve\n    print()\n    plt.figure()\n    plt.plot(fpr, tpr, color='darkorange', lw=2, label='ROC curve (area = %0.2f)' % roc_auc)\n    plt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--')\n    plt.xlim([0.0, 1.0])\n    plt.ylim([0.0, 1.05])\n    plt.xlabel('False Positive Rate')\n    plt.ylabel('True Positive Rate')\n    plt.title('Receiver Operating Characteristic - {}'.format(nome))\n    plt.legend(loc=\"lower right\")\n    plt.grid(False)\n    plt.show()\n    print()\n    print()\n    \n    # Accuracy\n    acc = accuracy_score(y_test, y_pred)\n    print(\"Accuracy:\", acc)\n    print()\n    print()\n\n    # Confusion Matrix\n    cm = confusion_matrix(y_test, y_pred)\n    print()\n    print()\n    \n    # Plotting the confusion matrix with Seaborn\n    plt.figure(figsize=(8, 6))\n    sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')\n    plt.title('Confusion Matrix - {}'.format(nome))\n    plt.xlabel('Foretold')\n    plt.ylabel('TRUE') \n    plt.xticks(ticks=[0.5, 1.5], labels=['Non-Protein', 'Protein'])\n    plt.yticks(ticks=[0.5, 1.5], labels=['Non-Protein', 'Protein'])\n    plt.show()\n    print()\n    print()\n\n    # Classification Report\n    print(\"Classification Report:\")\n    print(classification_report(y_test, y_pred))","metadata":{"execution":{"iopub.status.busy":"2024-06-02T17:58:45.916587Z","iopub.execute_input":"2024-06-02T17:58:45.917312Z","iopub.status.idle":"2024-06-02T18:04:43.876401Z","shell.execute_reply.started":"2024-06-02T17:58:45.917258Z","shell.execute_reply":"2024-06-02T18:04:43.875417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **1. Confusion matrix**\n\n**The confusion matrix provided for the LightGBM model presents the following values:**\n\n- True Negatives (Non-Protein correctly predicted): 5458\n\n- False Positives (Non-Protein predicted as Protein): 451\n\n- False Negatives (Protein predicted as Non-Protein): 867\n\n- True Positives (Protein correctly predicted): 5224\n\n**Analyzing these values, we can make the following observations:**\n\n**Accuracy:**\n\nAccuracy is the proportion of correct predictions (both positive and negative) over the total cases.\n\n- Accuracy = (True Positives + True Negatives) / Total\n\n- Accuracy = (5224 + 5458) / (5224 + 5458 + 867 + 451) = 10682 / 12000 ≈ 89.02%\n\n**Precision for the \"Protein\" class:**\n\n- Precision is the proportion of true positives over the total positive predictions.\n\n- Precision = True Positives / (True Positives + False Positives)\n\n- Precision = 5224 / (5224 + 451) ≈ 92.05%\n\n**Recall (Sensitivity) for the \"Protein\" class:**\n\n- Recall is the proportion of true positives over the total actual positive cases.\n\n- Recall = True Positives / (True Positives + False Negatives)\n\n- Recall = 5224 / (5224 + 867) ≈ 85.75%\n\n**F1-Score for the \"Protein\" class:**\n\n- F1-Score is the harmonic mean between precision and recall, giving a balanced measure of precision and recall.\n\n- F1-Score = 2 (Precision Recall) / (Precision + Recall)\n\n- F1-Score = 2 (0.9205 0.8575) / (0.9205 + 0.8575) ≈ 88.77%\n\n**Confusion Matrix Interpretation:**\n\n- The model performs well overall with a high accuracy and precision rate. However, there is a considerable amount of false negatives (867), indicating that the model sometimes fails to correctly identify the \"Protein\" class. The false positive rate (451) is relatively lower compared to false negatives.\n\n- In summary, the LightGBM model has shown excellent performance in terms of accuracy and precision. However, there may be room for improvement in reducing false negatives to increase recall and consequently the F1-Score.\n\n## **2. Classification Report**\n\n**Precision:** Precision is the proportion of instances classified as positive that are actually positive. For class 0, precision is 86%, meaning 86% of instances classified as 0 are actually 0. For class 1, precision is 92%.\n\n**Recall:** Recall is the proportion of actual positive instances that were correctly detected by the model. For class 0, recall is 92%, meaning the model correctly identifies 92% of instances that are actually 0. For class 1, recall is 86%.\n\n**F1-score:** F1-score is the harmonic mean of precision and recall. It's useful when classes are imbalanced. Both for class 0 and class 1, the F1-score is 89%, suggesting a good balance between precision and recall.\n\n**Support:** Support is the number of instances of each class in the dataset. There are 5909 instances of class 0 and 6091 instances of class 1.\n\n**Accuracy:** Accuracy is the proportion of all instances correctly classified by the model. In this case, accuracy is 89%, meaning the model correctly classifies 89% of all instances in the dataset.\n\n**Macro avg and Weighted avg:** These are the averages of the metrics for all classes. In the case of macro avg, it's the unweighted average of the metrics for each class. In the case of weighted avg, it's the weighted average of the metrics, where each class contributes a weight equal to its support. Both are 89% for precision, recall, and F1-score, reflecting consistent performance across all classes.\n\n## **3. ROC Curve**\n\n**Components of the ROC Curve**\nX-Axis (False Positive Rate - FPR): Represents the proportion of actual negatives that were incorrectly classified as positives by the model. It is calculated as:\n\nwhere FP is the number of false positives and TN is the number of true negatives.\n\n**Y-Axis (True Positive Rate - TPR):** Represents the proportion of actual positives that were correctly classified as positives by the model. It is also known as sensitivity or recall. It is calculated as:\n\n- where TP is the number of true positives and FN is the number of false negatives.\n\n**ROC Curve:** The orange line in the image represents the ROC curve of the LGBM model. Each point on the curve represents a (FPR, TPR) pair for a specific decision threshold.\n\n**Diagonal Line (Reference Line):** The dashed blue line represents the performance of a random classifier. This line corresponds to an area under the curve (AUC) of 0.5. If the model's ROC curve is close to this line, it indicates that the model is no better than a random classifier.\n\n**Area Under the Curve (AUC):** The AUC value summarizes the ROC curve into a single number. The higher the AUC, the better the model's performance. In the image, the AUC is 0.96, indicating that the LGBM model has excellent performance in distinguishing between positive and negative classes.\n\n**Interpretation of the ROC Curve**\n\n**Excellent Performance:** The AUC of 0.96 indicates that the model has excellent performance. An AUC of 1.0 represents a perfect classifier, while an AUC of 0.5 represents a random classifier.\n\n**High and Leftward ROC Curve:** The model's ROC curve is well above the diagonal line and close to the top-left corner of the plot. This indicates that the model has a high true positive rate (TPR) and a low false positive rate (FPR) for most decision thresholds.\n\n## Conclusion\n\nThe provided ROC curve demonstrates that the LGBM model has excellent performance in the binary classification task, with a high ability to distinguish between positive and negative classes, as indicated by the high AUC value of 0.96.","metadata":{}},{"cell_type":"markdown","source":"# 7. Model result","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import average_precision_score\n\n# Dictionary to store accuracy and average precision metrics\nresultados_metricas = {'Model': [], \n                       'Accuracy': [],\n                       'Average Precision': []\n                      }\n\n# Iterating over each model\nfor nome, modelo in models.items():\n    \n    # Model training\n    modelo.fit(X_train, y_train)\n\n    # Prediction on test set\n    y_pred = modelo.predict(X_test)\n\n    # Calculating accuracy\n    acc = accuracy_score(y_test, y_pred)\n    \n    # Calculating average precision\n    avg_precision = average_precision_score(y_test, y_pred)\n\n    # Storing accuracy and average precision results in the dictionary\n    resultados_metricas['Model'].append(nome)\n    resultados_metricas['Accuracy'].append(acc)\n    resultados_metricas['Average Precision'].append(avg_precision)\n\n# Creating DataFrame with the results\ndf_metricas = pd.DataFrame(resultados_metricas)\n\n# Displaying the DataFrame sorted by accuracy column in descending order\ndf_metricas_sorted = df_metricas.sort_values(by='Accuracy', ascending=False)\ndf_metricas_sorted","metadata":{"execution":{"iopub.status.busy":"2024-06-02T18:04:43.877503Z","iopub.execute_input":"2024-06-02T18:04:43.877770Z","iopub.status.idle":"2024-06-02T18:10:10.715164Z","shell.execute_reply.started":"2024-06-02T18:04:43.877746Z","shell.execute_reply":"2024-06-02T18:10:10.714138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**LGBM (Light Gradient Boosting Machine):**\n\nAccuracy: 0.890167 Average Precision: 0.861750 Analysis: This model has the highest accuracy and average precision, indicating that it performs the best among all the models in terms of both classification accuracy and precision. Logistic Regression:\n\nAccuracy: 0.866333 Average Precision: 0.824446 Analysis: Logistic Regression shows strong performance, being the second-best model with high accuracy and average precision, making it a reliable choice.\n\n**XGBoost:**\n\nAccuracy: 0.834333 Average Precision: 0.810085 Analysis: XGBoost also performs well with decent accuracy and precision, ranking third. It’s a robust choice for many classification tasks. Gradient Boosting Classifier:\n\nAccuracy: 0.830333 Average Precision: 0.804084 Analysis: Similar to XGBoost, this model has competitive performance, making it another good option for classification tasks.\n\n**KNN (K-Nearest Neighbors):**\n\nAccuracy: 0.798833 Average Precision: 0.727347 Analysis: KNN has lower accuracy and precision compared to the top models, indicating it may not generalize as well for this particular dataset.\n\n**AdaBoost:**\n\nAccuracy: 0.781833 Average Precision: 0.733364 Analysis: AdaBoost shows moderate performance but is not as effective as the top-performing boosting algorithms.\n\n**Naive Bayes:**\n\nAccuracy: 0.742417 Average Precision: 0.679529 Analysis: Naive Bayes has relatively low accuracy and precision, suggesting it may not capture the complexity of the data as well as other models.\n\n**Decision Tree Classifier:**\n\nAccuracy: 0.695167 Average Precision: 0.683770\n\nAnalysis: This model has the lowest accuracy and average precision, indicating it’s the least effective model for this classification task.\n\n## **Conclusion:** \n\n- LGBM is the best-performing model based on both accuracy and average precision, followed by Logistic Regression and XGBoost. Decision Tree Classifier and Naive Bayes are the least effective models for this task.","metadata":{}},{"cell_type":"markdown","source":"# 8. Saving models","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Importing libraries\nimport lightgbm as lgb\nfrom lightgbm import LGBMClassifier\n\n# Define the hyperparameters\nparams = {'boosting_type': 'gbdt',\n          'objective': 'binary',\n          'metric': 'binary_logloss',\n          'device': 'gpu',  # Usar GPU\n          'gpu_platform_id': 0,\n          'gpu_device_id': 0,\n         }\n\n# Create the model\nmodel_LGBMClassifier = LGBMClassifier(**params)\n\n# Train the model\nmodel_LGBMClassifier.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-06-02T18:10:10.716639Z","iopub.execute_input":"2024-06-02T18:10:10.717029Z","iopub.status.idle":"2024-06-02T18:10:19.856299Z","shell.execute_reply.started":"2024-06-02T18:10:10.716991Z","shell.execute_reply":"2024-06-02T18:10:19.855315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Use the model to make probability predictions\nprob_predictions = model_LGBMClassifier.predict_proba(X_test)\n\n# If the positive class is the second class, you would use prob_predictions[:, 1]\npositive_probabilities = prob_predictions[:, 1]\n\n# Create a pandas DataFrame with the predicted probabilities\ndf = pd.DataFrame(\n    {'Id': range(1, len(X_test) + 1), \n     'binds': positive_probabilities})\n\n# Model 1 - LGBM\n# Save the DataFrame to a CSV file\ndf.to_csv('/kaggle/working/submission_NGCDFG.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-06-02T18:10:19.857420Z","iopub.execute_input":"2024-06-02T18:10:19.857805Z","iopub.status.idle":"2024-06-02T18:10:21.539604Z","shell.execute_reply.started":"2024-06-02T18:10:19.857758Z","shell.execute_reply":"2024-06-02T18:10:21.538644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9.Conclusion\n\nIn this BELKA competition, the main objective is to create new drugs based on molecules. In this notebook, we performed data preprocessing using the One Hot Encoder technique, which transforms categorical data into numerical data. This step is crucial for machine learning models to work efficiently with the data.\n\nSubsequently, we divided the data into training and testing sets, ensuring that the model could be properly evaluated. The target variable for the model was carefully selected to ensure accurate predictions.\n\nWe tested eight different machine learning models to identify the most effective one. Among them, LightGBM stood out, achieving an accuracy of 89%. This result made LightGBM the most suitable model for the project due to its high precision.\n\nFor future improvements to the project, we suggest optimizing hyperparameters using advanced techniques such as Grid Search, Random Search, or Bayesian Optimization. Additionally, implementing pipelines and cross-validation could increase the robustness of the model. Another promising approach would be applying classification neural networks, which have the potential to offer even better accuracy.\n\nBesides the development and evaluation of the models, we also focused on visualizing the proteins generated by the LightGBM machine learning model. This visualization allows us to observe the predictions of the proteins and the molecules generated, providing valuable insights for the creation of new drugs.\n\nIn summary, this project not only demonstrated the effectiveness of the LightGBM model in predicting the efficacy of new molecules but also established a solid foundation for future improvements and applications in drug discovery.\n\n**--->Social networks<---**\n\nLinkedIn \n\nGitHub \n\nPortfolio ","metadata":{}},{"cell_type":"markdown","source":"# 10. Reference\n\nBlevins, A. D. (2024). Leash Tutorial: ECFPs and Random Forest. Available at: https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}