{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Exploring the patients data and starter model","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"In this notebook, I will be going over the train.csv file and perform EDA as well as create a basic model using the patients data. In a realistic scenario, the pictures in the data of the competition should be used for modelling.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Importing necessary libraries\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\nfrom lightgbm import LGBMClassifier\nfrom sklearn.preprocessing import LabelEncoder\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Reading in the data\ndf = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Taking a look at the first 5 rows of data\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Getting the shape of the dataset\ndf.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are a total of 33126 rows of data and 8 columns including 2 target columns, i.e. benign_maginant and target.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df.dtypes.value_counts().sort_values(ascending=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Checking the data types, it is clear that most of the columns contain objects.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Now, seeing how many unique categorical types of data are in the columns having 'object' as a datatype.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df.select_dtypes('object').apply(pd.Series.nunique, axis = 0)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, finding the number of photos present of each patient.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['patient_id'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, the max number of images for a single patient is 115 and the minimum number of images for a single patient is 2.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['patient_id'].value_counts().mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The mean number of images for the 2056 patients is 16.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['target'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are large number of cases in the dataset that are benign and low number of cases in the dataset that are malignant. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Creating a helper function for further analysis.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_analysis(col_name, df, plot_kind='bar'):\n    \"\"\"\n    Function to plot two subplots containing Joe's and non-Joe's counts of data points for a given feature.\n    :param col_name: Column name of the feature to be analysed\n    :param df: DataFrame containing the source of data\n    :plot_kind: Line plot or Bar Plot\n    :return True: Boolean indicating that the analysis has been plotted\n    \"\"\"\n    \n    df_benign = df[df['target']==0]\n    df_malignant = df[df['target']!=0]\n    fig, axs = plt.subplots(2,figsize=(26,8))\n    fig.suptitle('Difference between ' + col_name + ' of patients in benign and malignant cases ')\n    axs[0].set_title('Benign')\n    axs[0].set_ylabel('Number of cases', fontsize=12)\n    axs[1].set_title('Malignant')\n    axs[1].set_ylabel('Number of cases', fontsize=12)\n    axs[1].set_xlabel(col_name, fontsize=12)\n    if plot_kind == 'line':\n        axs[0].plot(df_benign[col_name].value_counts().index, df_benign[col_name].value_counts().values)\n        axs[1].plot(df_malignant[col_name].value_counts().index, df_malignant[col_name].value_counts().values)\n    elif plot_kind == 'bar':\n        axs[0].bar(df_benign[col_name].value_counts().index, df_benign[col_name].value_counts().values)\n        axs[1].bar(df_malignant[col_name].value_counts().index, df_malignant[col_name].value_counts().values)\n    return True","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_analysis('sex', df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks the proportion of male and female patients is quite similar for the benign and malignant cases","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_analysis('anatom_site_general_challenge', df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The proportion of anatom_site_general_challenge is also very similar in proportion.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_analysis('diagnosis', df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are a large number of unknown diagnosis for benign cases but in malignant cases, it is all melanoma.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_analysis('age_approx', df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The mean age of all the patients in both cases is around 40-60.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Data Modelling","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Checking for null values show that there are some missing values for sex, age_approx and anatom_site_general_challenge. Filling the null values with the mostly observed value.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Performing Data pre-processing on the training data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def data_preparation(df, evaluation = False):\n    # Filling in the missing values\n    df['sex'] = df['sex'].fillna('male')\n    df['age_approx'] = df['age_approx'].fillna(df['age_approx'].mean())\n    df['anatom_site_general_challenge'] = df['anatom_site_general_challenge'].fillna('torso')\n\n    # Label encoding sex\n    labelencoder = LabelEncoder()\n    df['sex'] = labelencoder.fit_transform(df['sex'])\n\n    df = pd.get_dummies(df, columns=['anatom_site_general_challenge'])\n\n    if evaluation:\n        X_return = df.drop(['image_name', 'patient_id'], axis = 1)\n        return X_return\n    else:\n        X_return = df.drop(['image_name', 'patient_id', 'benign_malignant','target','diagnosis'], axis = 1)\n        return X_return, df['target']\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X, y = data_preparation(df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Train test split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, random_state=42, stratify = y)\n\n# Gradient Boosting Model\nclf = LGBMClassifier(\n            objective='binary',\n            n_estimators=100000,\n            num_leaves=10,\n            learning_rate=0.1,\n            max_depth=16,\n            subsample_for_bin= 200000,\n            subsample=1,\n            subsample_freq= 200,\n            silent=-1,\n            verbose=-1,\n            min_split_gain=0.0001,\n            min_child_samples=800,\n            )\n\n# Training the model\nclf.fit(X_train, y_train, eval_set=[(X_train, y_train), (X_test, y_test)], eval_metric='auc', verbose=200, early_stopping_rounds=1000)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df  = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Performing data pre-processing","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Preprocessing test data\neval_X = data_preparation(test_df, evaluation = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Getting the prediction probability\nprediction_list = clf.predict_proba(eval_X)\nfinal_pred_list = [a[1] for a in prediction_list]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Appending to the test dataframe\ntest_df['target'] = final_pred_list","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Getting the final output file","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[['image_name','target']].to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"I will try my best to use the pictures next to get better accuracy for this competition.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}