{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\nExploratory Data Analysis (EDA) is the process of exploring a dataset and gaining an understanding of its main characteristics. The dataprep.eda package makes this process easier by allowing users to explore important characteristics using simple APIs. Each API enables the user to analyse the dataset at various levels, from high to low, and from various perspectives.\n\n\n\n<img src=\"https://seleritysas.com/wp-content/uploads/2019/12/shutterstock_606583169.jpg\" width=\"700px\">","metadata":{}},{"cell_type":"markdown","source":"\n\n## Data Description\n### Files\nThis is a synchronous rerun code competition. The provided test set is a small representative set of files (copied from the training set) to demonstrate the format of the private test set. When you submit your notebook, Kaggle will rerun your code on the test set, which contains unseen images.\n\n* train.csv - the training set, contains full history of clinical information\n* test.csv - the test set, contains only the baseline measurement\n* train/ - contains the training patients' baseline CT scan in DICOM format\n* test/ - contains the test patients' baseline CT scan in DICOM format\n* sample_submission.csv - demonstrates the submission format\n\n<img src=\"https://drparthivshah.com/wp-content/uploads/2020/02/shutterstock_1326795209_0.jpg\" width=\"700px\">\n\n\n\nDataset Link\n\n\n\n##### [Here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data)","metadata":{}},{"cell_type":"code","source":"!pip install dataprep","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-05-25T11:58:48.539017Z","iopub.execute_input":"2021-05-25T11:58:48.539380Z","iopub.status.idle":"2021-05-25T11:58:55.785912Z","shell.execute_reply.started":"2021-05-25T11:58:48.539347Z","shell.execute_reply":"2021-05-25T11:58:55.784587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install autoviz","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-05-25T11:58:55.788067Z","iopub.execute_input":"2021-05-25T11:58:55.788476Z","iopub.status.idle":"2021-05-25T11:59:02.773724Z","shell.execute_reply.started":"2021-05-25T11:58:55.788437Z","shell.execute_reply":"2021-05-25T11:59:02.772199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install xlrd","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-05-25T11:59:02.776082Z","iopub.execute_input":"2021-05-25T11:59:02.776389Z","iopub.status.idle":"2021-05-25T11:59:09.438703Z","shell.execute_reply.started":"2021-05-25T11:59:02.776358Z","shell.execute_reply":"2021-05-25T11:59:09.436979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import Lab\nimport numpy as np\nimport pandas as pd\n\nimport seaborn as sns\nfrom dataprep.eda import *\nfrom dataprep.eda import plot\nfrom dataprep.eda import plot_correlation\nfrom dataprep.eda import plot_missing\n\nimport matplotlib.pyplot as plt\nimport warnings\nwarnings.filterwarnings('ignore')\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport seaborn as sns\nfrom dataprep.eda import *\nfrom dataprep.eda import plot\nfrom dataprep.eda import plot_correlation\nfrom dataprep.eda import plot_missing\n\nimport matplotlib.pyplot as plt\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:09.443253Z","iopub.execute_input":"2021-05-25T11:59:09.443589Z","iopub.status.idle":"2021-05-25T11:59:09.452633Z","shell.execute_reply.started":"2021-05-25T11:59:09.443555Z","shell.execute_reply":"2021-05-25T11:59:09.451501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\n\ndf_test = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:09.454536Z","iopub.execute_input":"2021-05-25T11:59:09.454919Z","iopub.status.idle":"2021-05-25T11:59:09.481403Z","shell.execute_reply.started":"2021-05-25T11:59:09.454886Z","shell.execute_reply":"2021-05-25T11:59:09.480218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:09.482829Z","iopub.execute_input":"2021-05-25T11:59:09.483155Z","iopub.status.idle":"2021-05-25T11:59:09.507907Z","shell.execute_reply.started":"2021-05-25T11:59:09.483122Z","shell.execute_reply":"2021-05-25T11:59:09.506822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:09.509462Z","iopub.execute_input":"2021-05-25T11:59:09.509787Z","iopub.status.idle":"2021-05-25T11:59:09.527033Z","shell.execute_reply.started":"2021-05-25T11:59:09.509754Z","shell.execute_reply":"2021-05-25T11:59:09.525770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alldata = pd.concat([df_train, df_test])\nalldata","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:09.530360Z","iopub.execute_input":"2021-05-25T11:59:09.530717Z","iopub.status.idle":"2021-05-25T11:59:09.557090Z","shell.execute_reply.started":"2021-05-25T11:59:09.530659Z","shell.execute_reply":"2021-05-25T11:59:09.555752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from autoviz.AutoViz_Class import AutoViz_Class\nAV = AutoViz_Class()\ntarget='SmokingStatus'\ndf = AV.AutoViz(filename=\"\",sep=',', depVar=target, dfte=alldata, header=0, verbose=1, \n                 lowess=False,  max_rows_analyzed=150000, max_cols_analyzed=30)","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:09.559447Z","iopub.execute_input":"2021-05-25T11:59:09.559859Z","iopub.status.idle":"2021-05-25T11:59:17.608525Z","shell.execute_reply.started":"2021-05-25T11:59:09.559823Z","shell.execute_reply":"2021-05-25T11:59:17.607190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate the Profiling Report\n\nimport pandas_profiling as pp\nprofile = pp.ProfileReport(\n    alldata, title=\"OSIC Pulmonary Fibrosis Progression\", html={\"style\": {\"full_width\": True}}, sort=\"None\"\n)","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:17.610614Z","iopub.execute_input":"2021-05-25T11:59:17.611080Z","iopub.status.idle":"2021-05-25T11:59:17.688661Z","shell.execute_reply.started":"2021-05-25T11:59:17.611026Z","shell.execute_reply":"2021-05-25T11:59:17.687594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"profile.to_widgets()","metadata":{"execution":{"iopub.status.busy":"2021-05-25T11:59:17.690203Z","iopub.execute_input":"2021-05-25T11:59:17.690626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"profile","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(alldata)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(alldata, 'SmokingStatus', 'Sex')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(alldata, 'Age', 'Percent')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(alldata, 'Age', 'FVC')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alldata","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alldata[\"Sex\"] = alldata[\"Sex\"].replace({\"Male\":1, 'Female':0})","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alldata[\"SmokingStatus\"] = alldata[\"SmokingStatus\"].replace({\"Ex-smoker\":2,\"Currently smokes\":1, 'Never smoked':0})","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alldata","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alldata.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(alldata)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}