{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":46105,"databundleVersionId":5087314,"sourceType":"competition"},{"sourceId":5127677,"sourceType":"datasetVersion","datasetId":2978370}],"dockerImageVersionId":30407,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"color:white; \n            background-color:white; \n            background-image:url('https://www.nicepng.com/png/full/392-3922851_28-collection-of-sign-language-clipart-pictures-sign.png');\n            background-position: center; \n            background-repeat: no-repeat; \n            background-size: cover;\n            height:300px;\">\n    <span style=\"color:white\">  g  </span>\n    <h1 style=\"color:black; font-size:4em; margin: 1% auto 1% 5%\">Isolated Sign Language Recognition</h1>\n    <h2 style=\"color:black; font-size:2em; margin: 5% auto 1% 7%\">Step 2 - Data Preprocessing</h2>\n    <span style=\"color:white\">   </span>\n</div>\n\n","metadata":{}},{"cell_type":"markdown","source":"**LINK**\n\nDOWNLOAD THIS DATASET : https://www.kaggle.com/datasets/cristaliss/islr-hands-only-dataset","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:24.697293Z","iopub.execute_input":"2023-03-10T11:41:24.698317Z","iopub.status.idle":"2023-03-10T11:41:24.704066Z","shell.execute_reply.started":"2023-03-10T11:41:24.69827Z","shell.execute_reply":"2023-03-10T11:41:24.702455Z"}}},{"cell_type":"markdown","source":"# 1. Introduction\n## Sign Language Recognition\nSign language recognition is a field of research that seeks to enable machines to understand and interpret sign language. Sign language relies on gestures and facial expressions to convey meaning, and is used by people with hearing impairments to communicate. \n\nBy leveraging machine learning algorithms and computer vision techniques, researchers have been able to develop models that can recognize and interpret sign language. These models have a variety of applications in educational, healthcare, and other settings, and are helping to bridge the communication gap between the deaf and hearing communities.\n\n## Preparing data\n\nprepare and preprocess data machine learning importance\nPreprocessing and preparing data is an important step in any machine learning project, as it helps to ensure that the data is of a high quality and ready to be used with a machine learning model. Data preprocessing involves cleaning, formatting, and transforming the data so that it is ready to be used in a machine learning algorithm. This is important because the algorithms used in machine learning are designed to work with a specific type of data. Without preprocessing, the machine learning algorithms may not be able to make accurate predictions or provide the insights needed. Additionally, preprocessing can help to reduce the amount of noise and bias in the data, which can improve the accuracy of the machine learning models.","metadata":{}},{"cell_type":"markdown","source":"# 2. Installing and importing dependencies","metadata":{}},{"cell_type":"markdown","source":"\nConstants and libraries are essential components of any Python program. Constants are variables whose value cannot be changed and libraries are collections of code that can be reused to perform specific tasks. \n\nIn Python, constants are usually declared and assigned in a module and libraries are initialized with the import statement. Using constants and libraries can help improve the readability and maintainability of a program, as well as help to ensure that code is being reused efficiently.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os \nimport json\nfrom tqdm import tqdm\n\n# Visualize\nimport sys\nimport csv\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:39.972571Z","iopub.execute_input":"2023-03-10T11:41:39.972952Z","iopub.status.idle":"2023-03-10T11:41:39.978361Z","shell.execute_reply.started":"2023-03-10T11:41:39.972918Z","shell.execute_reply":"2023-03-10T11:41:39.977238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ASL_PATH = \"../input/asl-signs/\"\n\nTRAIN_PATH = os.path.join(ASL_PATH,'train.csv')\n\nSIGN_MAPPING_PATH = os.path.join(ASL_PATH,'sign_to_prediction_index_map.json')\n\nLANDMARKS_BASE_PATH = os.path.join(ASL_PATH,'train_landmark_files')","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:40.147268Z","iopub.execute_input":"2023-03-10T11:41:40.147928Z","iopub.status.idle":"2023-03-10T11:41:40.153205Z","shell.execute_reply.started":"2023-03-10T11:41:40.14789Z","shell.execute_reply":"2023-03-10T11:41:40.152027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Loading data\n\nBefore training a model, it is essential to first load the necessary data. This data may come from a variety of sources such as files, databases, or other external sources. In order to get the most out of the model, it is important to make sure that the data is of high quality and is in a format that is suitable for the model. This may involve pre-processing steps such as cleaning, normalizing, or transforming the data. Once the data is loaded, it can then be used to train the model by providing it with the necessary inputs and outputs. After the model is trained, it can then be tested against new data to ensure that it is performing as expected.\n\nIn this case, we are going to load *train.csv* file, which contains our mapping between the action occurred in a sequence (*sequence_id*) and the sign (*expected sign to predict*).\n\nThen, we will load the *sign_to_prediction_index_map.json*, which match a number to each type of sign.\nFinally, some *sequences* will be load. We will try to make a visual approach to them.","metadata":{}},{"cell_type":"code","source":"def read_dict(file_path):\n    path = os.path.expanduser(file_path)\n    with open(path, \"r\") as f:\n        dict_json = json.load(f)\n    return dict_json\n\ndef get_data(data_path):\n    train_data = pd.read_csv(f'{data_path}train.csv')\n    \n    sign_codes = read_dict(f'{data_path}sign_to_prediction_index_map.json')\n    sign_codes = dict([(sign_codes[key], key) for key in sign_codes])\n    \n    sequence_sample = pd.read_parquet(f'{data_path}train_landmark_files/2044/1001950812.parquet')\n    \n    mediapipe_holistic_data = []\n    \n    #JUST 300 -> IDEA OF https://www.kaggle.com/code/renzophellan/understanding-the-data-creating-a-model\n    for row in tqdm(train_data.head(3).itertuples()):\n        parquet_path = os.path.join(data_path, row.path)\n        data = pd.read_parquet(parquet_path)\n        mediapipe_holistic_data.append(data)\n    \n    #mediapipe_holistic_data =  pd.concat(mediapipe_holistic_data)\n    return train_data, sign_codes, mediapipe_holistic_data","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:40.495523Z","iopub.execute_input":"2023-03-10T11:41:40.496131Z","iopub.status.idle":"2023-03-10T11:41:40.504096Z","shell.execute_reply.started":"2023-03-10T11:41:40.496093Z","shell.execute_reply":"2023-03-10T11:41:40.503074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training data","metadata":{}},{"cell_type":"markdown","source":"To read *train.csv* file, you can use the *read_csv* function from the *pandas* library. This function will read the contents of the CSV file into a dataframe, which is a two-dimensional data structure. Once the dataframe is created, you can then access various columns and rows in the dataset. Here is an example of how to use readcsv:\n\n```python\nimport pandas as pd\n\ndf = pd.read_csv('train.csv')\n```\n\nThis will read the contents of the train.csv file into a dataframe called 'df'. You can then access the elements in the dataframe using the column and row indices. You can also use the head() and tail() methods to preview the first or last few rows in the dataframe.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(TRAIN_PATH, encoding='utf8',header=0)","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:41.948513Z","iopub.execute_input":"2023-03-10T11:41:41.949184Z","iopub.status.idle":"2023-03-10T11:41:42.046334Z","shell.execute_reply.started":"2023-03-10T11:41:41.949146Z","shell.execute_reply":"2023-03-10T11:41:42.045358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Sign-index mapping data","metadata":{}},{"cell_type":"markdown","source":"The json.load() function is used to read a JSON file and convert its contents into a Python dictionary. To use it, you would open the file, then pass the open file object to the json.load() function. Here is an example of how to do this:\n\n```python\nimport json\n\n# open the json file\nwith open(\"data.json\") as json_file:\n    # load the json data\n    data = json.load(json_file)\n    # do something with the data\n```\n\nThe json.load() function will return the contents of the file as a Python dictionary, which you can then access and manipulate.","metadata":{}},{"cell_type":"code","source":"def read_dict(path):\n    with open(path, \"r\") as f:\n        dict_json = json.load(f)\n    return dict_json","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:42.876934Z","iopub.execute_input":"2023-03-10T11:41:42.877567Z","iopub.status.idle":"2023-03-10T11:41:42.883058Z","shell.execute_reply.started":"2023-03-10T11:41:42.877527Z","shell.execute_reply":"2023-03-10T11:41:42.881678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sign_mapping = read_dict(SIGN_MAPPING_PATH)\nsign_mapping = dict([(key, sign_mapping[key]) for key in sign_mapping])","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:41:55.094142Z","iopub.execute_input":"2023-03-10T11:41:55.095039Z","iopub.status.idle":"2023-03-10T11:41:55.102195Z","shell.execute_reply.started":"2023-03-10T11:41:55.094988Z","shell.execute_reply":"2023-03-10T11:41:55.101194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Parquet Landmark Data","metadata":{}},{"cell_type":"markdown","source":"Now, we are loading parquet landmark data. These files contains Mediapide's Holistic Data\n\nMediaPipe's Holistic Landmark Data provides a set of landmarks that can be used to accurately identify and track a hand's position (face, pose... as well), orientation, and size in 3D space. The data is generated using a deep learning model that has been trained on a large, publicly available dataset of 3D images. The data can be used to create 3D models, track a user's hands/pose/face in 3D space, and even generate 3D augmented reality applications. Additionally, MediaPipe Holistic's data can be used to generate accurate cropping of the face, data augmentation processes, and more.\n\n\nLet's examine and example!!","metadata":{}},{"cell_type":"code","source":"sequence_sample = pd.read_parquet(f'{LANDMARKS_BASE_PATH}/2044/1001950812.parquet')\nsequence_sample","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.359946Z","iopub.execute_input":"2023-03-08T09:33:14.360326Z","iopub.status.idle":"2023-03-08T09:33:14.407019Z","shell.execute_reply.started":"2023-03-08T09:33:14.360281Z","shell.execute_reply":"2023-03-08T09:33:14.405864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. EDA and preprocessing process\n\nExploratory Data Analysis (EDA) and Data Pre-processing are two important processes when working with data. EDA involves analyzing a dataset to gain insights and identify patterns, while data pre-processing involves transforming and preparing data for analysis or machine learning models. EDA can involve visualizing data, computing descriptive statistics, and testing assumptions of the data. Pre-processing includes data cleaning, normalization, feature engineering, and feature selection. Both processes are necessary for ensuring the data is of good quality and can be used for machine learning models.","metadata":{}},{"cell_type":"markdown","source":"## Train file","metadata":{}},{"cell_type":"code","source":"train_df.head(n=5)","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.408805Z","iopub.execute_input":"2023-03-08T09:33:14.409556Z","iopub.status.idle":"2023-03-08T09:33:14.421937Z","shell.execute_reply.started":"2023-03-08T09:33:14.409512Z","shell.execute_reply":"2023-03-08T09:33:14.420632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.424147Z","iopub.execute_input":"2023-03-08T09:33:14.42489Z","iopub.status.idle":"2023-03-08T09:33:14.446672Z","shell.execute_reply.started":"2023-03-08T09:33:14.424848Z","shell.execute_reply":"2023-03-08T09:33:14.445356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, there are **no NULL values** in the data set. This means that all of the data points are valid and can be used for analysis, making your analysis more accurate and reliable.\n\nOther information we can get from watching this data is **column**'s info:\n- ***path*** (string): path to the parquet file of each row\n- ***participant_id*** (integer): id of the person who participate in the sequence\n- ***sequence_id*** (integer): id of the sequence\n- ***sign*** (string): sign made in the sequence ","metadata":{}},{"cell_type":"markdown","source":"The Pandas method **describe()** is used to calculate some statistical data like percentile, mean, and standard deviation of the numerical values of a Series or DataFrame. It produces a numerical summary of the data, including count, mean, standard deviation, minimum, maximum, 25th percentile, 50th percentile (median), 75th percentile, and interquartile range. The default output for describe() is a table with the data frame columns across the top, and the statistics across the side and over the rows.","metadata":{}},{"cell_type":"code","source":"train_df.describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.448384Z","iopub.execute_input":"2023-03-08T09:33:14.449203Z","iopub.status.idle":"2023-03-08T09:33:14.527383Z","shell.execute_reply.started":"2023-03-08T09:33:14.449164Z","shell.execute_reply":"2023-03-08T09:33:14.526112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.participant_id.unique().shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.530521Z","iopub.execute_input":"2023-03-08T09:33:14.531007Z","iopub.status.idle":"2023-03-08T09:33:14.539998Z","shell.execute_reply.started":"2023-03-08T09:33:14.530965Z","shell.execute_reply":"2023-03-08T09:33:14.538731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.sequence_id.unique().shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.541606Z","iopub.execute_input":"2023-03-08T09:33:14.542408Z","iopub.status.idle":"2023-03-08T09:33:14.555904Z","shell.execute_reply.started":"2023-03-08T09:33:14.542367Z","shell.execute_reply":"2023-03-08T09:33:14.554687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Conclusions\n\n\nFrom the previos information we can conclude the following:\n- There are **250 unique signs**\n- There are **21 unique participants**\n- Each **sequence_id is different** from others","metadata":{}},{"cell_type":"markdown","source":"Now, it is showed the most frequent signs of the dataset.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\ntrain_df['sign'].value_counts().head(20).plot(ax=ax, kind='bar',\n                                             title=\"Top 20 Frequent Signs\")","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.557884Z","iopub.execute_input":"2023-03-08T09:33:14.55844Z","iopub.status.idle":"2023-03-08T09:33:14.883534Z","shell.execute_reply.started":"2023-03-08T09:33:14.558401Z","shell.execute_reply":"2023-03-08T09:33:14.881323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Landmark data","metadata":{}},{"cell_type":"code","source":"def get_landmark_data(train_df, path,n=20):\n    mediapipe_holistic_data = []\n    \n    for row in tqdm(train_df.head(n).itertuples()):\n        parquet_path = os.path.join(path, row.path)\n        data = pd.read_parquet(parquet_path)\n        mediapipe_holistic_data.append(data)\n    \n    return mediapipe_holistic_data","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.884973Z","iopub.execute_input":"2023-03-08T09:33:14.885615Z","iopub.status.idle":"2023-03-08T09:33:14.892434Z","shell.execute_reply.started":"2023-03-08T09:33:14.885575Z","shell.execute_reply":"2023-03-08T09:33:14.891216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmarks_data = get_landmark_data(train_df, ASL_PATH)","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:14.894093Z","iopub.execute_input":"2023-03-08T09:33:14.894591Z","iopub.status.idle":"2023-03-08T09:33:15.147932Z","shell.execute_reply.started":"2023-03-08T09:33:14.894491Z","shell.execute_reply":"2023-03-08T09:33:15.146877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmarks_data[6]","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:15.150189Z","iopub.execute_input":"2023-03-08T09:33:15.151168Z","iopub.status.idle":"2023-03-08T09:33:15.170719Z","shell.execute_reply.started":"2023-03-08T09:33:15.15111Z","shell.execute_reply":"2023-03-08T09:33:15.169745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Creating a NEW dataset\n\nCreating a new dataset from original data involves several steps. First, you need to collect the data from its original source, either through manual data entry or by importing a file. Once the data is collected, it can then be cleaned and organized into a format that is suitable for the task at hand. This may involve transforming the data into a more convenient structure, such as a data frame or matrix, or performing data aggregation. Finally, any additional calculations or processing can be done in order to prepare the data for analysis.","metadata":{}},{"cell_type":"code","source":"train_df.iloc[0]","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:15.174498Z","iopub.execute_input":"2023-03-08T09:33:15.174769Z","iopub.status.idle":"2023-03-08T09:33:15.182537Z","shell.execute_reply.started":"2023-03-08T09:33:15.174743Z","shell.execute_reply":"2023-03-08T09:33:15.181417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are dropping pose and face coordenates. \nDropping face and pose data from dataset can be beneficial for a number of reasons. Firstly, it can reduce the size of the dataset, which can save storage space and reduce computational time when processing the dataset. Additionally, it can reduce the complexity of the dataset, making it easier to analyze and interpret. Furthermore, dropping this data can also reduce the risk of a data breach, as face and pose data can be particularly sensitive. Finally, if the data is not relevant to the task at hand, dropping it can help to reduce the noise in the dataset and improve the accuracy of the results.","metadata":{}},{"cell_type":"code","source":"landmarks_data[0].type.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:15.184004Z","iopub.execute_input":"2023-03-08T09:33:15.184644Z","iopub.status.idle":"2023-03-08T09:33:15.196617Z","shell.execute_reply.started":"2023-03-08T09:33:15.184604Z","shell.execute_reply":"2023-03-08T09:33:15.195638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can check, almost all data is from the face and pose. \n\nPose and face are both important elements in sign language. The body position and facial expressions of the signer convey a lot of information, including emotions, emphasis, and emphasis on certain words or phrases. This helps to make sign language more expressive and communicative. Additionally, the pose and facial expressions can be used to differentiate between different signs and to provide a more accurate representation of what is being said.\n\nHowever, if we want to make a quickly training, as these two types are the most we have in the dataset, can overshadow hands data.\n\nThat is why we are going to drop their rows, just to try if training can run faster and without bias.\n","metadata":{}},{"cell_type":"code","source":"def get_landmark_by_index(train_df, index):\n    parquet_path = train_df.iloc[index].path\n    parquet_path = os.path.join(ASL_PATH, parquet_path)\n    data = pd.read_parquet(parquet_path)\n    data = data[(data.type!='face') & (data.type!='pose')].reset_index()\n    data = data.loc[:,['frame','type','landmark_index','x','y','z']]\n    \n    data = data.groupby(['type','landmark_index']).std()\n    data.fillna(0, inplace=True)\n    data = data.apply(list).reset_index()\n    data = data.drop(columns=['type','landmark_index','frame'], axis=1)\n    \n    return data","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:15.199058Z","iopub.execute_input":"2023-03-08T09:33:15.199559Z","iopub.status.idle":"2023-03-08T09:33:15.208382Z","shell.execute_reply.started":"2023-03-08T09:33:15.199519Z","shell.execute_reply":"2023-03-08T09:33:15.207222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The previous cell makes the preprocessing of data:\n1. read file \n2. drop face and pose data\n3. group by type and landmark to calculate standard deviation\n4. drop not interesting columns","metadata":{}},{"cell_type":"code","source":"X = np.array([get_landmark_by_index(train_df, i) \n              for i, row in train_df.iterrows()], dtype=object)\n\ny = np.array(train_df.sign.map(sign_mapping))","metadata":{"execution":{"iopub.status.busy":"2023-03-08T09:33:15.209905Z","iopub.execute_input":"2023-03-08T09:33:15.210355Z","iopub.status.idle":"2023-03-08T10:25:32.839402Z","shell.execute_reply.started":"2023-03-08T09:33:15.210319Z","shell.execute_reply":"2023-03-08T10:25:32.838308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we apply all this process to all our training dataset.","metadata":{}},{"cell_type":"code","source":"print(f'X shape: {X.shape} ; y shape: {y.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-03-08T10:25:32.841083Z","iopub.execute_input":"2023-03-08T10:25:32.841455Z","iopub.status.idle":"2023-03-08T10:25:32.850608Z","shell.execute_reply.started":"2023-03-08T10:25:32.841416Z","shell.execute_reply":"2023-03-08T10:25:32.849556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X[0]","metadata":{"execution":{"iopub.status.busy":"2023-03-08T10:25:32.852797Z","iopub.execute_input":"2023-03-08T10:25:32.853089Z","iopub.status.idle":"2023-03-08T10:25:32.864016Z","shell.execute_reply.started":"2023-03-08T10:25:32.853046Z","shell.execute_reply":"2023-03-08T10:25:32.863018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y","metadata":{"execution":{"iopub.status.busy":"2023-03-08T10:25:32.865356Z","iopub.execute_input":"2023-03-08T10:25:32.866246Z","iopub.status.idle":"2023-03-08T10:25:32.876726Z","shell.execute_reply.started":"2023-03-08T10:25:32.866203Z","shell.execute_reply":"2023-03-08T10:25:32.875578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Now we have our dataset!!!!**\n\nDatasets that have been processed and cleaned can offer a wide range of benefits. For example, they can provide insights into patterns and trends that would otherwise remain hidden, and they can be easily used for advanced analytics, predictive modeling, and machine learning applications. Additionally, they can help reduce the amount of manual work required to clean and prepare data for analysis, as well as help to ensure data accuracy and consistency. Finally, processed and cleaned datasets can also help to reduce the amount of time needed to analyze data, leading to faster time to insights.","metadata":{}},{"cell_type":"markdown","source":"# 5. Save dataset","metadata":{}},{"cell_type":"code","source":"with open('/kaggle/working/X.npy', 'wb') as f:\n    np.save(f, X)\n\nwith open('/kaggle/working/y.npy', 'wb') as f:\n    np.save(f, y)","metadata":{"execution":{"iopub.status.busy":"2023-03-08T10:25:32.878088Z","iopub.execute_input":"2023-03-08T10:25:32.878643Z","iopub.status.idle":"2023-03-08T10:25:33.848714Z","shell.execute_reply.started":"2023-03-08T10:25:32.878599Z","shell.execute_reply":"2023-03-08T10:25:33.847265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open(\"/kaggle/working/y_index_mapping.json\", \"w\") as fp:\n    json.dump(sign_mapping, fp, indent=4)","metadata":{"execution":{"iopub.status.busy":"2023-03-10T11:47:43.690371Z","iopub.execute_input":"2023-03-10T11:47:43.691068Z","iopub.status.idle":"2023-03-10T11:47:43.696554Z","shell.execute_reply.started":"2023-03-10T11:47:43.691029Z","shell.execute_reply":"2023-03-10T11:47:43.695313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 👍 Thumbs up if you liked it","metadata":{}}]}