{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"# 📔**Competition goal**: For each **id** in the **evaluation set**, you should predict a value for **each of the 18,211 🧬genes🧬** named in the remaining columns. Each id corresponds to a c**ell_type / sm_name pair**, which you may identify from the id_map.csv file.","metadata":{}},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"code","source":"library(tidyverse) # metapackage of all tidyverse packages\nlibrary(arrow) # for reading parquet files\nlibrary(tictoc) # for tracking how long operations take\nlibrary(ggthemes) # for ggplot themes\npaste('There are ', length(list.files(path = '../input/open-problems-single-cell-perturbations')), 'files to read.')\nlist.files(path = '../input/open-problems-single-cell-perturbations')","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","execution":{"iopub.status.busy":"2023-09-26T23:37:55.152346Z","iopub.execute_input":"2023-09-26T23:37:55.154735Z","iopub.status.idle":"2023-09-26T23:37:57.134566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the .csv files\ntic('Reading the .csv files')\nadata_obs_meta <- read.csv('../input/open-problems-single-cell-perturbations/adata_obs_meta.csv')\nid_map <- read.csv('../input/open-problems-single-cell-perturbations/id_map.csv')\nmultiome_obs_meta <- read.csv('../input/open-problems-single-cell-perturbations/multiome_obs_meta.csv')\nmultiome_var_meta <- read.csv('../input/open-problems-single-cell-perturbations/multiome_var_meta.csv')\nsample_submission <- read.csv('../input/open-problems-single-cell-perturbations/sample_submission.csv')\ntoc()\n\npaste('The .csv files combined take up ', format(object.size(adata_obs_meta) + object.size(id_map) + object.size(multiome_obs_meta) + object.size(multiome_var_meta) + object.size(sample_submission), unit = 'MB'), 'in RAM')\n\n# Reading the .csv files took ~ 7 seconds and they take up ~ 125 MB in RAM","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:37:57.137443Z","iopub.execute_input":"2023-09-26T23:37:57.170745Z","iopub.status.idle":"2023-09-26T23:38:09.829604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The three **parquet files**, not surprisingly, take up a lot of space in **RAM**.  For now, we will read them in, determine how long it takes, and how much RAM they consume.  Then, we will remove them from memory.  This will allow us to continue with the rest of the notebook and not bust our RAM limit.  Future versions of this notebook will have any modifications to those files **re-written as output for future use**. ","metadata":{}},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we know initially about the **adata_train.parquet** file.\n\nUnaggregated count and normalized data in COO sparse-array format. A supplement to de_train. In addition to the fields in de_train, this data also has:\n\n* obs_id - This is a unique identifier assigned to each cell in the raw dataset.\n* gene - Corresponds to the columns of de_train.\n* count - The raw molecular counts for the gene expression data measured in the experiment as output by 10x CellRanger.\n* normalized_count - These counts have been library size normalized and log(X+1) transformed.","metadata":{}},{"cell_type":"code","source":"options(arrow.skip_nul = TRUE) # required to deal with null values, otherwise errors out during parquet read\ntic('Reading adata_train.parquet')\nadata_train.parquet <- read_parquet('../input/open-problems-single-cell-perturbations/adata_train.parquet', as_data_frame = TRUE)\ntoc()\nformat(object.size(adata_train.parquet), unit = 'Gb')\n\n# The reading of this .parquet file took ~ 2 minutes, and it consumes ~ 10.9 GB of RAM when read as a data_frame.","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:38:09.842040Z","iopub.execute_input":"2023-09-26T23:38:09.843578Z","iopub.status.idle":"2023-09-26T23:46:51.930795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 416442312 rows!\nadata_train.parquet %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:46:51.933831Z","iopub.execute_input":"2023-09-26T23:46:51.935410Z","iopub.status.idle":"2023-09-26T23:46:51.954324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Here's the structure of our data\nadata_train.parquet %>% str()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:46:51.957955Z","iopub.execute_input":"2023-09-26T23:46:51.959975Z","iopub.status.idle":"2023-09-26T23:46:51.995122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 240090 unique cells in the raw dataset\nadata_train.parquet %>% select(obs_id) %>% unique() %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:46:51.999004Z","iopub.execute_input":"2023-09-26T23:46:52.001130Z","iopub.status.idle":"2023-09-26T23:50:40.300293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of count values.  With 416442312 rows of data already in RAM, trying to plot the distribution of all those count values explodes the RAM usage and busts the RAM 30 GB limit.  \n# But, we can randomly sample 1,000 of those values for a decent look at the distribution:\n# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =8)\ndata.frame(count = sample(adata_train.parquet$count,size = 1000,replace=FALSE)) %>% ggplot(aes(x=count)) + geom_histogram(fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:40.303135Z","iopub.execute_input":"2023-09-26T23:50:40.304732Z","iopub.status.idle":"2023-09-26T23:50:41.730117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We'll do the same thing with the normalized_count to see it's distribution\n# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =8)\ndata.frame(normalized_count = sample(adata_train.parquet$normalized_count,size = 1000,replace=FALSE)) %>% ggplot(aes(x=normalized_count)) + geom_histogram(fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:41.732843Z","iopub.execute_input":"2023-09-26T23:50:41.734286Z","iopub.status.idle":"2023-09-26T23:50:42.143870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For now, we will remove this item so we can continue to read the other files and not bust our RAM limitations.  \n# Future versions will modify & re-save the .parquet file so we can manage RAM usage.\nrm(adata_train.parquet)\ngc() # not strictly necessary, but noted by the gc() documentation as potentially useful after removing large objects, as in this case ","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:42.147404Z","iopub.execute_input":"2023-09-26T23:50:42.148953Z","iopub.status.idle":"2023-09-26T23:50:44.888756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we know initially about the **de_train.parquet** file.  It is an aggregated differential expression data in dense array format.\n\n* genes A1BG, A1BG-AS1, …, ZZEF1 (numbering 18,211 in all) - Differential expression value (-log10(p-value) * sign(LFC)) for each gene.\n* cell_type - The annotated cell type of each cell based on RNA expression.\n* sm_name - The primary name for the (parent) compound (in a standardized representation) as chosen by LINCS. This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n* sm_lincs_id - The global LINCS ID (parent) compound (in a standardized representation). This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n* SMILES - Simplified molecular-input line-entry system (SMILES) representations of the compounds used in the experiment. This is a 1D representation of molecular structure. These SMILES are provided by Cellarity based on the specific compounds ordered for this experiment.\n* control - Boolean indicating whether this instance was used as a control.\n\nThe **de_train.parquet** file comprises the main competition data. It contains values for a number of cell_type / sm_name pairs. Your **goal is to predict corresponding values for the cell_type / sm_name pairs given in id_map.csv**.","metadata":{}},{"cell_type":"code","source":"tic('Reading de_train.parquet')\nde_train.parquet <- read_parquet('../input/open-problems-single-cell-perturbations/de_train.parquet', as_data_frame = TRUE)\ntoc()\nformat(object.size(de_train.parquet), unit = 'Gb')\n\n# The reading of this .parquet file took ~ 1 second, and it consumes ~ 0.1 GB of RAM.  Nice!  That's not bad at all. ","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:44.891542Z","iopub.execute_input":"2023-09-26T23:50:44.892976Z","iopub.status.idle":"2023-09-26T23:50:46.577103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Only 614 rows of data in this parquet file\nde_train.parquet %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:46.579739Z","iopub.execute_input":"2023-09-26T23:50:46.581178Z","iopub.status.idle":"2023-09-26T23:50:46.599004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is what our data looks like\nde_train.parquet %>% head()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:46.601715Z","iopub.execute_input":"2023-09-26T23:50:46.603208Z","iopub.status.idle":"2023-09-26T23:50:55.812724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's take a look at the **average value** for one of the genes (**A1BG**), and plot its average value vs. the sm_name.","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 20, repr.plot.height =15)\nde_train.parquet %>% group_by(sm_name) %>% mutate(avg_A1BG = mean(A1BG)) %>% select(sm_name, avg_A1BG) %>% unique() %>% \nggplot(aes(x=reorder(sm_name,avg_A1BG))) + geom_col(aes(y=avg_A1BG), fill = 'steelblue') + coord_flip() + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:55.815607Z","iopub.execute_input":"2023-09-26T23:50:55.817246Z","iopub.status.idle":"2023-09-26T23:50:57.658549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For consistency, we'll remove this df for now and future versions we'll see if we need to modify it at all, and then re-write changes.\nrm(de_train.parquet)\ngc() ","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:57.661938Z","iopub.execute_input":"2023-09-26T23:50:57.663474Z","iopub.status.idle":"2023-09-26T23:50:58.087176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we know about the **multiome_train.parquet** file -- it is **optional** additional **10x Multiome** data for each sample at baseline.\n\n* obs_id - Unique identifier for each observation. (Distinct from identifiers used in adata.)\n* location - This is a feature ID. If the feature_type in multiome_var_meta.csv is Gene Expression then this is a gene symbol. If feature_type is Peaks, then this is the genomic interval of the peak.\n* count - This is the raw molecular counts of the transcript for accessible DNA measurement as output by Cellranger-Arc.\n* normalized_count - If the feature_type in multiome_var_meta.csv is Gene Expression then this is library size normalized and log(X+1) transformed counts. If feature_type is Peaks, then this is ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF).","metadata":{}},{"cell_type":"code","source":"tic('Reading multiome_train.parquet')\nmultiome_train.parquet <- read_parquet('../input/open-problems-single-cell-perturbations/multiome_train.parquet', as_data_frame = TRUE)\ntoc()\nformat(object.size(multiome_train.parquet), unit = 'Gb')\n\n# The reading of this .parquet file took ~ 82 seconds, and it consumes ~ 5.7 GB of RAM when read as a dataframe.  ","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:50:58.089856Z","iopub.execute_input":"2023-09-26T23:50:58.091322Z","iopub.status.idle":"2023-09-26T23:58:21.194544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# There are 216251368 rows of data\nmultiome_train.parquet %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:58:21.197527Z","iopub.execute_input":"2023-09-26T23:58:21.198974Z","iopub.status.idle":"2023-09-26T23:58:21.215370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is what our dataset looks lie\nmultiome_train.parquet %>% head()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:58:21.218012Z","iopub.execute_input":"2023-09-26T23:58:21.219448Z","iopub.status.idle":"2023-09-26T23:59:52.524991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's look at the distribution of **normalized_count**","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =8)\nmultiome_train.parquet %>% ggplot(aes(x=normalized_count)) + geom_histogram(fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-26T23:59:52.527870Z","iopub.execute_input":"2023-09-26T23:59:52.529412Z","iopub.status.idle":"2023-09-27T00:03:11.486317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For consistency, we'll remove this df for now and future versions we'll see if we need to modify it at all, and then re-write changes.\nrm(multiome_train.parquet)\ngc()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:11.489608Z","iopub.execute_input":"2023-09-27T00:03:11.491340Z","iopub.status.idle":"2023-09-27T00:03:15.001681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we have from the competition hosts regarding the **adata_obs_meta.csv**:\n\n* library_id - A unique identifier for each library, which is a measurement made on pooled samples from each row of the plate. All cells from wells on the same row of the same \n* plate will share a library_id.\n* plate_name - A unique ID for all samples from the same plate.\n* well - The well location of the sample on each plate (this is standard across 96 well plate experiments). It is a concatenation of row and col.\n* row - Which row on the plate the sample came from.\n* col - Which column on the plate the sample came from.\n* donor_id - Identifies the donor source of the sample, one of three.\n* cell_type - The annotated cell type of each cell based on RNA expression. This matches the cell_type in the de_train.parquet.\n* cell_id - This is included for consistency with LINCS Connectivity Map metadata, which denotes a cell_id for each cell line.\n* sm_name - The primary name for the (parent) compound (in a standardized representation) as chosen by LINCS. This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n* sm_lincs_id - The global LINCS ID (parent) compound (in a standardized representation). This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n* SMILES - Simplified molecular-input line-entry system (SMILES) representations of the compounds used in the experiment. This is a 1D representation of molecular structure. These SMILES are provided by Cellarity based on the specific compounds ordered for this experiment.\n* dose_uM - Dose of the compound in on a micro-molar scale. This maps to the pert_idose field in LINCS.\n* timepoint_hr - Duration of treatment in hours. This maps to the pert_itime field in LINCS.\n* control - Whether this observation was used as a control, True or False.","metadata":{}},{"cell_type":"markdown","source":"## Taking a look at the **library_id** totals","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 15, repr.plot.height =10)\n# Grab the mean\nlibrary_id_mean <- adata_obs_meta %>% group_by(library_id) %>% summarize(totals = n()) %>% pull(totals) %>% mean()\n# Make the plot\nadata_obs_meta %>% group_by(library_id) %>% summarize(totals = n()) %>% ggplot(aes(x=reorder(library_id, totals))) + \ngeom_col(aes(y=totals), fill = 'steelblue') + geom_hline(yintercept = library_id_mean, color= 'red') + coord_flip() + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:15.004517Z","iopub.execute_input":"2023-09-27T00:03:15.006462Z","iopub.status.idle":"2023-09-27T00:03:15.482841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## There are only three options for the **donor_id** variable (0, 1, & 2).","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =6)\n# Make the plot\nadata_obs_meta %>% group_by(donor_id) %>% summarize(totals = n()) %>% ggplot(aes(x=reorder(donor_id, totals))) + \ngeom_col(aes(y=totals), fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:15.487072Z","iopub.execute_input":"2023-09-27T00:03:15.489456Z","iopub.status.idle":"2023-09-27T00:03:15.773509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## As we can see from the below plot, the '**T cells CD4+**' and '**NK cells**' feature most prominently.","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =6)\n# Make the plot\nadata_obs_meta %>% group_by(cell_type) %>% summarize(totals = n()) %>% ggplot(aes(x=reorder(cell_type, totals))) + \ngeom_col(aes(y=totals), fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:15.776541Z","iopub.execute_input":"2023-09-27T00:03:15.778185Z","iopub.status.idle":"2023-09-27T00:03:16.078457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Looking at the distribution of **dose_uM**, we see the **value of '1'** comprises the majority at **187,569**.","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =6)\n# Make the plot\nadata_obs_meta %>% ggplot(aes(x=dose_uM)) + geom_histogram(fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.081377Z","iopub.execute_input":"2023-09-27T00:03:16.083025Z","iopub.status.idle":"2023-09-27T00:03:16.520547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Lastly, **161,223** of the observations were **not in the control**, or **~ 80%** of the observations were **not used**.","metadata":{}},{"cell_type":"code","source":"# Adjust plot size\noptions(repr.plot.width = 8, repr.plot.height =6)\n# Make the plot\nadata_obs_meta %>% group_by(control) %>% summarize(totals = n()) %>% ggplot(aes(x=reorder(control, totals))) + \ngeom_col(aes(y=totals), fill = 'steelblue') + theme_clean()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.524336Z","iopub.execute_input":"2023-09-27T00:03:16.526026Z","iopub.status.idle":"2023-09-27T00:03:16.817502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we have from the competition hosts regarding the **multiome_var_meta.csv**:\n\n* location - This is a feature ID. If the feature_type is Gene Expression then this is a gene symbol. If feature_type is Peaks, then this is the genomic interval of the peak.\n* gene_id - This is an alternative unique feature ID. If the feature_type is Gene Expression then this is an Ensembl Stable Gene ID. If feature_type is Peaks, then this is the genomic interval of the peak.\n* feature_type - Denotes whether the feature is an RNA expression measurement or a Chromatin Accessibility measurement.\n* genome - The genome version used when running CellRanger-Arc\n* interval - The genomic coordinates of each feature on reference genome GRCh38. Genomic coordinates are directly related to the reference genome and include the chromosome name, start position, and end position in the following format: chr1:1234570-1234870.","metadata":{}},{"cell_type":"code","source":"# There are 158205 rows of data\nmultiome_var_meta %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.820410Z","iopub.execute_input":"2023-09-27T00:03:16.821908Z","iopub.status.idle":"2023-09-27T00:03:16.839359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is what our data looks like\nmultiome_var_meta %>% head()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.842118Z","iopub.execute_input":"2023-09-27T00:03:16.843562Z","iopub.status.idle":"2023-09-27T00:03:16.870553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we have from the competition hosts regarding the **multiome_obs_meta.csv**:\n\n* obs_id - Identifier corresponding to that in multiome_train.parquet.\n* cell_type - The annotated cell type of each cell based on RNA expression.\n* donor_id - Identifies the donor source of the sample, one of three.","metadata":{}},{"cell_type":"code","source":"# There are 25551 rows of data\nmultiome_obs_meta %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.873386Z","iopub.execute_input":"2023-09-27T00:03:16.874866Z","iopub.status.idle":"2023-09-27T00:03:16.891824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is what our data looks like\nmultiome_obs_meta %>% head()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.894748Z","iopub.execute_input":"2023-09-27T00:03:16.896366Z","iopub.status.idle":"2023-09-27T00:03:16.924798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what we have from the competition hosts regarding the **id_map.csv**:\n* Identifies the cell_type / sm_name pair to be predicted for the given id.","metadata":{}},{"cell_type":"code","source":"# There are 255 rows of data\nid_map %>% nrow()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.927874Z","iopub.execute_input":"2023-09-27T00:03:16.929738Z","iopub.status.idle":"2023-09-27T00:03:16.946378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is what our data looks like\nid_map %>% head()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.949351Z","iopub.execute_input":"2023-09-27T00:03:16.950848Z","iopub.status.idle":"2023-09-27T00:03:16.977534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***","metadata":{}},{"cell_type":"markdown","source":"## This is what our **sample_submission**.csv file looks like:","metadata":{}},{"cell_type":"code","source":"sample_submission %>% head()","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:16.980362Z","iopub.execute_input":"2023-09-27T00:03:16.981848Z","iopub.status.idle":"2023-09-27T00:03:25.377643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's write a sample **submission**","metadata":{}},{"cell_type":"code","source":"write_csv(sample_submission, 'submission.csv')","metadata":{"execution":{"iopub.status.busy":"2023-09-27T00:03:25.380600Z","iopub.execute_input":"2023-09-27T00:03:25.382153Z","iopub.status.idle":"2023-09-27T00:03:34.073247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### This is just a start & I will continue to update this as time permits.  If it was helpful, please **🔼upvote🔼**!  And if you have ways to make it better, please leave a **🖊️comment🖊️** -- thanks!","metadata":{}}]}