{
  "id": 460960,
  "title": "SMILES😘 Data Science Competition: A Deep Dive",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460960",
  "author_name": "Jocelyn Dumlao",
  "post_date": "2023-12-12T01:40:36.672000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Introduction:</h1>\n<p>In the pursuit of advancing single-cell data science and catalyzing drug discovery, the SMILES😘 competition introduces a groundbreaking dataset. Developed for the competition, this dataset features human peripheral blood mononuclear cells (PBMCs) and includes 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map. The experiment, replicated in three healthy human donors, provides meticulous measurements of single-cell gene expression profiles following 24 hours of treatment.</p>\n<h1>Exploratory Data Analysis (EDA):</h1>\n<h2>Distribution of Donors:</h2>\n<ul>\n<li>Visualized the frequency of each donor in the dataset using a count plot.</li>\n<li>The plot offers an overview of how samples are distributed across different donors.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F75a8c97f33b9542e9d1e14f9ff873d54%2FDistDonors.png?generation=1702342943516621&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Distribution of Dose (uM):</h2>\n<ul>\n<li>Utilized a histogram to showcase the distribution of doses in microMolarity.</li>\n<li>This plot provides insight into how doses are spread across the dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F0fa42111ed6b467ad42bbc28cb55a711%2FDist-Dose(uM).png?generation=1702344900095659&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Average Dose at Different Timepoints:</h2>\n<ul>\n<li>Presented a bar plot illustrating the average dose (uM) at different timepoints.</li>\n<li>This plot helps visualize how doses vary with time.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F70275fa6f5b4d479338c4d4d9555932d%2FAveDoseAtDiff.png?generation=1702343097669094&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Distribution of Dose (uM) Summary:</h2>\n<ul>\n<li>Provided a summary of the distribution of doses using a box plot.</li>\n<li>The plot includes quartiles, median, and potential outliers.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Ff826df73e40e1c01448aba6a23d33bfc%2FDist-dose.png?generation=1702343026198358&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Dose (uM) and Timepoint (hours) Relationship:</h2>\n<ul>\n<li>Visualized the relationship between dose (uM) and timepoint (hours) using a pair plot.</li>\n<li>The diagonal shows kernel density estimates.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F54fb101ac78660f4934bf3b46bbd5ab4%2FDose%20and%20Timep.png?generation=1702345034468959&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Distribution of Cell Types:</h2>\n<ul>\n<li>Displayed a count plot to visualize the occurrences of each cell type.</li>\n<li>This plot gives an overview of the distribution of cell types in the dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F8e094f4cf2aa946d695e1dc43eaf2a56%2FDistCellType.png?generation=1702343745263825&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Correlation Heatmap:</h2>\n<ul>\n<li>Presented a heatmap visualizing the correlation between dose (uM) and timepoint (hours).</li>\n<li>Values closer to 1 indicate a stronger positive correlation.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F33c5b64b658ff3e6749524c57f7993e9%2FCorHeatmap.png?generation=1702343800299076&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Control Distribution:</h2>\n<ul>\n<li>Represented the distribution of 'control' values (True or False) using a pie chart.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F601d850c48e2ae4a88c4c3107f6e51f9%2FDist-pie.png?generation=1702343949783969&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Dose (uM) vs. Timepoint (hours):</h2>\n<ul>\n<li>Illustrated the relationship between dose (uM) and timepoint (hours) with a scatter plot.</li>\n<li>Different colors represent different control values.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F3e5768bd41279903c51770bb626b5998%2FDose%20vs%20Timepoint.png?generation=1702344027609287&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Dose (uM) by Control:</h2>\n<ul>\n<li>Used a violin plot to show the distribution of doses for each control category.</li>\n<li>Allows for a comparison of dose distributions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F6ace487693a5a1c21cfd07bca506abd6%2FDosebycontrol.png?generation=1702344109580758&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h1>Feature Extraction:</h1>\n<h2>Morgan Fingerprints:</h2>\n<ul>\n<li>Introduced Morgan fingerprints as a numerical representation suitable for T-SNE and PCA.</li>\n<li>Provided insights into the structure and application of Morgan fingerprints in chemoinformatics and drug discovery.</li>\n</ul>\n<h2>T-SNE and PCA Visualization:</h2>\n<p>Applied T-SNE and PCA to the features for dimensionality reduction and visualization.</p>\n<h2>PCA Visualization:</h2>\n<ul>\n<li>Utilized PCA for dimensionality reduction and visualization.</li>\n<li>The scatter plot displays the reduced features in two dimensions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fb8c9af44df865b3fb9e7da392cc29e8e%2FPCA-vis.png?generation=1702344215751186&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Molecule Visualizations:</h2>\n<ul>\n<li>Generated images for a subset of molecules using RDKit.</li>\n</ul>\n<h2>Molecule Visualizations:</h2>\n<ul>\n<li>Added a new column 'Molecule' with RDKit Mol objects.</li>\n<li>Visualized a subset of molecules with RDKit.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fe6bb7c2f447b811fe06151aad047f3ac%2FRIDkit.png?generation=1702344347921123&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h1>Prediction Results:</h1>\n<h2>Prediction Results for Various Models:</h2>\n<ul>\n<li>Presented results for linear regression, logistic regression, decision tree, random forest, SVM, KNN, K-Means, naive Bayes, and neural network models.</li>\n<li>Included metrics such as mean absolute error, mean squared error, and R-squared (R2).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F58a4cc1d79d614594e8e98f1bb411ebd%2FResults.png?generation=1702344308542403&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h1>Conclusion:</h1>\n<p>In conclusion, this comprehensive analysis of the SMILES😘 competition dataset provides valuable insights into the distribution of donors, doses, timepoints, and cell types. The exploration of Morgan fingerprints, T-SNE, PCA, and molecule visualizations adds depth to the understanding of the dataset. The prediction results offer a benchmark for various models, highlighting their performance in the context of the competition's objectives. This write-up serves as a resource for researchers and data scientists engaged in single-cell data analysis and drug discovery.</p>\n<p>SMILES😘: <a href=\"https://www.kaggle.com/code/jocelyndumlao/smiles/notebook\" target=\"_blank\">https://www.kaggle.com/code/jocelyndumlao/smiles/notebook</a></p>",
  "messages": [
    {
      "id": 2558204,
      "postDate": "2023-12-12T01:40:36.673Z",
      "content": "<h1>Introduction:</h1>\n<p>In the pursuit of advancing single-cell data science and catalyzing drug discovery, the SMILES😘 competition introduces a groundbreaking dataset. Developed for the competition, this dataset features human peripheral blood mononuclear cells (PBMCs) and includes 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map. The experiment, replicated in three healthy human donors, provides meticulous measurements of single-cell gene expression profiles following 24 hours of treatment.</p>\n<h1>Exploratory Data Analysis (EDA):</h1>\n<h2>Distribution of Donors:</h2>\n<ul>\n<li>Visualized the frequency of each donor in the dataset using a count plot.</li>\n<li>The plot offers an overview of how samples are distributed across different donors.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F75a8c97f33b9542e9d1e14f9ff873d54%2FDistDonors.png?generation=1702342943516621&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Distribution of Dose (uM):</h2>\n<ul>\n<li>Utilized a histogram to showcase the distribution of doses in microMolarity.</li>\n<li>This plot provides insight into how doses are spread across the dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F0fa42111ed6b467ad42bbc28cb55a711%2FDist-Dose(uM).png?generation=1702344900095659&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Average Dose at Different Timepoints:</h2>\n<ul>\n<li>Presented a bar plot illustrating the average dose (uM) at different timepoints.</li>\n<li>This plot helps visualize how doses vary with time.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F70275fa6f5b4d479338c4d4d9555932d%2FAveDoseAtDiff.png?generation=1702343097669094&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Distribution of Dose (uM) Summary:</h2>\n<ul>\n<li>Provided a summary of the distribution of doses using a box plot.</li>\n<li>The plot includes quartiles, median, and potential outliers.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Ff826df73e40e1c01448aba6a23d33bfc%2FDist-dose.png?generation=1702343026198358&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Dose (uM) and Timepoint (hours) Relationship:</h2>\n<ul>\n<li>Visualized the relationship between dose (uM) and timepoint (hours) using a pair plot.</li>\n<li>The diagonal shows kernel density estimates.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F54fb101ac78660f4934bf3b46bbd5ab4%2FDose%20and%20Timep.png?generation=1702345034468959&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Distribution of Cell Types:</h2>\n<ul>\n<li>Displayed a count plot to visualize the occurrences of each cell type.</li>\n<li>This plot gives an overview of the distribution of cell types in the dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F8e094f4cf2aa946d695e1dc43eaf2a56%2FDistCellType.png?generation=1702343745263825&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Correlation Heatmap:</h2>\n<ul>\n<li>Presented a heatmap visualizing the correlation between dose (uM) and timepoint (hours).</li>\n<li>Values closer to 1 indicate a stronger positive correlation.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F33c5b64b658ff3e6749524c57f7993e9%2FCorHeatmap.png?generation=1702343800299076&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Control Distribution:</h2>\n<ul>\n<li>Represented the distribution of 'control' values (True or False) using a pie chart.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F601d850c48e2ae4a88c4c3107f6e51f9%2FDist-pie.png?generation=1702343949783969&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Dose (uM) vs. Timepoint (hours):</h2>\n<ul>\n<li>Illustrated the relationship between dose (uM) and timepoint (hours) with a scatter plot.</li>\n<li>Different colors represent different control values.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F3e5768bd41279903c51770bb626b5998%2FDose%20vs%20Timepoint.png?generation=1702344027609287&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Dose (uM) by Control:</h2>\n<ul>\n<li>Used a violin plot to show the distribution of doses for each control category.</li>\n<li>Allows for a comparison of dose distributions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F6ace487693a5a1c21cfd07bca506abd6%2FDosebycontrol.png?generation=1702344109580758&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h1>Feature Extraction:</h1>\n<h2>Morgan Fingerprints:</h2>\n<ul>\n<li>Introduced Morgan fingerprints as a numerical representation suitable for T-SNE and PCA.</li>\n<li>Provided insights into the structure and application of Morgan fingerprints in chemoinformatics and drug discovery.</li>\n</ul>\n<h2>T-SNE and PCA Visualization:</h2>\n<p>Applied T-SNE and PCA to the features for dimensionality reduction and visualization.</p>\n<h2>PCA Visualization:</h2>\n<ul>\n<li>Utilized PCA for dimensionality reduction and visualization.</li>\n<li>The scatter plot displays the reduced features in two dimensions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fb8c9af44df865b3fb9e7da392cc29e8e%2FPCA-vis.png?generation=1702344215751186&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h2>Molecule Visualizations:</h2>\n<ul>\n<li>Generated images for a subset of molecules using RDKit.</li>\n</ul>\n<h2>Molecule Visualizations:</h2>\n<ul>\n<li>Added a new column 'Molecule' with RDKit Mol objects.</li>\n<li>Visualized a subset of molecules with RDKit.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fe6bb7c2f447b811fe06151aad047f3ac%2FRIDkit.png?generation=1702344347921123&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h1>Prediction Results:</h1>\n<h2>Prediction Results for Various Models:</h2>\n<ul>\n<li>Presented results for linear regression, logistic regression, decision tree, random forest, SVM, KNN, K-Means, naive Bayes, and neural network models.</li>\n<li>Included metrics such as mean absolute error, mean squared error, and R-squared (R2).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F58a4cc1d79d614594e8e98f1bb411ebd%2FResults.png?generation=1702344308542403&amp;alt=media\" alt=\"image\"></li>\n</ul>\n<h1>Conclusion:</h1>\n<p>In conclusion, this comprehensive analysis of the SMILES😘 competition dataset provides valuable insights into the distribution of donors, doses, timepoints, and cell types. The exploration of Morgan fingerprints, T-SNE, PCA, and molecule visualizations adds depth to the understanding of the dataset. The prediction results offer a benchmark for various models, highlighting their performance in the context of the competition's objectives. This write-up serves as a resource for researchers and data scientists engaged in single-cell data analysis and drug discovery.</p>\n<p>SMILES😘: <a href=\"https://www.kaggle.com/code/jocelyndumlao/smiles/notebook\" target=\"_blank\">https://www.kaggle.com/code/jocelyndumlao/smiles/notebook</a></p>",
      "rawMarkdown": "# Introduction:\n\nIn the pursuit of advancing single-cell data science and catalyzing drug discovery, the SMILES😘 competition introduces a groundbreaking dataset. Developed for the competition, this dataset features human peripheral blood mononuclear cells (PBMCs) and includes 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map. The experiment, replicated in three healthy human donors, provides meticulous measurements of single-cell gene expression profiles following 24 hours of treatment.\n\n# Exploratory Data Analysis (EDA):\n\n## Distribution of Donors:\n* Visualized the frequency of each donor in the dataset using a count plot.\n* The plot offers an overview of how samples are distributed across different donors.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F75a8c97f33b9542e9d1e14f9ff873d54%2FDistDonors.png?generation=1702342943516621&alt=media)\n\n## Distribution of Dose (uM):\n* Utilized a histogram to showcase the distribution of doses in microMolarity.\n* This plot provides insight into how doses are spread across the dataset.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F0fa42111ed6b467ad42bbc28cb55a711%2FDist-Dose(uM).png?generation=1702344900095659&alt=media)\n## Average Dose at Different Timepoints:\n* Presented a bar plot illustrating the average dose (uM) at different timepoints.\n* This plot helps visualize how doses vary with time.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F70275fa6f5b4d479338c4d4d9555932d%2FAveDoseAtDiff.png?generation=1702343097669094&alt=media)\n\n## Distribution of Dose (uM) Summary:\n* Provided a summary of the distribution of doses using a box plot.\n* The plot includes quartiles, median, and potential outliers.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Ff826df73e40e1c01448aba6a23d33bfc%2FDist-dose.png?generation=1702343026198358&alt=media)\n\n## Dose (uM) and Timepoint (hours) Relationship:\n* Visualized the relationship between dose (uM) and timepoint (hours) using a pair plot.\n* The diagonal shows kernel density estimates.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F54fb101ac78660f4934bf3b46bbd5ab4%2FDose%20and%20Timep.png?generation=1702345034468959&alt=media)\n\n## Distribution of Cell Types:\n* Displayed a count plot to visualize the occurrences of each cell type.\n* This plot gives an overview of the distribution of cell types in the dataset.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F8e094f4cf2aa946d695e1dc43eaf2a56%2FDistCellType.png?generation=1702343745263825&alt=media)\n\n## Correlation Heatmap:\n* Presented a heatmap visualizing the correlation between dose (uM) and timepoint (hours).\n* Values closer to 1 indicate a stronger positive correlation.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F33c5b64b658ff3e6749524c57f7993e9%2FCorHeatmap.png?generation=1702343800299076&alt=media)\n\n## Control Distribution:\n* Represented the distribution of 'control' values (True or False) using a pie chart.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F601d850c48e2ae4a88c4c3107f6e51f9%2FDist-pie.png?generation=1702343949783969&alt=media)\n## Dose (uM) vs. Timepoint (hours):\n* Illustrated the relationship between dose (uM) and timepoint (hours) with a scatter plot.\n* Different colors represent different control values.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F3e5768bd41279903c51770bb626b5998%2FDose%20vs%20Timepoint.png?generation=1702344027609287&alt=media)\n\n## Dose (uM) by Control:\n* Used a violin plot to show the distribution of doses for each control category.\n* Allows for a comparison of dose distributions.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F6ace487693a5a1c21cfd07bca506abd6%2FDosebycontrol.png?generation=1702344109580758&alt=media)\n\n# Feature Extraction:\n## Morgan Fingerprints:\n* Introduced Morgan fingerprints as a numerical representation suitable for T-SNE and PCA.\n* Provided insights into the structure and application of Morgan fingerprints in chemoinformatics and drug discovery.\n\n## T-SNE and PCA Visualization:\nApplied T-SNE and PCA to the features for dimensionality reduction and visualization.\n## PCA Visualization:\n* Utilized PCA for dimensionality reduction and visualization.\n* The scatter plot displays the reduced features in two dimensions.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fb8c9af44df865b3fb9e7da392cc29e8e%2FPCA-vis.png?generation=1702344215751186&alt=media)\n\n## Molecule Visualizations:\n* Generated images for a subset of molecules using RDKit.\n## Molecule Visualizations:\n* Added a new column 'Molecule' with RDKit Mol objects.\n* Visualized a subset of molecules with RDKit.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fe6bb7c2f447b811fe06151aad047f3ac%2FRIDkit.png?generation=1702344347921123&alt=media)\n# Prediction Results:\n## Prediction Results for Various Models:\n* Presented results for linear regression, logistic regression, decision tree, random forest, SVM, KNN, K-Means, naive Bayes, and neural network models.\n* Included metrics such as mean absolute error, mean squared error, and R-squared (R2).\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F58a4cc1d79d614594e8e98f1bb411ebd%2FResults.png?generation=1702344308542403&alt=media)\n\n# Conclusion:\n\nIn conclusion, this comprehensive analysis of the SMILES😘 competition dataset provides valuable insights into the distribution of donors, doses, timepoints, and cell types. The exploration of Morgan fingerprints, T-SNE, PCA, and molecule visualizations adds depth to the understanding of the dataset. The prediction results offer a benchmark for various models, highlighting their performance in the context of the competition's objectives. This write-up serves as a resource for researchers and data scientists engaged in single-cell data analysis and drug discovery.\n\nSMILES😘: https://www.kaggle.com/code/jocelyndumlao/smiles/notebook\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2558204": "# Introduction:\n\nIn the pursuit of advancing single-cell data science and catalyzing drug discovery, the SMILES😘 competition introduces a groundbreaking dataset. Developed for the competition, this dataset features human peripheral blood mononuclear cells (PBMCs) and includes 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map. The experiment, replicated in three healthy human donors, provides meticulous measurements of single-cell gene expression profiles following 24 hours of treatment.\n\n# Exploratory Data Analysis (EDA):\n\n## Distribution of Donors:\n* Visualized the frequency of each donor in the dataset using a count plot.\n* The plot offers an overview of how samples are distributed across different donors.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F75a8c97f33b9542e9d1e14f9ff873d54%2FDistDonors.png?generation=1702342943516621&alt=media)\n\n## Distribution of Dose (uM):\n* Utilized a histogram to showcase the distribution of doses in microMolarity.\n* This plot provides insight into how doses are spread across the dataset.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F0fa42111ed6b467ad42bbc28cb55a711%2FDist-Dose(uM).png?generation=1702344900095659&alt=media)\n## Average Dose at Different Timepoints:\n* Presented a bar plot illustrating the average dose (uM) at different timepoints.\n* This plot helps visualize how doses vary with time.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F70275fa6f5b4d479338c4d4d9555932d%2FAveDoseAtDiff.png?generation=1702343097669094&alt=media)\n\n## Distribution of Dose (uM) Summary:\n* Provided a summary of the distribution of doses using a box plot.\n* The plot includes quartiles, median, and potential outliers.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Ff826df73e40e1c01448aba6a23d33bfc%2FDist-dose.png?generation=1702343026198358&alt=media)\n\n## Dose (uM) and Timepoint (hours) Relationship:\n* Visualized the relationship between dose (uM) and timepoint (hours) using a pair plot.\n* The diagonal shows kernel density estimates.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F54fb101ac78660f4934bf3b46bbd5ab4%2FDose%20and%20Timep.png?generation=1702345034468959&alt=media)\n\n## Distribution of Cell Types:\n* Displayed a count plot to visualize the occurrences of each cell type.\n* This plot gives an overview of the distribution of cell types in the dataset.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F8e094f4cf2aa946d695e1dc43eaf2a56%2FDistCellType.png?generation=1702343745263825&alt=media)\n\n## Correlation Heatmap:\n* Presented a heatmap visualizing the correlation between dose (uM) and timepoint (hours).\n* Values closer to 1 indicate a stronger positive correlation.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F33c5b64b658ff3e6749524c57f7993e9%2FCorHeatmap.png?generation=1702343800299076&alt=media)\n\n## Control Distribution:\n* Represented the distribution of 'control' values (True or False) using a pie chart.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F601d850c48e2ae4a88c4c3107f6e51f9%2FDist-pie.png?generation=1702343949783969&alt=media)\n## Dose (uM) vs. Timepoint (hours):\n* Illustrated the relationship between dose (uM) and timepoint (hours) with a scatter plot.\n* Different colors represent different control values.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F3e5768bd41279903c51770bb626b5998%2FDose%20vs%20Timepoint.png?generation=1702344027609287&alt=media)\n\n## Dose (uM) by Control:\n* Used a violin plot to show the distribution of doses for each control category.\n* Allows for a comparison of dose distributions.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F6ace487693a5a1c21cfd07bca506abd6%2FDosebycontrol.png?generation=1702344109580758&alt=media)\n\n# Feature Extraction:\n## Morgan Fingerprints:\n* Introduced Morgan fingerprints as a numerical representation suitable for T-SNE and PCA.\n* Provided insights into the structure and application of Morgan fingerprints in chemoinformatics and drug discovery.\n\n## T-SNE and PCA Visualization:\nApplied T-SNE and PCA to the features for dimensionality reduction and visualization.\n## PCA Visualization:\n* Utilized PCA for dimensionality reduction and visualization.\n* The scatter plot displays the reduced features in two dimensions.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fb8c9af44df865b3fb9e7da392cc29e8e%2FPCA-vis.png?generation=1702344215751186&alt=media)\n\n## Molecule Visualizations:\n* Generated images for a subset of molecules using RDKit.\n## Molecule Visualizations:\n* Added a new column 'Molecule' with RDKit Mol objects.\n* Visualized a subset of molecules with RDKit.\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2Fe6bb7c2f447b811fe06151aad047f3ac%2FRIDkit.png?generation=1702344347921123&alt=media)\n# Prediction Results:\n## Prediction Results for Various Models:\n* Presented results for linear regression, logistic regression, decision tree, random forest, SVM, KNN, K-Means, naive Bayes, and neural network models.\n* Included metrics such as mean absolute error, mean squared error, and R-squared (R2).\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7212220%2F58a4cc1d79d614594e8e98f1bb411ebd%2FResults.png?generation=1702344308542403&alt=media)\n\n# Conclusion:\n\nIn conclusion, this comprehensive analysis of the SMILES😘 competition dataset provides valuable insights into the distribution of donors, doses, timepoints, and cell types. The exploration of Morgan fingerprints, T-SNE, PCA, and molecule visualizations adds depth to the understanding of the dataset. The prediction results offer a benchmark for various models, highlighting their performance in the context of the competition's objectives. This write-up serves as a resource for researchers and data scientists engaged in single-cell data analysis and drug discovery.\n\nSMILES😘: https://www.kaggle.com/code/jocelyndumlao/smiles/notebook\n"
  }
}