{
  "id": 460128,
  "title": "Convert feature-added table data to image data and use Timm models.",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460128",
  "author_name": "masaishi",
  "post_date": "2023-12-08T00:59:05.760000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>My one approach is to convert feature-added table data to image data and use PyTorch Image Models (Timm).</p>\n<p><a href=\"https://www.kaggle.com/code/masaishi/opscp-m-nb019/notebook\" target=\"_blank\">Code</a></p>\n<p>Private Score<br>\n0.816</p>\n<p>Public Score<br>\n0.620</p>\n<p>Score is bad, but it's my personal favorite.</p>\n<p>I shared other codes to <a href=\"https://github.com/masaishi/kaggle-OPSCP\" target=\"_blank\">kaggle-OPSCP repo on Github</a></p>\n<h2>Feature Engineering:</h2>\n<p>My initial step involved enriching the dataset with calculated statistics for each gene and per 'sm_name' and 'cell_type'. This included mean, standard deviation, minimum, maximum, median, skewness, kurtosis, and mean-to-standard deviation ratios. These new features significantly increased the data dimensions.</p>\n<pre><code> ():\n    \n    stat_df = df.groupby(group_col)[cols].apply(stat_func)\n    stat_df.columns = [  col  cols]\n     stat_df.reset_index().astype({group_col: })\n\n\ncell_type_mean = calculate_statistic(all_de_train, , genes,  x: x.mean(), )\ncell_type_std = calculate_statistic(all_de_train, , genes,  x: x.std(), )\n\n</code></pre>\n<h2>Converting Data to Image Format:</h2>\n<p>I converted the data into an image format to handle this massive feature space. This conversion was necesally for applying image-based models.</p>\n<pre><code> ():\n        square_side = (math.ceil(math.sqrt(data.shape[])))\n        padding_size = square_side **  - data.shape[]\n\n        data_padded = np.pad(data, ((, ), (, padding_size)), , constant_values=)\n\n        \n        data_reshaped = data_padded.reshape(-, square_side, square_side, )\n\n        \n        data_rgb = np.repeat(data_reshaped, , -)\n\n         data_rgb\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6426215%2F984375a0877013033e4cbcc82405ffb3%2Foutput.png?generation=1701996123906415&amp;alt=media\" alt=\"Sample image\"></p>\n<h2>Fine-Tuning PyTorch Image Models (timm):</h2>\n<p>I utilized the PyTorch Image Models library, which offers many pre-trained models. I fine-tuned these models, such as efficientnet_v2_l and regnet_y_800mf.</p>",
  "messages": [
    {
      "id": 2553029,
      "postDate": "2023-12-08T00:59:05.760Z",
      "content": "<p>My one approach is to convert feature-added table data to image data and use PyTorch Image Models (Timm).</p>\n<p><a href=\"https://www.kaggle.com/code/masaishi/opscp-m-nb019/notebook\" target=\"_blank\">Code</a></p>\n<p>Private Score<br>\n0.816</p>\n<p>Public Score<br>\n0.620</p>\n<p>Score is bad, but it's my personal favorite.</p>\n<p>I shared other codes to <a href=\"https://github.com/masaishi/kaggle-OPSCP\" target=\"_blank\">kaggle-OPSCP repo on Github</a></p>\n<h2>Feature Engineering:</h2>\n<p>My initial step involved enriching the dataset with calculated statistics for each gene and per 'sm_name' and 'cell_type'. This included mean, standard deviation, minimum, maximum, median, skewness, kurtosis, and mean-to-standard deviation ratios. These new features significantly increased the data dimensions.</p>\n<pre><code> ():\n    \n    stat_df = df.groupby(group_col)[cols].apply(stat_func)\n    stat_df.columns = [  col  cols]\n     stat_df.reset_index().astype({group_col: })\n\n\ncell_type_mean = calculate_statistic(all_de_train, , genes,  x: x.mean(), )\ncell_type_std = calculate_statistic(all_de_train, , genes,  x: x.std(), )\n\n</code></pre>\n<h2>Converting Data to Image Format:</h2>\n<p>I converted the data into an image format to handle this massive feature space. This conversion was necesally for applying image-based models.</p>\n<pre><code> ():\n        square_side = (math.ceil(math.sqrt(data.shape[])))\n        padding_size = square_side **  - data.shape[]\n\n        data_padded = np.pad(data, ((, ), (, padding_size)), , constant_values=)\n\n        \n        data_reshaped = data_padded.reshape(-, square_side, square_side, )\n\n        \n        data_rgb = np.repeat(data_reshaped, , -)\n\n         data_rgb\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6426215%2F984375a0877013033e4cbcc82405ffb3%2Foutput.png?generation=1701996123906415&amp;alt=media\" alt=\"Sample image\"></p>\n<h2>Fine-Tuning PyTorch Image Models (timm):</h2>\n<p>I utilized the PyTorch Image Models library, which offers many pre-trained models. I fine-tuned these models, such as efficientnet_v2_l and regnet_y_800mf.</p>",
      "rawMarkdown": "My one approach is to convert feature-added table data to image data and use PyTorch Image Models (Timm).\n\n[Code](https://www.kaggle.com/code/masaishi/opscp-m-nb019/notebook)\n\nPrivate Score\n0.816\n\nPublic Score\n0.620\n\nScore is bad, but it's my personal favorite.\n\nI shared other codes to [kaggle-OPSCP repo on Github](https://github.com/masaishi/kaggle-OPSCP)\n\n## Feature Engineering:\nMy initial step involved enriching the dataset with calculated statistics for each gene and per 'sm_name' and 'cell_type'. This included mean, standard deviation, minimum, maximum, median, skewness, kurtosis, and mean-to-standard deviation ratios. These new features significantly increased the data dimensions.\n```python\ndef calculate_statistic(df, group_col, cols, stat_func, stat_name):\n    \"\"\"\n    Function to calculate statistics for given columns grouped by a specific column.\n    \"\"\"\n    stat_df = df.groupby(group_col)[cols].apply(stat_func)\n    stat_df.columns = [f'{group_col}_{stat_name}_{col}' for col in cols]\n    return stat_df.reset_index().astype({group_col: str})\n\n# Example usage for calculating various statistics for 'cell_type'\ncell_type_mean = calculate_statistic(all_de_train, \"cell_type\", genes, lambda x: x.mean(), 'mean')\ncell_type_std = calculate_statistic(all_de_train, \"cell_type\", genes, lambda x: x.std(), 'std')\n# Additional statistics like min, max, median, skew, kurtosis, and ratio_mean_std are also calculated\n```\n\n## Converting Data to Image Format:\nI converted the data into an image format to handle this massive feature space. This conversion was necesally for applying image-based models.\n```python\ndef convert_to_image_format(data, output_size=(224, 224)):\n\t\tsquare_side = int(math.ceil(math.sqrt(data.shape[1])))\n\t\tpadding_size = square_side ** 2 - data.shape[1]\n\t\t\n\t\tdata_padded = np.pad(data, ((0, 0), (0, padding_size)), 'constant', constant_values=0)\n\t\t\n\t\t# We are reshaping to [N, H, W, C] because the final torch tensor needs to be [N, C, H, W]\n\t\tdata_reshaped = data_padded.reshape(-1, square_side, square_side, 1)\n\t\t\n\t\t# Expand the last dimension to three channels by repeating the data\n\t\tdata_rgb = np.repeat(data_reshaped, 3, -1)\n\t\t\n\t\treturn data_rgb\n```\n![Sample image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6426215%2F984375a0877013033e4cbcc82405ffb3%2Foutput.png?generation=1701996123906415&alt=media)\n\n## Fine-Tuning PyTorch Image Models (timm):\nI utilized the PyTorch Image Models library, which offers many pre-trained models. I fine-tuned these models, such as efficientnet_v2_l and regnet_y_800mf.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2553029": "My one approach is to convert feature-added table data to image data and use PyTorch Image Models (Timm).\n\n[Code](https://www.kaggle.com/code/masaishi/opscp-m-nb019/notebook)\n\nPrivate Score\n0.816\n\nPublic Score\n0.620\n\nScore is bad, but it's my personal favorite.\n\nI shared other codes to [kaggle-OPSCP repo on Github](https://github.com/masaishi/kaggle-OPSCP)\n\n## Feature Engineering:\nMy initial step involved enriching the dataset with calculated statistics for each gene and per 'sm_name' and 'cell_type'. This included mean, standard deviation, minimum, maximum, median, skewness, kurtosis, and mean-to-standard deviation ratios. These new features significantly increased the data dimensions.\n```python\ndef calculate_statistic(df, group_col, cols, stat_func, stat_name):\n    \"\"\"\n    Function to calculate statistics for given columns grouped by a specific column.\n    \"\"\"\n    stat_df = df.groupby(group_col)[cols].apply(stat_func)\n    stat_df.columns = [f'{group_col}_{stat_name}_{col}' for col in cols]\n    return stat_df.reset_index().astype({group_col: str})\n\n# Example usage for calculating various statistics for 'cell_type'\ncell_type_mean = calculate_statistic(all_de_train, \"cell_type\", genes, lambda x: x.mean(), 'mean')\ncell_type_std = calculate_statistic(all_de_train, \"cell_type\", genes, lambda x: x.std(), 'std')\n# Additional statistics like min, max, median, skew, kurtosis, and ratio_mean_std are also calculated\n```\n\n## Converting Data to Image Format:\nI converted the data into an image format to handle this massive feature space. This conversion was necesally for applying image-based models.\n```python\ndef convert_to_image_format(data, output_size=(224, 224)):\n\t\tsquare_side = int(math.ceil(math.sqrt(data.shape[1])))\n\t\tpadding_size = square_side ** 2 - data.shape[1]\n\t\t\n\t\tdata_padded = np.pad(data, ((0, 0), (0, padding_size)), 'constant', constant_values=0)\n\t\t\n\t\t# We are reshaping to [N, H, W, C] because the final torch tensor needs to be [N, C, H, W]\n\t\tdata_reshaped = data_padded.reshape(-1, square_side, square_side, 1)\n\t\t\n\t\t# Expand the last dimension to three channels by repeating the data\n\t\tdata_rgb = np.repeat(data_reshaped, 3, -1)\n\t\t\n\t\treturn data_rgb\n```\n![Sample image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6426215%2F984375a0877013033e4cbcc82405ffb3%2Foutput.png?generation=1701996123906415&alt=media)\n\n## Fine-Tuning PyTorch Image Models (timm):\nI utilized the PyTorch Image Models library, which offers many pre-trained models. I fine-tuned these models, such as efficientnet_v2_l and regnet_y_800mf."
  }
}