JAEN
RESEARCH 04

Bioinformatics Analysis

BIOINFORMATICS ANALYSIS

Combining artificial intelligence with established bioinformatics methods to systematically analyze single-cell multi-omics data.

Research Topics  ›  Bioinformatics Analysis

背景と意義

Background

This research combines artificial intelligence with established bioinformatics methods to systematically analyze single-cell multi-omics data. We target diverse omics layers — single-cell transcriptomics, epigenomics, and spatial transcriptomics — focusing on optimizing analysis algorithms, developing omics-integration models, and building general-purpose computational pipelines. This approach aims to provide more robust and efficient computational tools for multi-layered understanding of complex biological processes and disease research.

The advent of single-cell technologies has enabled life-science research to unravel the complexity of biological systems at the cellular level. scRNA-seq analyzes gene expression profiles, scATAC-seq captures chromatin accessibility to estimate transcriptional regulatory potential, and spatial transcriptomics preserves molecular information together with spatial context within tissue. These advanced methods offer new perspectives for exploring cellular heterogeneity, developmental trajectories, and tissue microenvironments. However, single-cell data is typically high-dimensional, sparse, and noisy, with substantial discrepancies across omics types — challenges that conventional analysis methods cannot fully address. Introducing artificial intelligence and deep learning is essential here, enabling automatic extraction of complex latent features, mitigation of noise and batch effects, and construction of a unified representation space across different omics. This approach advances cell-type identification, reconstruction of developmental processes, inference of gene regulatory networks, and elucidation of spatial tissue structure, offering a new computational paradigm for biology and precision medicine.

技術アプローチ

Methods

Automatic Cell-Type Recognition Model Based on scRNA-seq

Repeatedly using similar scRNA-seq samples to determine cell types is enormously time- and labor-intensive. Efficiently transferring annotation information from a reference dataset to a query dataset enables an innovative strategy for cell-type recognition, with deep learning's strength in high-dimensional feature extraction playing a key technical role. We propose a multi-layer feature-extraction model (Figure 1) that introduces a multi-head attention mechanism to capture complex inter-cellular relationships, substantially improving discrimination accuracy for cell types with subtle differences. Multiple fully-connected layers combined with nonlinear activation modules further strengthen deep, nonlinear feature learning. This method enables high-accuracy automatic annotation of new samples obtained under the same experimental conditions as the reference dataset, while also providing reliable auxiliary information for manual annotation, improving both the efficiency and consistency of cell-type recognition.

Algorithm Development for Single-Cell Multi-Omics Integration Models

The core of multi-omics integration is mapping different omics data into a unified low-dimensional embedding space, achieving efficient information representation and comprehensive downstream analysis. Severe heterogeneity and data incompatibility across omics, however, have long been a major challenge. We propose an integrated computational framework for scRNA-seq and scATAC-seq (Figure 2) that uses two independent variational autoencoders (VAE) and graph autoencoders (GAE) for feature extraction and data reconstruction of each omics layer, and introduces a generative adversarial network (GAN) to effectively mitigate distributional differences between omics, achieving consistent alignment in latent space. This framework enables integrated execution of multiple functions, including cross-omics transfer of cell labels, data-modality conversion, and gene regulatory network inference.

Integrated Analysis of Single-Cell and Spatial Transcriptomics via Deep Learning

Single-cell RNA sequencing (scRNA-seq) analyzes cellular heterogeneity at high resolution and throughput, enabling detailed identification of cell types and states, but lacks spatial information, limiting analysis of cellular arrangement and interactions within tissue. Spatial transcriptomics (ST), meanwhile, preserves the spatial distribution of gene expression but has limits on resolution and the number of detectable genes. Integrating both enables a more precise reconstruction of tissue spatial maps by leveraging the strengths of each technology. We use deep learning to learn a shared latent representation between single-cell and spatial data, building correspondence between cellular characteristics and spatial location, aiming to reconstruct spatial tissue maps at near single-cell resolution. By incorporating explainable analysis, we also explore biological patterns inherent in the data in reverse, elucidating the spatial arrangement and interactions among cells — deepening understanding of the spatial molecular mechanisms involved in tissue development and disease progression.

代表的な研究成果

Selected Works